The product question was narrower than semantic similarity
ZenTree starts with a short representation of each tab, built from its title and useful parts of its URL. MiniLM turns that text into a 384-number embedding. Cosine similarity then tells the extension which vectors are close.
That answers:
Which tabs sound related?
The product needs a slightly different answer:
Which tabs belong to the same task right now?
Those questions overlap, but they are not interchangeable.
For example, these pages can all mention Chrome extensions:
- a developer guide for building an extension;
- a Chrome Web Store page for finding or publishing one;
- a help page for installing and managing extensions.
An embedder can reasonably put them near each other. A person working on ZenTree might still want the developer pages in one group and the user-facing pages in another. The correct decision depends on the task behind the tabs, not only on the words in the tabs.
That distinction was already part of ZenTree's design. MiniLM proposes semantic relationships, while ZenTree's deterministic grouping code combines that signal with other rules before deciding whether a group is safe to show. Fine-tuning MiniLM can change the semantic signal, but it cannot add task context that never appeared in the tab text.
Testing the Idea With Public Data
Instead of trying to simulate tab-history data, I decided to start with public data.
I converted the useful public session sources into triplets instead:
{
"anchor": "",
"positive": "",
"negative": ""
}
The anchor is the current tab being grouped. The positive is a tab that belongs with it. The negative is a confusing tab that should stay out. The first dataset I considered was frankjc2022/semantic-history-search from Hugging Face. It looked close because it contained browser-history-style pages, queries, and relevance judgments. The issue with it was its labels would describe how relevant it was to a search query, not whether two Chrome tabs belonged in the same task group.
I then found two datasets that preserved more of the user’s session and task structure. I used each dataset for a different kind of supervision.
TREC 2012 gave real human search-session structure:
- Queries
- Clicked result titles & URLs
- Task identifiers
Based on this dataset, I organized the public session data I was testing into JSON triplets:
{
"anchor": "hosting pampered chef",
"positive": "pampered chef",
"negative": "soralfun"
}
This way, every search query would be the anchor, a clicked result from the same task as the positive, and a clicked result from another task as the negative.
TREC gave me a useful starting signal. When a user searched for something and clicked a result, I could treat that page as relevant to the query.
Query: "cheap flights to Tokyo"
Click: "Tokyo flight deals"
However it does not explain how different queries relate to each other inside the same browsing session. For example, it does not clearly label whether:
"cheap flights to Tokyo"
"Tokyo hotels"
"buy luggage"
are all part of one task or whether the user drifted into a separate activity.
This ambiguity is something that the second dataset, UMass’s Task-Aware Query Recommendation aimed to solve. It used query contexts from search sessions and marked earlier queries as either onTask or offTask relative to the current query.
In the same Tokyo example, if the current query was “cheap flights to Tokyo” (anchor), an on-task query might be “Tokyo hotels” (positive), while an off-task query might be “best wireless headphones” (negative). That gave me a clearer signal about whether activity stayed within the same task or drifted elsewhere.
This still did not show Chrome tab-group data, but it was closer to ZenTree’s problem than a simple clicked-page label. So instead of favoring one dataset over the other, I kept the TREC examples and added the UMass ones so the final dataset included both query-to-page relevance and explicit task-continuity labels.
Together, these two sources produced 20,827 triplets. Before using them to train MiniLM, I split them by whole research task rather than by individual row:
- 16,224 examples for training
- 4,603 examples held back for evaluation
This was a task-disjoint split, meaning every example from the same research task stayed on the same side. Otherwise, MiniLM could see part of a task during training and then be tested on another example from that same task. That would make the score look better than its performance on genuinely new tasks.
16,224 + 4,603 = 20,827
After creating the TREC + UMass dataset, I wanted to test a different kind of label instead of assuming that search-session data was the only public option. Would browser-history retrieval data teach MiniLM something that TREC + UMass missed?
Mozilla’s public History Search Retrieval dataset gave me that second training path. It contained synthetic browser-history profiles, documents, search queries, and query-to-page relevance links. Synthetic here describes Mozilla’s browser profiles, not a set of questions I invented. Mozilla provided the profiles and queries; I used its documents and relevance links to derive triplets.
For each relevance link, I turned the search query into the anchor, a relevant history page into the positive, and another non-relevant page from the same profile into the negative. That created 37,144 local triplets:
- 31,273 examples were used for training
- 5,871 examples were held out for testing
I used those rows to fine-tune a separate MiniLM model. I did not merge them into the TREC + UMass training set.

The Mozilla training path contained more examples, but that does not mean its labels were automatically closer to ZenTree’s task-grouping problem. Mozilla taught query-to-history relevance. TREC + UMass taught clicked-page relevance and on-task versus off-task query relationships.
Comparing The Public-Data Models
I ended up with three models to compare: the original MiniLM, a MiniLM fine-tuned on TREC + UMass, and a MiniLM fine-tuned on Mozilla. A better result had to name the test, because the tests answered different questions.
The Mozilla model and the TREC + UMass model started from the same base MiniLM, but their training labels meant different things:
- TREC + UMass: clicked-page relevance plus on-task and off-task query labels
- Mozilla: query-to-history relevance labels
The first test was Mozilla’s held-out history-search data. It compared the original MiniLM with the Mozilla-trained MiniLM on the same query-to-history task that produced the Mozilla labels. The Mozilla-trained model moved from 78.59% to 83.82%, an increase of 5.23 percentage points. That was a real improvement on Mozilla’s retrieval task. It did not yet tell me whether Mozilla had learned ZenTree’s grouping decision.
The direct comparison happened on the same 4,603 held-out TREC + UMass rows. Original MiniLM scored 99.02%, TREC + UMass scored 98.91%, and Mozilla scored 99.15%. Mozilla was 0.13 percentage points above the base model and 0.24 points above the TREC + UMass model. The score was already very high, so those differences were small.
The figure below keeps the two result types separate. Its left panel is Mozilla’s own history-search test; its right panel uses the same held-out TREC + UMass rows for all three models.

The harder check was a 30-case task-boundary worksheet that I kept out of all public-data training. All three models scored 25 out of 30. Each negative was written to sound related to the anchor while still belonging to another immediate task. This was closer to ZenTree’s grouping decision, but it was still an author-created synthetic worksheet. The tie meant neither public-data fine-tune gave a clear improvement on that closer test.
Training The TREC + UMass MiniLM Model
MiniLM is an encoder, not a chat model. The training job did not teach it to write a response. It changed its embedding space so an anchor tab would be closer to its positive tab than its negative tab.
I used CachedMultipleNegativesRankingLoss. In each batch, the positive tabs from other triplets also act as additional negatives, so the model sees more confusing alternatives without needing new rows. I allowed every MiniLM weight to update. The run made one pass through the 16,224 TREC + UMass training triplets on an NVIDIA A100 and took about 31 seconds.
Evaluation used the 4,603 held-out triplets. Held-out accuracy asks whether the anchor was more similar to its positive than to its negative. Mean similarity gap is the average difference between those two cosine scores. Higher is better.
| Model | Held-out accuracy | Mean similarity gap |
|---|
| Original MiniLM | 99.02% | 0.6344 |
| TREC + UMass fine-tune | 98.91% | 0.6216 |
On that held-out test, the TREC + UMass model fell from 99.02% to 98.91%. Its mean similarity gap also fell from 0.6344 to 0.6216. The change was small, but it went in the wrong direction, so I did not deploy it.
This was still useful. It confirmed that the trainer, evaluation code, and browser input cleanup could work together. It also showed the label problem clearly: a page that matches a query, or a query marked onTask, is not automatically another tab that belongs in the same current Chrome group.
I trained it on the mistake I wanted to fix
The public-data models did not fix the task-boundary mistake I cared about, so I went back to the five worksheet rows that the original MiniLM got wrong. In those cases, the wrong tab shared a broad topic with the anchor, while the correct tab helped complete the next step of the same task.
Examples included:
- a MacBook dock grouped with MacBook battery repair instead of monitor compatibility;
- jazz practice grouped with a general music page instead of a metronome guide;
- a regression homework tab pulled toward a general pandas page instead of the relevant Kaggle material;
- a basmati-rice task confused with a different cooking or shopping task.
I wrote 25 new triplets around those failure patterns. The original 30 worksheet rows stayed out of the training file. The point was to see whether a small, deliberate set of examples could change those exact errors without letting the model memorize the test.
On the original worksheet, the targeted model improved from 25/30 to 28/30. It corrected three of the five base-model mistakes. That was worth following up, but the worksheet and new triplets were both written by me, so I treated this as a smoke test.
To check that result with different wording, I built a second 30-case challenge set. None of its rows were copied from the targeted triplets.
| Model | Fresh challenge result | Mean similarity gap |
|---|
| Base MiniLM | 19/30, 63.33% | 0.0725 |
| Targeted-task model | 21/30, 70.00% | 0.1041 |
This was the strongest training result, but it did not become a deployment result. The targeted model improved on two author-created tests. It did not prove that it could infer a real person’s browsing task from their open tabs.
Running the trained model inside ZenTree
The model had to work in two different environments. Training happened in Python on the GPU server. ZenTree runs inside Chrome, so the candidate had to produce the same embeddings through JavaScript and WebAssembly.

The diagram separates the two paths. Anchor, positive, and negative triplets only exist while training changes MiniLM’s weights. In Chrome, ZenTree takes a tab title and cleaned URL, runs the exported encoder, then gives the similarity result to its existing grouping rules. Fine-tuning changes the model’s signal. It does not replace the rules that decide whether a group is safe.
My first ONNX export still included Sentence Transformers’ pooling head. Its output names did not match what ZenTree’s browser worker expected. ZenTree already does mean pooling itself, so I exported only the underlying transformer. That gave the worker the normal last_hidden_state output it already knows how to use.
I then compared the pooled and normalized PyTorch embeddings with the pooled and normalized ONNX embeddings on representative tab text. The largest absolute difference was about 1.5e-7, and their cosine similarity was effectively 1.0. This checked that the browser candidate was the model I had trained, instead of a different model caused by export behavior.
I packaged that candidate as a roughly 90 MB local ONNX model and gave ZenTree a reversible A/B switch. The original MiniLM stayed the default. The candidate only loaded when selected, and the model files stayed outside Git.
The browser test was useful because it stayed ambiguous
The first browser failure was a version mismatch, not a bad model. ZenTree uses Transformers.js 2.17.2. I had passed the newer dtype: "fp32" option, but this version uses quantized: false to request the full model. Since it ignored the unknown option, it looked for model_quantized.onnx even though the package only contained model.onnx.
After changing that setting and reloading the unpacked extension, the candidate loaded locally and returned groups.
One test tab was named “Install and manage extensions.” The original MiniLM placed it with Chrome Web Store pages. The targeted model placed it with Chrome Developers pages.
Neither output was automatically right. If I had that tab open because I was building ZenTree, the Chrome Developers group was better. If I was looking for or managing extensions, the Chrome Web Store group was better. The title and URL did not tell MiniLM which task I was doing.
That was the point of the browser test. The targeted model changed which semantic relationship it preferred. It did not gain access to the user’s current goal.
What I would build next
The next data should come from actual ZenTree actions, not a different public corpus:
- a group I accepted;
- a group I broke apart;
- a tab I moved into another group;
- a suggestion I rejected;
- tabs that were open together but intentionally kept separate.
Each of those actions says something more direct than a search click. An accepted group says which tabs belonged together. Breaking it apart says the opposite. A moved tab identifies a correction. Before training, the records would need to remove sensitive titles, search terms, login tokens, and document IDs.
The next loop would work like this

The diagram keeps MiniLM in the small, local role it is already good at: finding candidate neighbors. ZenTree’s rules and current context decide whether a group is confident enough to apply. If the answer is unclear, the safer option is to leave the tabs alone or ask.
Those corrections would produce a real task-oriented train/test set, with new tasks held out from training. That is the experiment I would trust before changing ZenTree’s default model.
Sources and related notes
© Kinshuk Goel | All Rights Reserved