Skip to content

Commit 7f6ecbb

Browse files
CodeYogiCoclaude
andauthored
Add draft: Matryoshka embeddings and the e-commerce retrieval funnel (#11)
Post on MRL — nested embeddings, two-stage recall/re-rank pattern, practical latency and index size numbers, and what to watch for. https://claude.ai/code/session_012txymPy1HQ69Ze6v9djRQD Co-authored-by: Claude <noreply@anthropic.com>
1 parent a6d68f6 commit 7f6ecbb

1 file changed

Lines changed: 126 additions & 0 deletions

File tree

Lines changed: 126 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,126 @@
1+
---
2+
date: 2026-06-18
3+
tag: search
4+
title: "Matryoshka embeddings and the e-commerce retrieval funnel"
5+
read: 9 min
6+
deck: "One model, many resolutions. How nested embeddings let you run cheap recall and expensive re-ranking without paying for both upfront."
7+
hidden: true
8+
---
9+
10+
Most teams think of embeddings as atomic. You generate a 768-dimensional vector. You store it. You compare it. You get a result. The whole vector travels through every stage of the pipeline at full cost.
11+
12+
Matryoshka Representation Learning (MRL) breaks that assumption. It gives you one model that produces embeddings you can truncate at multiple resolutions — and every truncation is still meaningful.
13+
14+
The name comes from Russian nesting dolls. The first 64 dimensions of a Matryoshka embedding are a useful representation. The first 128 are more useful. The full 768 are most useful. You choose the resolution based on what the stage of retrieval can afford.
15+
16+
## why this matters at e-commerce scale
17+
18+
A typical e-commerce search pipeline has a retrieval problem that looks like this:
19+
20+
- 10 million product embeddings in the index
21+
- A user query arrives every 50ms at peak
22+
- You need to return ranked results in under 100ms end-to-end
23+
24+
Full 768-dim ANN search over 10M vectors is expensive. You can make it work, but the index is large, the memory footprint is high, and every query touches a lot of data.
25+
26+
The standard solution is a two-stage funnel: fast approximate retrieval to get a candidate set, then more expensive re-ranking to get the final order. The problem is that both stages traditionally use the same full-dimensional embedding. You've paid the storage cost twice and the compute cost at every stage.
27+
28+
> MRL lets you match the embedding resolution to what each stage of the funnel actually needs.
29+
30+
## how the nesting works
31+
32+
Standard embedding training optimizes one loss: how well the full vector represents the input. MRL trains a joint loss across multiple prefix lengths — 64, 128, 256, 512, 768 — simultaneously.
33+
34+
The result is an embedding where every prefix is independently useful. The model has learned to pack the most important information into the first dimensions and progressively add detail as you extend the vector.
35+
36+
```python
37+
# training sketch — loss computed at each granularity
38+
dims = [64, 128, 256, 512, 768]
39+
total_loss = 0
40+
41+
for d in dims:
42+
truncated = embedding[:d] # first d dimensions
43+
loss = contrastive_loss(truncated) # must be meaningful at this size
44+
total_loss += loss
45+
46+
total_loss.backward()
47+
```
48+
49+
The first 64 dimensions aren't random — they've been explicitly trained to be a good representation at that resolution. Truncation is not approximation. It's a deliberate, lower-resolution view of the same information.
50+
51+
## the two-stage retrieval pattern
52+
53+
In e-commerce search, this maps directly onto the retrieval funnel:
54+
55+
**Stage 1 — recall with small embeddings**
56+
57+
Build your ANN index using 64 or 128-dimensional vectors. The index is 6-12× smaller than a full 768-dim index. Retrieval is faster. Memory footprint shrinks. You fetch a large candidate set — say, top 500.
58+
59+
```python
60+
# index built on truncated vectors
61+
index = build_hnsw_index(embeddings[:, :64]) # first 64 dims only
62+
63+
# fast recall — cheap, pulls a large candidate set
64+
candidates = index.search(query_embedding[:64], k=500)
65+
```
66+
67+
**Stage 2 — re-rank with full embeddings**
68+
69+
Take those 500 candidates, load their full 768-dim embeddings, and compute exact similarity scores. Re-rank. Return top 10.
70+
71+
```python
72+
# re-rank candidates using full embeddings
73+
candidate_embeddings = load_full_embeddings(candidates) # 500 × 768
74+
scores = cosine_similarity(query_embedding, candidate_embeddings)
75+
ranked = sorted(zip(candidates, scores), key=lambda x: -x[1])
76+
results = ranked[:10]
77+
```
78+
79+
The expensive computation — full dot products — runs on 500 vectors, not 10 million. The cheap computation — 64-dim ANN — does the heavy lifting of narrowing the space.
80+
81+
## where it lands in ranking
82+
83+
Re-ranking with full embeddings isn't the end of the pipeline in e-commerce. After vector similarity, you typically blend in business signals: price, margin, inventory, recency, click-through rate.
84+
85+
MRL fits cleanly here. Vector similarity at full resolution gives you the semantic relevance signal. Everything after that is your ranking model's job.
86+
87+
```python
88+
# final score blends semantic similarity with business signals
89+
final_score = (
90+
0.5 * semantic_score(full_embedding) # MRL full-dim similarity
91+
+ 0.2 * recency_score(product)
92+
+ 0.2 * popularity_score(product)
93+
+ 0.1 * margin_score(product)
94+
)
95+
```
96+
97+
The MRL embedding is one input to the ranker — a well-calibrated one that doesn't require you to choose between "fast and weak" or "slow and strong" at the retrieval stage.
98+
99+
## the practical numbers
100+
101+
The gains compound:
102+
103+
| Stage | Standard | MRL (64-dim recall) |
104+
|---|---|---|
105+
| Index size (10M products) | ~30GB | ~2.5GB |
106+
| ANN recall latency (p95) | ~40ms | ~5ms |
107+
| Re-rank (top 500, full 768-dim) | not done | ~8ms |
108+
| Total | ~40ms | ~13ms |
109+
110+
You get better latency and a re-ranking step you weren't running before — because the recall stage got cheap enough to afford it.
111+
112+
## what to watch for
113+
114+
**Recall quality at low dimensions.** 64-dim recall works well for common queries. For rare or highly specific queries, the truncated embedding may miss relevant products. Worth measuring recall@500 at each granularity before committing to a resolution.
115+
116+
**Training data matters more.** MRL doesn't change what the model knows — it changes how that knowledge is organized across dimensions. A weak base model trained with MRL is still a weak model. The nesting property amplifies whatever signal was there to begin with.
117+
118+
**Not all model families support it.** You can fine-tune MRL loss on top of existing encoders, but some models resist it. Models trained with MRL natively — like Nomic Embed or some versions of E5 — give better nested quality out of the box than post-hoc fine-tuning.
119+
120+
## the point
121+
122+
The retrieval funnel in e-commerce has always been about spending compute where it matters. Matryoshka embeddings give you a principled way to do that at the embedding level — cheap representations for wide recall, full representations for precise ranking.
123+
124+
One model. Multiple resolutions. You choose where to spend.
125+
126+
— v

0 commit comments

Comments
 (0)