|
| 1 | +--- |
| 2 | +date: 2026-06-18 |
| 3 | +tag: search |
| 4 | +title: "Matryoshka embeddings and the e-commerce retrieval funnel" |
| 5 | +read: 9 min |
| 6 | +deck: "One model, many resolutions. How nested embeddings let you run cheap recall and expensive re-ranking without paying for both upfront." |
| 7 | +hidden: true |
| 8 | +--- |
| 9 | + |
| 10 | +Most teams think of embeddings as atomic. You generate a 768-dimensional vector. You store it. You compare it. You get a result. The whole vector travels through every stage of the pipeline at full cost. |
| 11 | + |
| 12 | +Matryoshka Representation Learning (MRL) breaks that assumption. It gives you one model that produces embeddings you can truncate at multiple resolutions — and every truncation is still meaningful. |
| 13 | + |
| 14 | +The name comes from Russian nesting dolls. The first 64 dimensions of a Matryoshka embedding are a useful representation. The first 128 are more useful. The full 768 are most useful. You choose the resolution based on what the stage of retrieval can afford. |
| 15 | + |
| 16 | +## why this matters at e-commerce scale |
| 17 | + |
| 18 | +A typical e-commerce search pipeline has a retrieval problem that looks like this: |
| 19 | + |
| 20 | +- 10 million product embeddings in the index |
| 21 | +- A user query arrives every 50ms at peak |
| 22 | +- You need to return ranked results in under 100ms end-to-end |
| 23 | + |
| 24 | +Full 768-dim ANN search over 10M vectors is expensive. You can make it work, but the index is large, the memory footprint is high, and every query touches a lot of data. |
| 25 | + |
| 26 | +The standard solution is a two-stage funnel: fast approximate retrieval to get a candidate set, then more expensive re-ranking to get the final order. The problem is that both stages traditionally use the same full-dimensional embedding. You've paid the storage cost twice and the compute cost at every stage. |
| 27 | + |
| 28 | +> MRL lets you match the embedding resolution to what each stage of the funnel actually needs. |
| 29 | +
|
| 30 | +## how the nesting works |
| 31 | + |
| 32 | +Standard embedding training optimizes one loss: how well the full vector represents the input. MRL trains a joint loss across multiple prefix lengths — 64, 128, 256, 512, 768 — simultaneously. |
| 33 | + |
| 34 | +The result is an embedding where every prefix is independently useful. The model has learned to pack the most important information into the first dimensions and progressively add detail as you extend the vector. |
| 35 | + |
| 36 | +```python |
| 37 | +# training sketch — loss computed at each granularity |
| 38 | +dims = [64, 128, 256, 512, 768] |
| 39 | +total_loss = 0 |
| 40 | + |
| 41 | +for d in dims: |
| 42 | + truncated = embedding[:d] # first d dimensions |
| 43 | + loss = contrastive_loss(truncated) # must be meaningful at this size |
| 44 | + total_loss += loss |
| 45 | + |
| 46 | +total_loss.backward() |
| 47 | +``` |
| 48 | + |
| 49 | +The first 64 dimensions aren't random — they've been explicitly trained to be a good representation at that resolution. Truncation is not approximation. It's a deliberate, lower-resolution view of the same information. |
| 50 | + |
| 51 | +## the two-stage retrieval pattern |
| 52 | + |
| 53 | +In e-commerce search, this maps directly onto the retrieval funnel: |
| 54 | + |
| 55 | +**Stage 1 — recall with small embeddings** |
| 56 | + |
| 57 | +Build your ANN index using 64 or 128-dimensional vectors. The index is 6-12× smaller than a full 768-dim index. Retrieval is faster. Memory footprint shrinks. You fetch a large candidate set — say, top 500. |
| 58 | + |
| 59 | +```python |
| 60 | +# index built on truncated vectors |
| 61 | +index = build_hnsw_index(embeddings[:, :64]) # first 64 dims only |
| 62 | + |
| 63 | +# fast recall — cheap, pulls a large candidate set |
| 64 | +candidates = index.search(query_embedding[:64], k=500) |
| 65 | +``` |
| 66 | + |
| 67 | +**Stage 2 — re-rank with full embeddings** |
| 68 | + |
| 69 | +Take those 500 candidates, load their full 768-dim embeddings, and compute exact similarity scores. Re-rank. Return top 10. |
| 70 | + |
| 71 | +```python |
| 72 | +# re-rank candidates using full embeddings |
| 73 | +candidate_embeddings = load_full_embeddings(candidates) # 500 × 768 |
| 74 | +scores = cosine_similarity(query_embedding, candidate_embeddings) |
| 75 | +ranked = sorted(zip(candidates, scores), key=lambda x: -x[1]) |
| 76 | +results = ranked[:10] |
| 77 | +``` |
| 78 | + |
| 79 | +The expensive computation — full dot products — runs on 500 vectors, not 10 million. The cheap computation — 64-dim ANN — does the heavy lifting of narrowing the space. |
| 80 | + |
| 81 | +## where it lands in ranking |
| 82 | + |
| 83 | +Re-ranking with full embeddings isn't the end of the pipeline in e-commerce. After vector similarity, you typically blend in business signals: price, margin, inventory, recency, click-through rate. |
| 84 | + |
| 85 | +MRL fits cleanly here. Vector similarity at full resolution gives you the semantic relevance signal. Everything after that is your ranking model's job. |
| 86 | + |
| 87 | +```python |
| 88 | +# final score blends semantic similarity with business signals |
| 89 | +final_score = ( |
| 90 | + 0.5 * semantic_score(full_embedding) # MRL full-dim similarity |
| 91 | + + 0.2 * recency_score(product) |
| 92 | + + 0.2 * popularity_score(product) |
| 93 | + + 0.1 * margin_score(product) |
| 94 | +) |
| 95 | +``` |
| 96 | + |
| 97 | +The MRL embedding is one input to the ranker — a well-calibrated one that doesn't require you to choose between "fast and weak" or "slow and strong" at the retrieval stage. |
| 98 | + |
| 99 | +## the practical numbers |
| 100 | + |
| 101 | +The gains compound: |
| 102 | + |
| 103 | +| Stage | Standard | MRL (64-dim recall) | |
| 104 | +|---|---|---| |
| 105 | +| Index size (10M products) | ~30GB | ~2.5GB | |
| 106 | +| ANN recall latency (p95) | ~40ms | ~5ms | |
| 107 | +| Re-rank (top 500, full 768-dim) | not done | ~8ms | |
| 108 | +| Total | ~40ms | ~13ms | |
| 109 | + |
| 110 | +You get better latency and a re-ranking step you weren't running before — because the recall stage got cheap enough to afford it. |
| 111 | + |
| 112 | +## what to watch for |
| 113 | + |
| 114 | +**Recall quality at low dimensions.** 64-dim recall works well for common queries. For rare or highly specific queries, the truncated embedding may miss relevant products. Worth measuring recall@500 at each granularity before committing to a resolution. |
| 115 | + |
| 116 | +**Training data matters more.** MRL doesn't change what the model knows — it changes how that knowledge is organized across dimensions. A weak base model trained with MRL is still a weak model. The nesting property amplifies whatever signal was there to begin with. |
| 117 | + |
| 118 | +**Not all model families support it.** You can fine-tune MRL loss on top of existing encoders, but some models resist it. Models trained with MRL natively — like Nomic Embed or some versions of E5 — give better nested quality out of the box than post-hoc fine-tuning. |
| 119 | + |
| 120 | +## the point |
| 121 | + |
| 122 | +The retrieval funnel in e-commerce has always been about spending compute where it matters. Matryoshka embeddings give you a principled way to do that at the embedding level — cheap representations for wide recall, full representations for precise ranking. |
| 123 | + |
| 124 | +One model. Multiple resolutions. You choose where to spend. |
| 125 | + |
| 126 | +— v |
0 commit comments