Finds the brand, misses the occasion.
Gets the occasion, blurs the brand.
The brand and the occasion.
TakeawayKeyword matches tokens, vectors match meaning. In production you run both: Milvus does BM25 and dense in one query.
Each model turns its input into an array of numbers - the embedding's position in high-dimensional space. Anywhere from a few hundred to a few thousand Float32 values.
RAG
Ground LLMs in your own documents
Agent memory
Recall the right past conversation
Legal analysis
Surface relevant case law
Fraud detection
Spot the needle in a stack of needles
Song matching
Identify a tune from a whistle
Visual search
Find products that look like this photo
Autonomous driving
Detect erratic lane changes
Molecular discovery
Find molecules with similar shape
Cancer screening
Match diagnostic images to known cases
How do you measure similarity in multi-dimensional space?
Compare the query to every entity in the database. Exact, simple, O(N).
16 dresses, fine. A billion vectors? Not so much. Latency grows linearly and ~all comparisons are waste.
Every technique trades speed, accuracy & cost.
Of all the good stuff, how much did we find?
Of what we returned, how much was good?
Production notesRecall & precision are measured against exact search, not human judgement.
ThoughtWhat would happen if we filtered to size 38, under €150?
Hierarchical Navigable Small World. Multi-layer graph: top layers have long-range highways, lower layers have local connections. Start at the top, walk greedily closer, drop down a layer, repeat.
IVF clusters the vectors into nlist cells. At query time, only search within the nearest nprobe cells. A true neighbour just over the border of a cell you never open is simply gone.
Graph index, engineered for SSD. Minimises random reads, index billions of vectors on ~GBs of RAM.
Approximate nearest-neighbour algorithms all trade perfection for reduced latency and cost.
No matter what algorithm you use, embeddings are big. In RAM or on disk, size matters.
100M chunks and you are holding 1.23 TB before indexing overhead or a single query runs.
Round float32 → int8: 4x smaller embeddings, a tiny recall hit & almost no work.
Rotate the space to reduce error, then keep just the sign of each dimension - one bit.
Scalar quantisation shrinks every number, PQ shrinks the whole vector.
Every lost bit risks recall, but the curve is surprisingly forgiving.
Each algorithm can use quantisation to trade accuracy for significantly reduced latency and cost.
All those knobs. Can't the machine work it out?
Trade-offFull control and full responsibility, or one dial and trust the engine.
PCA finds the directions of greatest variance and keeps the top k. Fewer dimensions, full precision.
BenefitKeep one number instead of two and 94% of the variance - linear, fast, deterministic.
DrawbackMaximises for variance, not meaning: structure on a low-variance axis is discarded, and it must be refit when the data shifts.
The dimensions are ordered by importance, so a prefix is a complete vector.
MRL tunes the model so the dimensions are ordered by importance. OpenAI's text-embedding-3-large is 3072-D native, but you can ask for any prefix down to 256-D via the dimensions parameter.
BenefitOne model, pick the length per query - short prefix to shortlist fast, full vector to re-rank. Degrades gracefully.
DrawbackOnly works if the model was trained this way - an ordinary embedding survives a light trim, then falls off a cliff once you cut hard (the berry line).
Dimensionality reduction nudges any index toward fast and cheap.
Build time compression and dimensionality reduction both trade accuracy to buy speed and scale. Refinement wins accuracy back at query time.
nprobe of the lists using the 1-bit RaBitQ codes. Cheap, and on its own it stops at 0.778.refine_k × limit candidates with them. Same candidates, better distances: 0.986.nprobe 256 to 1024 buys 0.992, at a quarter of the throughput.Refinement spends a little cost & latency to buy accuracy back.
The catchThe harder you filter, the more of the graph you destroy.
There's no single fix - the right technique depends on how much survives the filter.
=# EXPLAIN ANALYZE SELECT * FROM dresses WHERE size = 38 AND price < 150; Index Scan using dresses_size_price_idx (cost=0.42..8.44 rows=12) (actual time=0.03..0.04 rows=12 loops=1) Index Cond: ((size = 38) AND (price < 150)) Execution Time: 0.05 ms
>>> client.search("dresses", data=[q], limit=5,
filter="size == 38 and price < 150")
The gapSQL fails loudly. Vector search fails silently, so we have to build the instrumentation back ourselves.
Score the index against exact search, on every deploy and every data change.
size == 38 and price < 150
Average over ~500 real queries, plot it daily.
Watch every layerEach layer up is slower and noisier, but closer to what matters. A dip at the bottom is cheap to catch.
Index recall: does the index match exact search? Answer-presence: is the answer in the top 10?
NQ-Open450 questions10M Wikipedia 2023-11mxbai-embed-large-v111 index configs
Evaluate the model firstThe embedding model and chunking set the ceiling. Measure LGTM@k using exact search before ANN.
The same question, three ways to feed the model. The search itself is never the expensive part.
No search, no payload. You pay for the answer, and nothing else.
Small results each time. But the whole transcript goes back every turn.
One fat payload of chunks, and a box that bills daily whether you ask or not.
A stronger model grades every answer, without knowing which method was used.
A blind model answered 25 of 40 questions on fastapi, and 0 on a repository published after training cutoff.
fastapi + agentic-hil40 questions eachSonnet 5 agentOpus 5 judgesclaude-context 0.1.15Milvus
Total agent spend over the whole run, divided by the answers the judge marked correct
fastapi + agentic-hil40 questions eachSonnet 5 agentOpus 5 judgesclaude-context 0.1.15Milvus
| Code it knows | Code it has never seen | |
|---|---|---|
| cost | every retrieval arm within ~15% | the index is ~40% cheaper |
| p95 latency | grep 32 s, indexed 46 s | grep 60 s, indexed 39 s |
No memory system vs Milvus memsearch index on a box at $86.07 a month
| Token $ per query | 98k tokens | 392k tokens | 1.2M tokens* |
|---|---|---|---|
| Replay, cache warm | $0.020 | $0.203 | $0.290 |
| Replay, cache cold | $0.393 | $1.569 | $3.839 |
| Index | $0.011 | $0.012 | $0.019 |
| Break-even, cache warm | 334 a day | 15 a day | 10 a day |
| Break-even, cache cold | 7.4 a day | 1.8 a day | 0.7 a day |
Break-even is the box's $2.83 a day divided by the saving per query. Warm uses the measured cache hit rates, 100%, 92% and 97%; cold is a miss on every query, which is what a five-minute cache sees at 15 queries a day. Mean billed cost per query; bars share one scale.
* 1.2M tokens is past the context window: replay keeps 78.5% of the history and answers 47% of questions correctly, against 100% for the index.
synthetic history98k / 392k / 1.2M tokensmemsearch, dense + BM25r8g.large, 16 GiBprices 2026-09-07
I assumed the worst-case - that you provision a VM just for the index
a box of your own
A dedicated r8g.large at $86.07 a month to hold the index. Break-even lands at 10 to 334 queries per day with a warm cache, and under 8 with a cold one.
a box you already run
Milvus Lite or Standalone on your laptop or spare space on an existing box. Same index, same prices, break-even on the first query.
no box at all
Serverless, metered per query with no floor to amortise. Every corpus combined fits into the Zilliz free-forever tier.
pip install pymilvusNotebooks, prototypes, CI, edge
docker compose upOne box in production
helm install milvusScale-out, with a platform team
pay per request + storageSmall or spiky traffic
provisioned compute, per hourSteady traffic, latency SLAs
storage + compute uptimeHuge dataset, occasional queries
Same engine in every box
Every lever in this talk spends recall, buys it back, or checks the balance.
Trade <10% recall for >100× speed, but only if the index is needed.
So far we've looked at the tech and the numbers, next we'll see what it looks like in production.
19:20 · Criteo
From Product Need to Distributed Vector Search
Mehdi & Peter on why their use case needed a distributed vector database, and the pain points on the way.
20:00 · Gorgias
RAG Design Patterns: Product Indexing at Scale
Mohamed & Othmane on two years of scalability pressure shaping a product index.