RAG at scale: index economics, freshness & multi-tenancy
What breaks only at scale — HNSW/IVF/IVFPQ memory economics, the build-vs-buy decision, the silent freshness graveyard, multi-tenant permission leaks, and caching/cost — with case studies from Perplexity, Harvey, Notion and LinkedIn.
Lesson 5 · Production
RAG at scale
The failures a demo never shows
At small scale, RAG bugs are loud — a wrong answer you can see. At scale, the dangerous failures are silent: recall decays while every dashboard stays green, a permission filter leaks one tenant’s docs into another’s answers, an index quietly costs 20× the RAM it needed. None of these show up in a notebook. This lesson is the part that separates “I built a RAG demo” from “I run RAG in production.”
Index economics: HNSW vs IVF vs IVFPQ
Past tens of millions of vectors, your ANN algorithm is a memory-budget decision first, a recall decision second. AWS’s billion-vector OpenSearch benchmark (1B × 128-d) makes the tradeoff brutally concrete:
Before the numbers, the mechanism, because interviewers ask you to explain why the memory differs by 20×. HNSW stores the full vectors plus a navigable graph (extra edges per node) — fast and high-recall, but it holds everything in RAM, so memory ≈ raw vectors × ~1.5–2. IVF clusters vectors into cells and at query time probes only the nearest few cells (the nprobe knob) — it stores full vectors too, so the memory saving is modest; the win is search speed and a tunable recall/latency dial. IVFPQ adds product quantization: each vector is chopped into sub-vectors, each sub-vector replaced by the id of its nearest centroid in a small codebook, so a 1,536-d float32 vector (~6 KB) compresses to a few dozen bytes. That’s the ~20× memory cut — and the recall cost, because you’re now searching lossy approximations, which is exactly why you pair IVFPQ with a reranker that re-scores the shortlist on full fidelity.
code
11 billion 128-d vectors, 1 replica23 Algorithm Memory p99 latency Recall Use when4 --------- -------- ----------- ------ ------------------------------5 HNSW ~1,408 GB 32-152 ms highest you can afford the RAM6 IVF ~1,126 GB < HNSW medium memory-tight, recall-tolerant7 IVFPQ ~70 GB 106-230 ms lowest massive footprint; pair w/ rerank89 HNSW vs IVFPQ = ~20x the memory. Index choice IS a procurement decision.
This is also where managed monoliths hit a wall. Harvey (legal AI) found Postgres + pgvector comfortable only to ~500k embeddings (sub-2 s p50); at tens of millions per tenant they moved to LanceDB IVF-PQ, which holds sub-2 s p50 at 15M rows with metadata filtering. The rule of thumb: if a live serving index will exceed ~10M vectors, plan for IVFPQ + reranking early — and budget for periodic recompute.
Interview angle. “We have 800M vectors and a tight RAM budget — which index?” is a near-guaranteed senior-RAG system-design question, and the strong answer is a decision, not a name: HNSW if you can afford the RAM (best recall, ~1.4 TB at 1B); IVFPQ if you can’t (~70 GB, ~20× less), recovering the lost recall with a reranker on the shortlist. Then add the second-order points — IVFPQ needs periodic recompute as data shifts, and metadata filtering interacts badly with some ANN indexes (a filter that excludes most candidates can wreck recall, so you may need pre-filtering or partitioned indexes). Quoting the ~20× memory delta and the rerank rescue is what lands it.
code
1Vector count -> the index decision (rough thresholds)23 Scale Comfortable choice Why4 -------------- ------------------------ ----------------------------------5 < 1M pgvector / managed HNSW simplest; RAM is a rounding error6 1M - 10M managed HNSW predictable cost, fast to ship7 10M - 100M HNSW (RAM-rich) or IVF watch RAM; IVF if memory-tight8 100M - 1B+ IVFPQ + reranker ~20x memory cut, recover w/ rerank9 any, privacy self-host in tenant VPC data plane is the security boundary1011 The threshold isn't just count -- it's RAM budget x privacy SLA x QPS x team size.
The “which vector database?” question is really four questions, and answering “Pinecone” or “self-host Vespa” without them is the junior tell. (1) Privacy/compliance — regulated or per-tenant-isolated data may legally have to stay in the customer’s cloud, which forces self-hosting in a VPC (Harvey). (2) Freshness SLA — seconds-fresh at high update rates pushes you toward an engine built for streaming updates (Perplexity on Vespa), away from rebuild-heavy indexes. (3) QPS and scale shape — managed services are predictable and cheap below ~10M vectors but get expensive and less tunable as you grow. (4) Team size — self-hosting an ANN engine is real operational load; a three-person team should buy, a platform team can build. The honest default: buy until one of privacy, freshness, scale, or cost forces you to build.
The freshness graveyard
Vector similarity ignores time. So a stale index fails invisibly: latency, throughput, and individual retrievals all look healthy while distributional recall drifts — one documented case fell from 0.92 to 0.74 with nothing in the dashboards to explain it. Three independent clocks drift:
01Document clock — the source changed. Rough shelf lives: API/SDK refs ~2 weeks, compliance ~6 months, architecture docs 1–2 years.
02Embedding clock — you upgraded the encoder; new vectors aren’t cosine-comparable to old ones (“representation shearing”).
03Chunking clock — you changed chunk boundaries, so a vector now represents different text than before.
Defenses the big systems use: CDC-driven incremental re-embedding (Debezium-style — re-embed only changed rows) instead of $12k/month full reindexes per TB; dual indexes (old + new in parallel, shift traffic once the new one validates on a held-out set); and shipping freshness as a first-class metric — top-k overlap and nearest-neighbour stability, plus a “last-verified” timestamp you can time-filter on. Never mix embedding-model generations in one index without a reindex plan.
The reason this is the lesson that separates demo-builders from operators: every other failure in RAG is visible and this one isn’t. A wrong answer you can see and fix. Stale recall produces plausible answers from outdated chunks — the system looks healthy on every dashboard while quietly regressing, and the documented 0.92 → 0.74 drop happened with nothing in the metrics to explain it. So the defense is telemetry that watches the distribution, not individual requests: top-k overlap (does today’s top-k for a fixed probe set match last week’s?) and nearest-neighbour stability catch drift that end-to-end accuracy misses. Interview angle. “How do you detect a stale or drifting index?” → ship freshness as a metric (top-k overlap on a probe set, a last-verified timestamp), not just latency/throughput; saying “I’d watch recall on a fixed probe set over time” is the senior answer.
python
1# Freshness as a first-class metric: watch top-k OVERLAP on a fixed probe set.2# A drop here flags silent drift that latency/throughput dashboards never show.3def topk_overlap(probe_queries, retrieve, baseline_topk, k=20):4 drift = []5 for q in probe_queries:6 today = {r["chunk_id"] for r in retrieve(q, k=k)}7 base = set(baseline_topk[q])8 overlap = len(today & base) / max(len(base), 1) # 1.0 = identical, < 1 = drift9 drift.append(overlap)10 return sum(drift) / len(drift)11# Alert when mean overlap falls below a threshold (e.g. 0.9) -> investigate a12# vendor checkpoint change, an embedding-generation mix, or stale documents.
The write path: deletes, updates, and consistency
Demos only ever insert; production systems delete and update, and the write path has its own scaling failures. Deletes must propagate or you leak — a document revoked from a user, or removed for compliance, that still lives in the index will keep surfacing in answers (and ANN indexes often tombstone rather than truly remove, so deleted vectors can still be navigated until a compaction runs — schedule it). Updates are a delete-plus-insert at the chunk level, and because a single edit can change which chunks a document produces, a naive “re-embed the doc” can orphan stale chunks under the old ids — clear the document’s old chunks before re-inserting. Consistency: there’s a window between “document changed” and “index reflects it,” so for freshness-critical answers you either accept eventual consistency (and surface a last-verified age) or read-through to the source of truth. Interview angle. “A doc was deleted but still shows up — why?” → tombstoned-not-compacted, or a delete that never propagated from the source through your CDC pipeline; name compaction and end-to-end delete propagation.
Multi-tenancy & the permission-leak trap
In B2B RAG, the scariest bug is one tenant seeing another’s data. Three real leak paths: (1) metadata-filter bypass — a missing/mis-set tenant_id filter returns forbidden docs; (2) prompt injection that pulls privileged context — a poisoned “public” doc flips the model into requesting documents the user can’t see; (3) cross-tenant churn on shared namespaces. The defenses are architectural:
02Single index + metadata filter + ID prefix — shared compute, full RBAC at query time, but leaks if filters are inconsistent.
03Per-customer VPC (Harvey) — embeddings + source never leave the tenant’s data plane; the vector DB is the security boundary.
The single most important architectural rule here is where the permission check lives: it must be a filter on the retrieval query, so a forbidden chunk is never fetched in the first place. The tempting anti-pattern — retrieve broadly, then tell the model (or a post-filter) to ignore documents the user can’t see — is unsafe because the chunk already entered the system: a prompt-injection payload, a citation request, or a summarization step can surface what was fetched. “Filter then retrieve” is safe; “retrieve then redact” is a breach waiting to happen.
python
1# SAFE: permissions are a filter on the search itself -> forbidden chunks never load.2def retrieve_safe(query, user, k=50):3 return hybrid_search(query, k=k, acl_filter=user.allowed_acls) # filter AT query time45# UNSAFE: fetch everything, then try to hide it. A clever prompt can surface it.6def retrieve_unsafe(query, user, k=50):7 hits = hybrid_search(query, k=k) # forbidden chunks ALREADY loaded8 return [h for h in hits if h["acl"] in user.allowed_acls] # too late -- breach risk910# Rule: a chunk that is never retrieved can never leak. Enforce ACLs pre-rerank.
Interview angle. Multi-tenant isolation is the highest-stakes question in enterprise RAG, and the wrong answer (“I’d add it to the system prompt” or “I’d redact the output”) is an instant fail. The senior answer names the three failure paths (filter bypass, injection that pulls privileged context, cross-tenant churn on shared namespaces) and the matching defenses (ACL filter at retrieval time; validate retrieved context, not just the user message; namespace- or VPC-level physical isolation for the strictest tenants). Frame the vector store as a security boundary, and pick the isolation level from the compliance requirement — namespace-per-tenant or per-customer VPC for regulated data, single-index-plus-metadata-filter only when you can guarantee filter consistency.
Latency, cost & caching at scale
At volume, cost is dominated by embedding, reranking, and generation; latency by ANN search, rerank, and prefill. The big levers: semantic query caching (short-circuit repeat queries — but scope per question-type and skip it where freshness matters, or you serve a stale regulated answer); KV / document caching (RAGCache reports up to 4× faster time-to-first-token and 2.1× throughput by reusing retrieved-document KV tensors); and model routing — classify the query and send the long tail of easy factoids to a small model, hard ones to a big one. One production report combined Redis semantic caching + routing for a ~5× LLM spend cut. And watch the long tail: ~5% of queries drive ~80% of retrievals, but the rare 95% are where users actually see failure — cache the hotspots, but budget the last 20% of latency for the tail.
python
1# A back-of-envelope cost model for a RAG feature -- the math to do in the interview.2qps_daily = 500_000 # queries/day3rerank_cost = 2.00 / 1000 # $/search (managed reranker)4in_tokens = 2_000 # retrieved context + question5out_tokens = 3006price_in = 2.50 / 1_000_000 # mid-tier generation model7price_out = 10.00 / 1_000_0008cache_hit_rate = 0.40 # semantic cache on repetitive traffic910gen = in_tokens * price_in + out_tokens * price_out11per_query = (rerank_cost + gen) * (1 - cache_hit_rate) # cache skips both for hits12monthly = qps_daily * 30 * per_query13print(round(monthly)) # caching + a small-model route on the easy tail move this most
The filtered-ANN gotcha (and the long tail)
A scaling failure that surprises people: metadata filtering can wreck ANN recall. HNSW and IVF navigate a graph or cluster structure built over all vectors; if you apply a restrictive filter (one tenant, one date) after the search, most of the returned neighbours get discarded and you may be left with far fewer than k — or the graph walk never reaches the eligible region at all. The fixes are architectural: pre-filtering (restrict the candidate set before the ANN walk, supported by some engines), partitioned indexes (a separate index per common filter value — e.g. per tenant, which doubles as isolation in L5), or over-fetching and filtering with a much larger k. Interview angle. “Your per-tenant recall is fine in aggregate but terrible for small tenants — why?” → post-filtering on a shared ANN index; small tenants have few eligible vectors, so the global graph walk rarely surfaces them — partition the index or pre-filter.
The case studies rhyme. Perplexity: one engine (Vespa), co-located compute, 200B+ URLs / 400 PB hot, tens of thousands of index updates/sec. Notion: 2 s → 350 ms by serving a smaller fine-tuned model, not by touching retrieval. LinkedIn: collapsed five legacy retrieval systems into one LLM dual-encoder for 1.3B users (75× throughput) by investing in signal quality, not field count. Harvey: per-tenant IVF-PQ inside the customer’s cloud for privacy. Different shapes, same discipline: RAG at scale is a systems problem, not a model problem.
At scale, RAG stops being a model problem and becomes a systems problem: memory budgets, freshness telemetry, tenant isolation, and the long tail. The teams that win treat the vector store like a database with an SLA, not like a magic box.
You must serve 800M vectors but your RAM budget is tight. Which index choice fits — and what’s the cost?
AHNSW — best recall, accept the RAM billBIVFPQ — ~20× less memory, then recover quality with a reranker on the shortlistCIt doesn’t matter — all ANN indexes use similar memory
A multi-tenant RAG returns correct answers in testing. What’s the highest-priority production risk to design against?
ASlightly higher latency from metadata filteringBPermission leaks — enforce ACL filters at retrieval time, before rerank/prompt assemblyCChoosing the wrong embedding model
Latency, throughput, and individual retrievals all look healthy, but users say answers are subtly out of date. Dashboards are green. What’s the likely failure and how would you have caught it?
AA capacity problem — add more replicasBSilent freshness drift (stale docs or a vendor checkpoint change); catch it with top-k-overlap telemetry on a fixed probe setCThe reranker is misconfigured
A 4-person startup is launching B2B RAG with ~2M vectors per customer and a strict data-residency requirement. Build or buy the vector store?
ASelf-host Vespa from day one for maximum controlBBuy a managed store, but pick one that supports per-tenant isolation / in-region (or VPC) deployment to satisfy residencyCUse one shared index with a metadata filter and ignore residency for now
You upgrade your embedding model and push the new vectors into the existing live index alongside the old ones. What happens?
ARecall improves immediately since the new model is betterBNothing — embeddings from different models are comparableCRecall shears silently because old and new vectors aren’t comparable; you must full-reindex behind a dual index and cut over
This is the lesson that decides senior RAG system-design rounds. Interviewers test whether you treat index choice as a memory/recall/cost decision, whether you know the silent failures (freshness drift, representation shearing, permission leaks) that never appear in a demo, and whether you reach for telemetry and architecture rather than a bigger model. Lead with the tradeoff and a number.
01“800M vectors, tight RAM — which index?” → IVFPQ (~20× less memory than HNSW) plus a reranker to recover the lost recall; HNSW only if RAM is plentiful.
02“Why does IVFPQ use ~20× less memory?” → product quantization replaces each vector with small codebook ids (a few dozen bytes vs ~6 KB), at a recall cost you rerank back.
03“Build or buy the vector DB?” → buy until privacy, freshness SLA, scale, or cost forces you to build; weigh team size — self-hosting an ANN engine is real ops load.
04“How do you detect a stale/drifting index?” → freshness as a metric: top-k overlap and NN-stability on a fixed probe set, plus a last-verified timestamp — not just latency/throughput.
05“You changed embedding model — what breaks?” → representation shearing; new and old vectors aren’t comparable, so it’s a full re-embed behind a dual index with atomic cutover.
06“How do you enforce per-tenant permissions?” → ACL filter at retrieval time (filter then retrieve), validate retrieved context against injection, namespace/VPC isolation for regulated tenants.
07“Where’s the catastrophic multi-tenant bug?” → cross-tenant leakage via filter bypass or injection pulling privileged context; the vector store is your security boundary.
08“How do you cut cost/latency at volume?” → semantic query caching (scope per question-type, skip where freshness matters), KV/document caching, and model routing of the easy tail.
Going deeper, the follow-ups that reward operators: “walk me through safely re-embedding 1 TB” (CDC incremental for changed rows; full re-embeds behind a dual index, ~$12k/month per TB, validate then cut over); “metadata filtering tanked your recall” (a filter that excludes most candidates interacts badly with HNSW/IVF — pre-filter or partition the index per common filter); “semantic cache served a stale regulated answer” (scope caching per question-type and disable it where freshness is legally required); and “the long tail is where users see failure” (~5% of queries drive ~80% of retrievals, but the rare 95% is where misses hurt — cache hotspots, budget latency for the tail). Always name the silent failure and the telemetry that catches it.
Could you whiteboard index choice, a freshness strategy, and tenant isolation for a 50M-vector multi-tenant RAG?
Not yetMostlyConfident
Takeaways
Index choice is a memory budget: HNSW (RAM-rich) vs IVFPQ (~20× less via product quantization, lean on reranking).
Build vs buy is four questions — privacy, freshness SLA, scale shape, team size; buy until one forces you to build.
Freshness fails silently — ship top-k-overlap telemetry on a probe set, CDC re-embeds, and dual-index cutovers.
Changing the embedding model shears recall — full re-embed behind a dual index, never mix generations.
Enforce ACLs at retrieval time (filter then retrieve); a chunk never retrieved can never leak.
Cache hotspots + route models for big cost cuts, but budget latency for the long tail.
Next: how to prove any of this works — evaluating and operating RAG without fooling yourself.