Lesson 7 of 7 · 44 min
Capstone: design a production RAG
Assemble the whole pipeline into a permission-aware, cited, evaluated assistant over private docs — the full ingest-to-answer architecture, the numbers to defend, and a rehearsal of the decisions an interviewer (or an incident) will push on.
Capstone
Design a production RAG end-to-end
- 01Shape? — chat over a closed corpus, enterprise search, or web-scale? (sets latency budget + dominant lever).
- 02Scale? — how many docs/chunks, what QPS? (sets index choice: managed HNSW vs IVFPQ + rerank).
- 03Freshness? — how stale can an answer be? (sets ingestion: nightly batch vs CDC streaming).
- 04Tenancy? — single org or multi-tenant with per-user permissions? (sets the isolation architecture + the ACL filter).
- 05Stakes? — is a wrong answer embarrassing or dangerous? (sets eval rigor, injection guards, human-in-the-loop).
The reference architecture
acl field → 3) embed + index (pick the shape and ANN algorithm for your scale) → 4) hybrid retrieve, filtered by the user’s permissions → 5) rerank to the best 5 → 6) generate grounded in cited chunks → 7) score on the golden set + run injection checks before shipping.1Production RAG -- the whole pipeline, with the lesson behind each stage23 INGEST (offline, idempotent)4 parse (layout/OCR-aware) ............................ L2 (parsing is stage zero)5 chunk (structure-aware) + contextualize ............ L2 (orphaned-chunk fix)6 attach metadata { doc_id, acl, last_verified } ..... L5 (ACL + freshness)7 embed + index (HNSW < 10M, else IVFPQ) ............. L1/L5 (index economics)8 | CDC re-embed changed rows; re-chunk behind a dual index9 QUERY (online, per request)10 (optional) query transform ......................... L4 (only if needed)11 hybrid retrieve k=50, ACL-FILTERED at search time .. L3 + L5 (filter then retrieve)12 rerank -> top 5 .................................... L4 (precision stage)13 ground + cite per claim ............................ L7 (grounding prompt)14 GATE (before any deploy)15 split eval on production-derived golden set ........ L616 injection tests on retrieved-context channel ....... L617 p99 latency + cost budget check .................... L51819 Offline path sets quality; online path sets latency; the gate stops regressions.1# The offline ingest pipeline -- idempotent, ACL-tagged, dual-index-safe.2def ingest(doc, source_acl, llm, embed, index):3 text = parse_layout_aware(doc) # L2: tables/scans survive4 chunks = structure_aware_split(text, target=512)5 contextualized = contextualize(text, chunks, llm) # L2: cache the doc -> ~90% off6 records = []7 for i, c in enumerate(contextualized):8 records.append({9 "text": c,10 "doc_id": doc.id, "chunk": i,11 "acl": source_acl, # L5: permission travels WITH the chunk12 "last_verified": now(), # L5: freshness you can time-filter on13 })14 index.upsert(embed([r["text"] for r in records]), records) # embeddings + BM2515# Re-run on UPDATE (idempotent). Re-chunk/re-embed only behind a dual index + cutover.1def answer(query, user):2 # 4) permission-aware retrieval: filter at search time, never after3 candidates = hybrid_search(query, k=50, acl_filter=user.allowed_acls) # BM25 + dense + RRF4 # 5) precision stage5 context = rerank(query, candidates, top_n=5)67 # 6) ground it: answer only from context, cite the chunk per claim8 prompt = f"""Answer ONLY from the context. Cite the chunk id for each claim.9If the context doesn't contain the answer, say you don't know.1011Context:12{format_with_ids(context)}1314Question: {query}"""15 return llm(prompt), [c["id"] for c in context] # answer + citationsKey idea
Grounding & citations: the answer-side contract
1# Grounding contract: answer only from context, cite per claim, allow abstention,2# then VERIFY the citations actually exist before returning.3GROUND = (4 "Answer ONLY from the numbered context. Cite the chunk id [n] after each claim. "5 "If the context does not contain the answer, reply exactly: I don't know."6)78def answer_grounded(query, context, llm):9 prompt = GROUND + "\n\nContext:\n" + format_with_ids(context) + "\n\nQ: " + query10 text = llm(prompt)11 valid_ids = {c["id"] for c in context}12 cited = extract_citations(text) # e.g. [3], [7]13 if not cited or not set(cited) <= valid_ids: # hallucinated / missing citation14 log_metric("ungrounded_or_bad_citation") # gate / retry / flag for review15 return text, cited16# Abstention + per-claim citation + citation validation = an auditable, low-hallucination answer.Sizing & the numbers you should be ready to defend
1Sizing a 50M-chunk multi-tenant assistant (numbers to say out loud)23 Index RAM (HNSW) 50M x 1,536 x 4B x ~1.8 ~= 450-600 GB -> consider IVFPQ4 IVFPQ alternative ~20x less ~= 25-35 GB + reranker rescue5 Query latency ANN ~10ms + rerank ~150ms + gen fits 1.5s chat SLA6 Ingest (one-time) corpus_tokens x embed $ + ctx ~$1.02/M (~90% off w/ cache)7 Re-embed (per TB) ~$12k/month -> CDC incremental + dual index89 These are order-of-magnitude. The point is to REASON in them, not memorize them.What breaks at 3am — the failure modes to pre-empt
Key idea
The decisions an interviewer will push on
- 01Chunk size — “how did you pick it?” → measured precision@5, started ~512, dropped to ~256 (L2).
- 02Hybrid vs dense — exact tokens / codes in queries? → BM25 + dense + RRF (L3).
- 03Reranker — what’s your latency budget? → retrieve 50, rerank 5, ~120 ms (L4).
- 04Index & memory — how many vectors, how much RAM? → HNSW vs IVFPQ + rerank (L5).
- 05Freshness — how do you avoid silent recall decay? → CDC re-embed, dual-index cutover, top-k-overlap telemetry (L5).
- 06Multi-tenancy — how do you stop leaks? → ACL filter pre-rerank, namespace/VPC isolation (L5).
- 07Eval — how do you know it’s good and won’t regress? → split eval + calibrated judge gate (L6).
- 08Safety — what about prompt injection via documents? → guardrail the retrieved context + whitelist tool calls (L6).
Common mistake
“Ship it when the answers look good.”
Rollout: ship it without breaking it
repoLLM Zoomcamp — end-to-end RAG projectDataTalksClub/llm-zoomcamprepoRAG_Techniques — reference implementations for every stageNirDiamant/RAG_TechniquespaperSeven Failure Points When Engineering a RAG System (a design checklist)Barnett et al. (arXiv)A production RAG isn’t the pipeline that returns a good answer once — it’s the one that’s versioned, eval-gated, canaried on the real distribution, and rollback-able, so the next change can’t silently break it.
Checkpoint
Where should you enforce per-user document permissions in a RAG system?
Checkpoint
An interviewer asks: “It works in the demo — how do you know it’s production-ready?” Best answer?
Checkpoint
An interviewer says “design a Q&A assistant over our internal docs.” What’s the strongest first move?
Checkpoint
Asked to size the index for 50M chunks of 1,536-d embeddings on HNSW, what do you say?
Checkpoint
Three months in, the assistant’s answers are subtly more outdated, but nothing was deployed and all dashboards are green. What does a strong design already have in place?
Checkpoint
A new B2B customer asks for the assistant over their private wiki, with exact part-number queries common, ~3M docs, a 1.5 s budget, strict isolation, and weekly doc updates. What’s the defensible end-to-end design?
Interview prep
- 01“Design a Q&A assistant over our docs.” → scope first (shape, scale, freshness, tenancy, stakes), then ingest → filtered-retrieve → rerank → grounded-cite → eval-gate.
- 02“How did you pick chunk size?” → measured precision@5; started ~512, dropped toward ~256; decoupled match-unit from read-unit (L2).
- 03“Hybrid or dense?” → BM25 + dense + RRF if queries carry codes/names/jargon; skip for pure paraphrase over a clean corpus (L3).
- 04“Reranker — where and what cost?” → on the fused shortlist before prompt assembly, retrieve 50 / rerank 5, ~120 ms (L4).
- 05“How many vectors, how much RAM?” → N × dim × 4B × ~1.8; HNSW under ~10M, IVFPQ + rerank beyond (L1/L5).
- 06“How do you stop tenant data leaking?” → ACL filter at retrieval time (filter then retrieve), validate retrieved context, namespace/VPC isolation (L5/L6).
- 07“How do you avoid silent recall decay?” → top-k-overlap telemetry, CDC re-embed, dual-index cutover (L5).
- 08“How do you know it’s production-ready?” → passes a split-eval gate on a production-derived golden set, clears injection tests, meets the p99/cost budget (L6).
Common mistake
The red flag that sinks candidates: designing the happy path and going silent on failure.
Could you whiteboard this whole system — and defend every decision — in a 45-minute interview?
You can now
- Scope a RAG design before drawing it — shape, scale, freshness, tenancy, stakes.
- Assemble a grounded, cited, permission-aware ingest→query→gate pipeline end-to-end.
- Size it with arithmetic (RAM, latency budget, ingest cost) and defend each decision with a metric.
- Name the silent failures unprompted — drift, leaks, shearing, the long tail, the cost cliff — and the defenses.
- Gate changes on a split eval + injection tests + a p99/cost budget, not vibes.
Up next track: Agents, Evals & LLMOps — give your RAG tools and ship it reliably.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.