Lesson 7 of 7 · 44 min

Capstone: design a production RAG

Assemble the whole pipeline into a permission-aware, cited, evaluated assistant over private docs — the full ingest-to-answer architecture, the numbers to defend, and a rehearsal of the decisions an interviewer (or an incident) will push on.

Capstone

Design a production RAG end-to-end

Time to assemble it: a question-answering assistant over a company’s private docs that cites its sources, only shows a user what they’re allowed to see, is gated by an eval, and is sized to a scale you can defend. This is the canonical RAG system-design interview — and a real product. Everything below is a callback to a specific lesson.
Treat this as the interview itself. A 45-minute RAG system-design round almost always opens with “design a Q&A assistant over our internal docs,” and the difference between a mid and a senior signal is structure: clarify the requirements (shape, scale, freshness, tenancy), sketch the ingest→answer pipeline, then defend each stage with a metric and a tradeoff. The single best opening move is to ask the scoping questions before drawing anything — it shows you know the same word “RAG” hides three different systems.
  1. 01Shape? — chat over a closed corpus, enterprise search, or web-scale? (sets latency budget + dominant lever).
  2. 02Scale? — how many docs/chunks, what QPS? (sets index choice: managed HNSW vs IVFPQ + rerank).
  3. 03Freshness? — how stale can an answer be? (sets ingestion: nightly batch vs CDC streaming).
  4. 04Tenancy? — single org or multi-tenant with per-user permissions? (sets the isolation architecture + the ACL filter).
  5. 05Stakes? — is a wrong answer embarrassing or dangerous? (sets eval rigor, injection guards, human-in-the-loop).

The reference architecture

End to end: 1) ingest + parse → 2) chunk (recursive/structure-aware) and contextualize each chunk, attaching metadata incl. an acl field → 3) embed + index (pick the shape and ANN algorithm for your scale) → 4) hybrid retrieve, filtered by the user’s permissions → 5) rerank to the best 5 → 6) generate grounded in cited chunks → 7) score on the golden set + run injection checks before shipping.
code
1Production RAG -- the whole pipeline, with the lesson behind each stage23  INGEST (offline, idempotent)4    parse (layout/OCR-aware) ............................ L2 (parsing is stage zero)5    chunk (structure-aware) + contextualize ............ L2 (orphaned-chunk fix)6    attach metadata { doc_id, acl, last_verified } ..... L5 (ACL + freshness)7    embed + index (HNSW < 10M, else IVFPQ) ............. L1/L5 (index economics)8       |  CDC re-embed changed rows; re-chunk behind a dual index9  QUERY (online, per request)10    (optional) query transform ......................... L4 (only if needed)11    hybrid retrieve k=50, ACL-FILTERED at search time .. L3 + L5 (filter then retrieve)12    rerank -> top 5 .................................... L4 (precision stage)13    ground + cite per claim ............................ L7 (grounding prompt)14  GATE (before any deploy)15    split eval on production-derived golden set ........ L616    injection tests on retrieved-context channel ....... L617    p99 latency + cost budget check .................... L51819  Offline path sets quality; online path sets latency; the gate stops regressions.
Notice the three-path framing — it’s a strong way to narrate the design. The offline ingest path determines quality (parsing, chunking, contextualization, embedding) and is where you spend money once; the online query path determines latency and must stay lean (filtered retrieve → rerank → generate); and the deploy gate is what makes the whole thing safe to change. Candidates who only draw the query path miss half the system — and all of the parts that actually break in production. Interview angle. Open your whiteboard with these three swim-lanes (ingest / query / gate); it instantly signals you think about quality, latency, and safety as separate concerns rather than one undifferentiated “pipeline.”
python
1# The offline ingest pipeline -- idempotent, ACL-tagged, dual-index-safe.2def ingest(doc, source_acl, llm, embed, index):3    text = parse_layout_aware(doc)                 # L2: tables/scans survive4    chunks = structure_aware_split(text, target=512)5    contextualized = contextualize(text, chunks, llm)   # L2: cache the doc -> ~90% off6    records = []7    for i, c in enumerate(contextualized):8        records.append({9            "text": c,10            "doc_id": doc.id, "chunk": i,11            "acl": source_acl,                     # L5: permission travels WITH the chunk12            "last_verified": now(),                # L5: freshness you can time-filter on13        })14    index.upsert(embed([r["text"] for r in records]), records)  # embeddings + BM2515# Re-run on UPDATE (idempotent). Re-chunk/re-embed only behind a dual index + cutover.
python
1def answer(query, user):2    # 4) permission-aware retrieval: filter at search time, never after3    candidates = hybrid_search(query, k=50, acl_filter=user.allowed_acls)  # BM25 + dense + RRF4    # 5) precision stage5    context = rerank(query, candidates, top_n=5)67    # 6) ground it: answer only from context, cite the chunk per claim8    prompt = f"""Answer ONLY from the context. Cite the chunk id for each claim.9If the context doesn't contain the answer, say you don't know.1011Context:12{format_with_ids(context)}1314Question: {query}"""15    return llm(prompt), [c["id"] for c in context]   # answer + citations

Grounding & citations: the answer-side contract

Retrieval gets the right chunks into the prompt; grounding is making the model actually answer from them and cite them — and it’s where failure point #4 (“not extracted”) lives. Three rules. (1) Instruct abstention: tell the model to answer only from the context and to say “I don’t know” when the answer isn’t there — a RAG that can’t say “I don’t know” will confidently hallucinate on every out-of-corpus question. (2) Cite per claim: pass chunks with stable ids and require a citation on each claim, so an answer is auditable and a user (or your faithfulness eval) can verify it against the source. (3) Verify the citations are real: models will occasionally cite an id that doesn’t support the claim, so validate that cited chunk ids exist and (in high-stakes settings) that the claim is entailed by the cited chunk. Interview angle. “How do you stop hallucination in RAG?” → ground-and-cite plus enforced abstention plus a faithfulness check — not “use a better model.”
python
1# Grounding contract: answer only from context, cite per claim, allow abstention,2# then VERIFY the citations actually exist before returning.3GROUND = (4    "Answer ONLY from the numbered context. Cite the chunk id [n] after each claim. "5    "If the context does not contain the answer, reply exactly: I don't know."6)78def answer_grounded(query, context, llm):9    prompt = GROUND + "\n\nContext:\n" + format_with_ids(context) + "\n\nQ: " + query10    text = llm(prompt)11    valid_ids = {c["id"] for c in context}12    cited = extract_citations(text)                 # e.g. [3], [7]13    if not cited or not set(cited) <= valid_ids:    # hallucinated / missing citation14        log_metric("ungrounded_or_bad_citation")    # gate / retry / flag for review15    return text, cited16# Abstention + per-claim citation + citation validation = an auditable, low-hallucination answer.

Sizing & the numbers you should be ready to defend

Interviewers reward candidates who can back their architecture with arithmetic. Be ready to do these on the spot: index RAM = N_chunks × dim × 4 bytes × ~1.5–2 for HNSW (so 50M × 1,536-d ≈ 300 GB raw, ~450–600 GB indexed — which justifies IVFPQ’s ~20× cut at larger scale); retrieval latency budget ≈ ANN (a few–tens of ms) + rerank (~100–300 ms) + generation prefill+decode, set against the shape’s SLA (chat 300 ms–2 s, RAG-augmented chat 1.5 s+); and ingest cost ≈ corpus_tokens × embedding price + contextualization (~$1.02/M doc tokens, ~90% off with prompt caching) — a one-time cost you re-pay only on re-chunk or re-embed. Quoting these turns “I’d add a reranker” into “I’d add a reranker; it costs ~120 ms and fits the 1.5 s budget.”
code
1Sizing a 50M-chunk multi-tenant assistant (numbers to say out loud)23  Index RAM (HNSW)   50M x 1,536 x 4B x ~1.8   ~= 450-600 GB  -> consider IVFPQ4  IVFPQ alternative  ~20x less                 ~= 25-35 GB    + reranker rescue5  Query latency      ANN ~10ms + rerank ~150ms + gen          fits 1.5s chat SLA6  Ingest (one-time)  corpus_tokens x embed $   + ctx ~$1.02/M (~90% off w/ cache)7  Re-embed (per TB)  ~$12k/month               -> CDC incremental + dual index89  These are order-of-magnitude. The point is to REASON in them, not memorize them.

What breaks at 3am — the failure modes to pre-empt

A senior design names how it fails before the interviewer asks. The five that recur: (1) silent freshness drift — stale or vendor-changed embeddings sag recall with green dashboards (defense: top-k-overlap telemetry, L5); (2) permission leak — a missing ACL filter or an injection that pulls privileged context (defense: filter-then-retrieve + validate retrieved content, L5/L6); (3) representation shearing — an in-place re-embed mixes incomparable vector generations (defense: dual index + atomic cutover, L2/L5); (4) the long tail — ~5% of queries drive ~80% of retrievals, but failures hide in the rare 95% (defense: canary on real traffic, error-analyze the tail, L6); and (5) cost cliff — a reasoning model or an uncapped reranker quietly multiplies the bill (defense: route the easy tail to a small model, cap candidates, cache hotspots, L5).

The decisions an interviewer will push on

  1. 01Chunk size — “how did you pick it?” → measured precision@5, started ~512, dropped to ~256 (L2).
  2. 02Hybrid vs dense — exact tokens / codes in queries? → BM25 + dense + RRF (L3).
  3. 03Reranker — what’s your latency budget? → retrieve 50, rerank 5, ~120 ms (L4).
  4. 04Index & memory — how many vectors, how much RAM? → HNSW vs IVFPQ + rerank (L5).
  5. 05Freshness — how do you avoid silent recall decay? → CDC re-embed, dual-index cutover, top-k-overlap telemetry (L5).
  6. 06Multi-tenancy — how do you stop leaks? → ACL filter pre-rerank, namespace/VPC isolation (L5).
  7. 07Eval — how do you know it’s good and won’t regress? → split eval + calibrated judge gate (L6).
  8. 08Safety — what about prompt injection via documents? → guardrail the retrieved context + whitelist tool calls (L6).

Rollout: ship it without breaking it

The last mile is operating it. Four practices the interviewer will respect: (1) version everything — pin the model, the prompt, the chunker, and the embedding model, and log every answer with those version ids so a regression is attributable; (2) gate in CI — run the split eval on the frozen golden set on every change, blocking a deploy that regresses past a margin (L6); (3) canary on real traffic — route a slice of production through the new version and compare per-span hit@k on the real distribution before full rollout, because offline sets miss the long tail (L6); and (4) have a rollback — a config-level switch back to the last-good version, and for index changes the dual-index cutover that lets you flip back atomically (L2/L5). The throughline of the whole track: every layer is testable and reversible, so a silent regression is caught in CI or canary, never by a user.
A production RAG isn’t the pipeline that returns a good answer once — it’s the one that’s versioned, eval-gated, canaried on the real distribution, and rollback-able, so the next change can’t silently break it.
repoLLM Zoomcamp — end-to-end RAG projectDataTalksClub/llm-zoomcamprepoRAG_Techniques — reference implementations for every stageNirDiamant/RAG_TechniquespaperSeven Failure Points When Engineering a RAG System (a design checklist)Barnett et al. (arXiv)

Checkpoint

Where should you enforce per-user document permissions in a RAG system?

AIn the prompt — tell the model to ignore docs the user can’t seeBAt retrieval time, as a metadata/ACL filter on the search (before rerank and prompt assembly)CAfter generation, by redacting the answer
Sign up free to answer and see why

Checkpoint

An interviewer asks: “It works in the demo — how do you know it’s production-ready?” Best answer?

AThe answers are fluent and the stakeholders are happyBIt passes a split-eval gate on a production-derived golden set, clears injection tests, and meets the p99 latency/cost budgetCIt uses the latest models and a managed vector DB
Sign up free to answer and see why

Checkpoint

An interviewer says “design a Q&A assistant over our internal docs.” What’s the strongest first move?

AStart drawing embed → retrieve → generate immediately to show you know the pipelineBAsk scoping questions first — shape, scale, freshness SLA, tenancy/permissions, stakes — then design to those constraintsCRecommend the latest models and a managed vector DB up front
Sign up free to answer and see why

Checkpoint

Asked to size the index for 50M chunks of 1,536-d embeddings on HNSW, what do you say?

A“A few gigabytes — embeddings are small”B“It depends on the model”C“~50M × 1,536 × 4 bytes ≈ 300 GB raw, ~450–600 GB with the HNSW graph — at this scale I’d weigh IVFPQ + a reranker”
Sign up free to answer and see why

Checkpoint

Three months in, the assistant’s answers are subtly more outdated, but nothing was deployed and all dashboards are green. What does a strong design already have in place?

AAuto-scaling to handle the loadBFreshness telemetry (top-k-overlap on a probe set) + CDC re-embedding + a last-verified timestamp — so silent drift is caught and correctedCA bigger embedding model swapped in last week
Sign up free to answer and see why

Checkpoint

A new B2B customer asks for the assistant over their private wiki, with exact part-number queries common, ~3M docs, a 1.5 s budget, strict isolation, and weekly doc updates. What’s the defensible end-to-end design?

ASingle shared index, dense-only, no reranker, nightly full reindex — keep it simpleBPer-tenant (or VPC) index; structure-aware + contextual chunks with ACL metadata; hybrid BM25+dense+RRF filtered at retrieval; rerank to 5; grounded cited answers; CDC weekly updates; split-eval gateCGraphRAG over the whole wiki with a 1M-token context fallback
Sign up free to answer and see why

Interview prep

The capstone is the interview. A RAG system-design round tests whether you scope before you draw, whether you can defend each stage with a metric and a tradeoff, whether you size with arithmetic, and whether you name the silent failures unprompted. The structure that wins: clarify → sketch the ingest/query/gate paths → defend each decision → pre-empt the failure modes.
  1. 01“Design a Q&A assistant over our docs.” → scope first (shape, scale, freshness, tenancy, stakes), then ingest → filtered-retrieve → rerank → grounded-cite → eval-gate.
  2. 02“How did you pick chunk size?” → measured precision@5; started ~512, dropped toward ~256; decoupled match-unit from read-unit (L2).
  3. 03“Hybrid or dense?” → BM25 + dense + RRF if queries carry codes/names/jargon; skip for pure paraphrase over a clean corpus (L3).
  4. 04“Reranker — where and what cost?” → on the fused shortlist before prompt assembly, retrieve 50 / rerank 5, ~120 ms (L4).
  5. 05“How many vectors, how much RAM?” → N × dim × 4B × ~1.8; HNSW under ~10M, IVFPQ + rerank beyond (L1/L5).
  6. 06“How do you stop tenant data leaking?” → ACL filter at retrieval time (filter then retrieve), validate retrieved context, namespace/VPC isolation (L5/L6).
  7. 07“How do you avoid silent recall decay?” → top-k-overlap telemetry, CDC re-embed, dual-index cutover (L5).
  8. 08“How do you know it’s production-ready?” → passes a split-eval gate on a production-derived golden set, clears injection tests, meets the p99/cost budget (L6).
Going deeper, the curveballs that separate offers: “the corpus is 500 pages per query — long context or RAG?” (RAG for cost/latency/attribution; long context only for holistic single-doc tasks — L4); “users mostly ask multi-hop questions” (decomposition / agentic retrieval for that slice, not globally — L4); “walk me through changing the chunker in prod” (dual index, validate on held-out, atomic cutover — never in place — L2/L5); and “the bot gave confidently wrong advice” (narrow the candidate set before retrieval, the Uber Genie −60% move, plus faithfulness eval — L3/L6). Tie every answer back to a metric, a number, and the failure it prevents.

Could you whiteboard this whole system — and defend every decision — in a 45-minute interview?

Not yetMostlyYes

You can now

  • Scope a RAG design before drawing it — shape, scale, freshness, tenancy, stakes.
  • Assemble a grounded, cited, permission-aware ingest→query→gate pipeline end-to-end.
  • Size it with arithmetic (RAM, latency budget, ingest cost) and defend each decision with a metric.
  • Name the silent failures unprompted — drift, leaks, shearing, the long tail, the cost cliff — and the defenses.
  • Gate changes on a split eval + injection tests + a p99/cost budget, not vibes.

Up next track: Agents, Evals & LLMOps — give your RAG tools and ship it reliably.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.