Lesson 6 of 6 · 49 min

Worked design: enterprise LLM/RAG assistant

The fastest-growing prompt, worked end to end. Retrieval-conditioned generation over millions of internal docs: hybrid search, rerankers, chunking, grounding and citations, the eval problem without ground-truth labels, guardrails, multi-tenancy and permissioning, the scoring rubric, and the curveballs.

The prompt with the weakest industrial conventions

The LLM/RAG assistant is the fastest-growing prompt area — and the research flags it as the one with the weakest industrial conventions, where a fourth axis appears that classic ML doesn’t have: evaluation without ground-truth labels. A real ML-engineer write-up names the question candidates dread most — “how do you calculate the accuracy of your RAG / agentic system?” This lesson works the enterprise assistant end to end: the hybrid-retrieval pipeline, rerankers, chunking, grounding, the eval problem, guardrails, multi-tenancy, and the curveballs that separate offers.
Building Production-Ready RAG ApplicationsJerry Liu (LlamaIndex) — AI Engineer

Framing: retrieval-conditioned generation, not open-domain

A representative prompt: “design an enterprise assistant grounding answers in 50M internal documents (HR, engineering, sales, support); p99 query latency < 2s, citation accuracy > 90%, hallucination < 5%.” Frame it as retrieval-conditioned generation, not open-domain generation. The target function is expected answer utility − hallucination penalty + cost(tokens) + latency penalty. The modern production stack, per the Applied-AI practitioner guide: ingest → chunk → embed → vector store → hybrid search (BM25 + dense) → rerank → LLM generate → evaluate. The senior framing leads evaluation-first, because in RAG you can lift one stage and tank the whole answer.
code
1ENTERPRISE RAG (50M docs, p99 <2s, citation >90%, hallucination <5%)23  INGEST   docs (S3/Confluence/SharePoint) -> OCR/clean -> CHUNKER4              structure-aware, 10-20% overlap, 128-1024 tokens5                 |6  EMBED    BGE-large / text-embedding-3 / Cohere embed7                 |8  INDEX    HYBRID: BM25 (OpenSearch) + dense ANN (Faiss/Pinecone)9                 |10  QUERY    rewrite (HyDE/multi-query, +200-500ms, +15-30% recall)11                 -> hybrid retrieve 50-100 candidates (RRF fusion)12                 -> CROSS-ENCODER RERANK ($0.001-0.002/q) -> 5-1013                 -> prompt builder WITH CITATIONS (per-tenant ACL filter)14                 -> LLM generate -> answer + citations + audit log1516  Offline: Recall@10, RAGAS faithfulness/context-precision. Online: citation17  accuracy (human), hallucination rate, thumbs-up. Counter: hallucination rate.

Hybrid search and the reranker: the two highest-ROI moves

The single biggest quality improvement in production RAG is hybrid search — BM25 + dense retrieval — which beats either alone by 15–30% on retrieval accuracy. The mechanism: BM25 nails exact-token matches (product names, error codes, employee IDs, SKUs, regulatory references) while dense retrieval nails semantic paraphrase (synonyms, multilingual). Fuse them with Reciprocal Rank Fusion (RRF), weights tuned on a labeled eval set. Interview angle. The naive RAG answer is “embed, cosine top-K, done” — it works until users ask about exact error codes and SKUs, after which BM25 dominates. Saying “hybrid by default, not pure vector” is the senior signal.
The cross-encoder reranker is the highest-ROI single component. It reads (query, document) jointly and outputs a calibrated relevance score, narrowing 50–100 candidates down to 5–10 before the LLM sees them. Cohere Rerank 3.5 costs $0.001–$0.002 per query; BGE-Reranker is open-source at ~100ms. The cost math is unambiguous: a $0.002 rerank that lets you drop the LLM input context is dominated by the token savings (LLM input runs ~$0.02+). The reranker also raises faithfulness by removing near-miss distractors. Say it plainly: a reranker is the cheapest way to buy accuracy in an LLM system.
A reranker is the single highest-ROI component in an LLM system, and hybrid search is the single biggest quality jump — the two cheapest ways to buy accuracy before you ever touch the model. The senior RAG answer narrows the candidate set; the junior one stuffs the window.

Chunking: the latent hyper-parameter most candidates skip

Chunking is a hyper-parameter that quietly caps retrieval quality, and most interviewees underspend on it. The published best practice: structure-aware chunking, 10–20% token overlap, sizes from 128 tokens (precision) to 1024 (broader context). The high-leverage enterprise trick is parent-document retrieval: embed small 400-token chunks for precise matching but generate against the 2000-token parent section for context. Interview angle. The senior move is to say you’d tune chunk size with offline retrieval recall on a labeled eval set before touching the LLM — chunking is a retrieval decision, validated with retrieval metrics, not a vibe.
Embedding-model choice is a real tradeoff to name: OpenAI text-embedding-3-large (3072-dim), BGE-large-en-v1.5 (1024-dim), Cohere embed-v3 (1024-dim). Open-source is cheaper at scale but trails managed by ~1–3 NDCG points. The vector DB is the other axis: Pinecone/Weaviate managed vs Milvus/Qdrant self-hosted — self-hosting roughly halves infra cost but requires SRE, observability, and sharding. For 50M docs, a managed vector DB runs $5K–$25K/month, which is a real number to put in the cost story.
code
1ENTERPRISE RAG -- numbers and metric targets to quote23  COMPONENT / METRIC          NUMBER                   SOURCE4  -------------------------   ----------------------   --------------5  managed vector DB (50M)     $5K-$25K / month         Applied-AI6  cross-encoder rerank        $0.001-$0.002 / query    Applied-AI7  rerank lift (fin. data)     +23.4% vs hybrid search  Cohere (Rerank 3.5)8  multi-query rewrite         +200-500ms latency       Applied-AI9  hybrid (BM25+dense) lift    +15-30% retrieval acc    Pinecone10  p99 query latency (MVP)     < 2 seconds              Applied-AI1112  TARGETS: retrieval recall@10 >0.8 | RAGAS faithfulness >0.85 |13           citation accuracy >90% | hallucination <5% | thumbs-up >70%14  Cheapest call: stable prefix (system + context) FIRST, user turn LAST.

The eval problem: measuring accuracy without ground-truth labels

This is the axis classic ML doesn’t have, and the one candidates most often improvise badly. You cannot label 50M-document answers by hand, so you need reference-free evaluation. The de-facto framework is RAGAS, with the metrics worth naming: faithfulness (is the answer grounded in the retrieved context?), answer relevancy, and context precision/recall (did retrieval surface the right chunks?). The strong answer proposes a layered eval: RAGAS faithfulness > 0.85 offline, citation accuracy > 90% via human spot-check, and an LLM-as-judge with a rubric prompt plus a sampling-and-audit loop online. Quote the targets — hallucination < 5%, thumbs-up > 70% — as the bar, not a vibe.
Interview angle. “How do you measure hallucination without ground truth?” is the curveball. Decompose it: a faithfulness check (does every claim trace to a retrieved chunk?), a citation-grounding requirement (the answer must cite source-doc IDs with offsets, the cheapest enterprise-compliance win), and a refusal classifier for out-of-distribution questions. The senior framing: you measure faithfulness and retrieval recall continuously with LLM-as-judge, sample for human audit, and tie the hallucination rate back to a business metric (CSAT, cost-per-ticket-deflection) — “hallucination is at 5%, can we ship?” is answered in business terms, not model terms.

Guardrails, multi-tenancy, and the failure modes that bite

Enterprise grounding has a hard requirement classic RAG demos skip: per-tenant permissioning. The assistant must answer HR questions only from docs the asking user may see, which means ACL filtering at retrieval time (filter the candidate set by the user’s document permissions before the LLM sees anything) — index isolation or per-tenant metadata filters. The silent failure mode to pre-empt is a permission leak: a chunk from a restricted doc surfacing in someone else’s answer. Guardrails wrap the rest: citation enforcement, a refusal path for OOD questions, PII redaction, and an audit log.
Two more failure modes earn senior points when named unprompted. Freshness / recall drift: when documents change, embeddings and reranker MAP decay silently — you need incremental re-indexing keyed on chunk modification dates and recall telemetry, or the assistant confidently cites a stale policy. The cost cliff: LLM token cost plus rerank dominate, so caching matters — an embedding-query cache at the retrieval plane and an LLM-output cache at generation, with a >30% hit rate typical for support workloads, plus prefix caching on the stable system-prompt-plus-context prefix (~90% cost / ~85% latency from Lesson 3). Multi-query rewriting adds +200–500ms, so enable it only where recall demands it.

Interview prep

The RAG prompt is graded on the pipeline shape, the eval-without-labels story, and the enterprise requirements (permissions, citations, cost). Lead evaluation-first. Answer each in 60–90 seconds.
  1. 01“Design an enterprise RAG assistant.” → ingest→chunk→embed→hybrid index→rewrite→retrieve→rerank→generate-with-citations→eval; ACL-filter at retrieval; lead evaluation-first.
  2. 02“Pure vector or hybrid search?” → hybrid (BM25 + dense, RRF) — 15–30% better; BM25 for exact tokens (error codes, SKUs), dense for paraphrase.
  3. 03“What’s the highest-ROI component?” → the cross-encoder reranker — $0.001–0.002/query narrows 50–100 to 5–10, raising faithfulness; token savings dominate its cost.
  4. 04“How do you chunk?” → structure-aware, 10–20% overlap, 128–1024 tokens; parent-document retrieval (embed 400, generate 2000); tune on offline retrieval recall first.
  5. 05“How do you measure accuracy without ground truth?” → RAGAS faithfulness + context precision/recall offline; LLM-as-judge + human spot-check + citation accuracy online.
  6. 06“How do you measure / bound hallucination?” → faithfulness (claims trace to chunks) + mandatory citations + refusal classifier for OOD; tie the rate to CSAT/deflection.
  7. 07“How do you ground to private docs without leaking?” → ACL filter the candidate set at retrieval by the user’s permissions; index isolation / per-tenant metadata; audit log.
  8. 08“Big-context model — skip retrieval?” → no; a medium model + reranker beats a 200k-token call at ~20% cost, and context rot degrades quality as the window fills.
Going deeper, the curveballs that decide the round: “hallucination is at 5%, can we ship?” (answer in business terms — tie it to CSAT and cost-per-deflection, propose a refusal threshold and human-escalation for low-confidence answers); “a user uploads a PDF that gets indexed at runtime” (ingestion pipeline with OCR, PII scrubbing, per-tenant isolation, and a security review — never index untrusted content into a shared space); “multi-turn with memory” (summarization step with a decay policy, and the cost of carrying history in every prompt); and “walk me through changing the chunker in prod” (dual index, validate on a held-out eval set, atomic cutover — never re-chunk in place). Tie every answer to a metric, a number, and the failure it prevents.
articleEnterprise RAG Architecture: A Practitioner’s Guide (hybrid, rerank, chunking, cost)Applied AIdocsRAGAS — reference-free RAG evaluation metrics (faithfulness, context precision/recall)RAGASarticleIntroducing Rerank 3.5 (+23.4% vs hybrid search, the cross-encoder rerank primary)CoherearticleChunking Strategies for LLM Applications (the full taxonomy)Pinecone

Checkpoint

Your enterprise RAG assistant retrieves well on conceptual questions but fails when users ask about exact error codes and SKU IDs. What’s the fix?

ASwitch to a larger embedding modelBAdd BM25 lexical search alongside dense retrieval (hybrid search, fused via RRF) — BM25 nails exact-token matches like codes and IDs, dense handles paraphrase; hybrid beats either alone by 15–30%CIncrease the number of retrieved chunks to 200DLower the chunk size to 64 tokens
Sign up free to answer and see why

Checkpoint

The interviewer asks: “how do you measure the accuracy of this RAG system when you have no labeled answers for 50M documents?” Strongest answer?

AReport the average cosine similarity of retrieved chunksBUse a reference-free eval stack — RAGAS faithfulness and context precision/recall offline, plus LLM-as-judge with a rubric and human spot-checks online — and track citation accuracy and hallucination rate against targetsCWait until you can hand-label all 50M documentsDTrust the LLM since modern models rarely hallucinate
Sign up free to answer and see why

Checkpoint

A teammate proposes skipping retrieval entirely and stuffing all relevant docs into a 200k-token context window per query. Why push back?

ALarge-context models don’t exist yetBIt’s more accurate but slowerCYou pay for every token (attention is super-linear), quality degrades as the window fills (context rot / Lost-in-the-Middle), and a medium model + reranker often beats it at ~20% of the costDCitations become impossible with retrieval
Sign up free to answer and see why

Checkpoint

Your assistant answers HR and finance questions for all employees from one shared index. Security flags that a junior employee got an answer citing a restricted compensation doc. Root cause and fix?

AThe LLM leaked training data; switch modelsBNo ACL filtering at retrieval — the candidate set wasn’t filtered by the user’s document permissions before the LLM saw it; fix with per-tenant/permission metadata filters (or index isolation) applied at retrieval time, plus an audit logCThe reranker scored the restricted doc too high; lower its weightDThe chunk size was too large; reduce it
Sign up free to answer and see why

Checkpoint

Users report the assistant confidently cites an outdated policy that was revised last month. What production gap does this reveal?

AThe LLM needs fine-tuning on the new policyBA freshness / recall-drift gap — when documents change, the index and reranker silently decay; fix with incremental re-indexing keyed on chunk modification dates plus recall telemetryCThe hybrid search weights are wrong; retune RRFDThe context window is too small to hold the new policy
Sign up free to answer and see why

Could you design and defend an enterprise RAG assistant — pipeline, eval-without-labels, guardrails, multi-tenancy — in a 45-minute round?

Not yetGetting thereConfident

Takeaways

  • Frame RAG as retrieval-conditioned generation and lead evaluation-first; the pipeline is ingest→chunk→embed→hybrid-retrieve→rerank→generate-with-citations→eval.
  • Hybrid search (BM25 + dense, RRF) beats pure vector by 15–30%; the cross-encoder reranker is the highest-ROI component ($0.001–0.002/q, 50–100→5–10).
  • Chunking is a tunable hyper-parameter — structure-aware, 10–20% overlap, 128–1024 tokens, parent-document retrieval — tuned on offline retrieval recall first.
  • Measure accuracy without labels via RAGAS (faithfulness, context precision/recall) + LLM-as-judge + human spot-checks; bound hallucination with citations + a refusal path, tied to CSAT.
  • ACL-filter the candidate set at retrieval to prevent permission leaks; watch freshness drift (re-index on modification dates) and the cost cliff (caching, prefix caching).
  • Never stuff a giant context to skip retrieval — a medium model + reranker beats it at ~20% cost and avoids context rot; name the failure modes unprompted.

You can now frame, work, and defend the full menu of ML/LLM system-design prompts end to end. Re-run the weak spots with a timer and a graded mock.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.