Prove it works without fooling yourself: the IR retrieval metrics, split retrieval vs generation eval, where RAGAS metrics lie, LLM-as-judge biases & calibration, golden sets, observability, drift, and prompt-injection in the retrieval path.
Every RAG demo looks great on the three questions you tried. Evaluation is how you know it works on the next thousand — and how you catch a “small” prompt tweak or a silent vendor embedding update that quietly breaks retrieval. The cardinal mistake is grading the system as one black box: when a query fails you won’t know if retrieval missed, the reranker cut the right chunk, or the generator ignored good context.
Why this is the lesson that gets you hired
Hamel Husain’s repeated observation is that evals are the single highest-leverage, most-neglected skill in applied LLM work — and interviewers know it. Anyone can wire up a retriever and a prompt; the senior engineer is the one who can say “here is how I’d know it works, here is where my metrics lie to me, and here is how I gate a deploy so a silent regression never reaches users.” This lesson is that answer.
The retrieval metrics: recall@k, MRR, NDCG, hit rate
Before generation metrics, nail the IR metrics for the retrieval stage, because interviewers ask you to define and choose among them. Recall@k — is the gold chunk in the top-k? — is the ceiling metric: a miss here is unrecoverable downstream, so it’s what the first stage must maximize. Hit rate@k is recall@k for a single gold doc (did we get it, yes/no). MRR (mean reciprocal rank) rewards putting the one right answer high — 1/rank of the first relevant hit — so it’s the metric when there’s a single correct chunk and position matters. NDCG@k handles graded relevance and discounts by position, so it’s the right choice when multiple chunks are relevant to differing degrees (the standard for reranking quality). The senior move: pick the metric to match the task — recall@k for the first stage’s ceiling, MRR/NDCG for ranking quality after a reranker.
code
1Retrieval metrics -- what each one rewards, when to use it23 Metric Rewards Use when4 ----------- ----------------------------- ------------------------------5 Recall@k gold chunk anywhere in top-k first-stage CEILING (must-have)6 Hit rate@k single gold doc retrieved Y/N one correct doc per query7 MRR the FIRST relevant hit, high one right answer, position matters8 NDCG@k graded relevance, discounted multiple relevant; rerank quality9 Precision@k share of top-k that's relevant context noise / token budget1011 Rule: maximize recall@k in stage 1; track MRR/NDCG to judge the reranker.
Interview angle. “Which retrieval metric would you report and why?” is a quiet competence check. The weak answer names one metric for everything; the strong answer says “recall@k as the must-pass ceiling for the retriever — a miss is unrecoverable — and NDCG@k (or MRR for single-answer queries) to judge whether the reranker is actually ordering well.” Tying each metric to the failure it exposes is the signal.
Split the eval: retrieval vs generation
Instrument every stage and run two families of eval. Retrieval: context precision (was the retrieved context relevant?) and context recall (did you fetch all the needed context?) — held independent of the generator. Generation: faithfulness/groundedness (is every claim supported by the context?) and answer relevancy (does it address the question?) — held with retrieval fixed. A real case: end-to-end accuracy dropped after a chunking change, but splitting the metrics showed recall@5 actually rose while precision collapsed — a parser was fragmenting tables into noise. Without the split, the team would have reverted a real improvement.
The logic of the split is a 2×2 you should be able to draw on a whiteboard. Retrieval good + generation good = working. Retrieval bad + generation good = the generator faithfully answered from the wrong context (fix chunking/hybrid/rerank — L2–L4). Retrieval good + generation bad = the right context was there and the model ignored it or hallucinated past it (fix the grounding prompt, or it’s a faithfulness problem). Both bad = start with retrieval, since generation can’t exceed the context it’s given. A single end-to-end score collapses all four into “it failed,” which is why teams tune the wrong stage for weeks. Interview angle. “Your answers are wrong — how do you debug?” → split the eval, locate the failing stage with the 2×2, then name the lever. That structure alone outperforms most candidates.
python
1# The split: grade retrieval and generation SEPARATELY so you can localize a failure.2def evaluate_split(eval_set, retrieve, generate):3 ret_recall, gen_faithful = [], []4 for q, gold_chunk_ids, gold_answer in eval_set:5 ctx = retrieve(q, k=5)6 got_ids = {c["chunk_id"] for c in ctx}7 ret_recall.append(1 if got_ids & set(gold_chunk_ids) else 0) # RETRIEVAL8 ans = generate(q, ctx)9 gen_faithful.append(is_grounded(ans, ctx)) # GENERATION (judge)10 # Read them TOGETHER: low recall + high faithful = right answer, wrong context.11 return sum(ret_recall) / len(eval_set), sum(gen_faithful) / len(eval_set)
Where the metrics lie
Treat each metric as one alarm, not truth. The most-cited example: default RAGAS faithfulness returned null on a large share of hard prompts — by corpus:
code
1RAGAS faithfulness "no output" rate (Cleanlab benchmark)23 FinanceBench (numeric, multi-step) .... 83.5% <- entailment breaks here4 DROP (discrete reasoning) ............. 58.9%5 CovidQA ............................... 21.2%6 PubMedQA .............................. 0.1%78 Lesson: never ship a RAG safety story on faithfulness alone for9 numeric / multi-hop corpora; pair it with a confidence-aware detector.
python
1from ragas import evaluate2from ragas.metrics import faithfulness, answer_relevancy, context_precision, context_recall34# dataset rows: question, answer, contexts, ground_truth5result = evaluate(dataset, metrics=[context_precision, context_recall, # retrieval6 faithfulness, answer_relevancy]) # generation7print(result) # per-metric scores you can gate a deploy on -- but read the caveats above
Why does faithfulness null out rather than score low on hard corpora? Because RAGAS faithfulness works by decomposing the answer into atomic claims and checking each for entailment against the context — and on numeric, multi-step, or discrete-reasoning answers (FinanceBench, DROP) the claim extraction and entailment step itself fails or returns nothing, so you get no score, not a low one. The danger is treating a null as a pass. The senior reading: any metric is one alarm with a known blind spot — faithfulness is blind to numeric/multi-hop, answer-relevancy says nothing about correctness, recall says nothing about whether the model used what you fetched. You triangulate, and for safety-critical numeric corpora you add a confidence-aware hallucination detector rather than trusting entailment alone.
The deploy gate: turn metrics into a CI check
Metrics only protect you if they block a bad deploy. The production pattern: freeze a golden set, define thresholds per metric, and run the eval in CI on every prompt, model, chunker, or index change — failing the build if any metric regresses past a margin against the frozen baseline. This is what converts “evals” from a notebook exercise into the thing that catches a silent regression before users do. Pin the model and prompt versions alongside the gate, so when something does drift you know exactly what changed.
python
1# The eval gate: fail CI if any metric regresses past a margin vs the frozen baseline.2BASELINE = {"context_recall": 0.88, "faithfulness": 0.92, "answer_relevancy": 0.90}3MARGIN = 0.03 # tolerate small noise; block real regressions45def gate(candidate_scores):6 regressions = {7 m: (BASELINE[m], candidate_scores[m])8 for m in BASELINE9 if candidate_scores[m] < BASELINE[m] - MARGIN10 }11 if regressions:12 raise SystemExit("EVAL GATE FAILED: " + str(regressions)) # block the deploy13 print("eval gate passed") # safe to ship; bump the pinned prompt/model version14# Run on EVERY prompt/model/chunker/index change -- silent regressions die here.
The discipline that actually moves quality (Hamel Husain / Shreya Shankar): open-code real production traces (read them, label failures), group failures into a taxonomy (axial coding), and stop at saturation. Prefer binary pass/fail over 1–5 Likert — reviewers regress to the middle and “3 vs 4” means nothing across annotators. Bootstrap with synthetic Q&A from your own chunks, but watch contamination (a model can answer questions it generated) and always build the gold set from the same retriever/index production runs — one legal team read 0.91 recall offline and 0.4 in production because their gold contexts came from chunks the live retriever never surfaced.
That 0.91-vs-0.4 gap is the most important cautionary tale in RAG eval, so understand the mechanism. They built the gold set by asking an LLM to write questions whose answers lived in chunks they hand-picked — but those hand-picked chunks weren’t the ones the live retriever actually surfaced for those questions. So the offline eval graded retrieval against a corpus the production system never used, and reported a fantasy. The rule that prevents it: generate gold questions, but harvest the gold contexts from real production retrieval runs, so your eval measures the system you actually ship. Interview angle. “How do you build a golden set?” The senior answer leads with provenance — error-analysis on real traces, binary labels, and gold contexts drawn from the production retriever — not “I’d have an LLM generate Q&A pairs,” which is the answer that hides the 0.4.
python
1# Error analysis -> taxonomy: the loop that actually moves RAG quality.2# 1) Read real production traces (open coding); label each failure in plain words.3traces = load_production_traces(n=100)4labels = [{"id": t.id, "failure": annotate(t)} for t in traces] # human reads + labels56# 2) Group the free-text failures into a taxonomy (axial coding) until saturation.7# Typical buckets map straight to lessons:8# retrieval_miss -> L2/L3 ranked_too_low -> L4 ignored_context -> grounding9# stale_doc -> L5 injection -> L6 wrong_format -> structured out10# 3) Count buckets -> fix the biggest one first. Re-run. Stop finding NEW buckets = done.11# Prefer BINARY pass/fail per trace; 1-5 Likert collapses to "3" across annotators.
LLM-as-judge: biases & calibration
LLM judges are cheap and indispensable — and quietly biased. Position bias: GPT-4o picks the first of two identical responses ~64% of the time. Verbosity bias: longer scores higher. Self-preference: models favour their own family. Mitigations that work: swap positions and average; use a multi-judge ensemble for high-stakes gates; and run a calibration loop — hand-label a gray-zone set, fold corrections into the judge prompt as few-shot, and track agreement with Cohen’s κ (one team went 62% → 0.78). The hard rule: data used to validate a judge must be hand-verified ground truth.
The conceptual trap with LLM-as-judge is treating the judge’s score as ground truth instead of as another model output that itself needs evaluating. A judge you haven’t validated against human labels can be confidently, systematically wrong — and because it’s automated, it scales that error across your whole eval set. So the discipline is: validate the judge against a small hand-labelled set, report Cohen’s κ (agreement corrected for chance — raw % agreement flatters because two raters agree by luck), iterate the judge prompt until κ clears a bar (~0.7+), and only then trust it to gate deploys. Interview angle. “How do you know your LLM judge is any good?” → measure its agreement with human ground truth via κ and calibrate it; “I trust GPT-4 to grade” without that step is the answer that fails. Also prefer narrow, binary judge questions (“is every claim supported by the context? yes/no”) over a vague 1–10 quality score — they’re far more reliable.
Operate it: observability, drift & injection
In production you trace per-stage spans (retrieval, rerank, context-build, generation) so a regression is attributable. The drift that bites isn’t in your git log — a vendor silently updates an embedding checkpoint and recall sags two weeks later, caught only by per-span hit@k, not end-to-end accuracy. Ship changes behind shadow/canary measured on the production query distribution (offline-only eval over-samples easy queries and misses the revenue-driving long tail).
Two operational disciplines separate teams that sleep at night. (1) Per-stage spans. Log retrieval candidates + scores, rerank order, the assembled context, and the final answer for every request, so when an answer is wrong you can replay exactly which stage failed instead of guessing. (2) Distribution-aware rollout. Offline eval sets skew easy (they’re hand-curated and small), so they systematically miss the hard long tail where users actually churn — the only honest test of an embedding swap, a prompt change, or a new reranker is shadow/canary traffic measured per-span on the real query mix. Interview angle. “Offline scores improved — ship it?” → no; shadow/canary on production traffic watching per-span hit@k, because embedding swaps are exactly the silent-drift case and offline sets flatter the change. Saying “my offline numbers went up” without the canary is the junior answer.
01Per-stage spans — retrieval candidates+scores, rerank order, assembled context, final answer; replay any failure to the exact stage.
02Online quality proxies — abstention rate, citation-validity rate, “I don’t know” rate, thumbs-down; cheap signals that move before accuracy does.
03Freshness/drift — top-k-overlap on a fixed probe set + last-verified age; catches silent vendor checkpoint changes (L5).
04Cost/latency per span — TTFT, ANN/rerank/gen breakdown, $/query; spot the cost cliff (a reasoning model, an uncapped reranker) early.
05Versioned logs — model/prompt/chunker/embedding ids on every request, so a regression is attributable and rollback is a config flip.
You’re swapping embedding vendors; offline RAGAS scores improve. Safest way to ship?
AShip it — offline scores went upBShadow/canary on real production traffic, watching per-span hit@k on the long tail before full rolloutCTrust the vendor’s benchmark numbers
You build a golden set by having an LLM write Q&A from chunks you hand-picked. Offline recall is 0.91; production recall is 0.40. What went wrong?
AThe production retriever is simply worse and needs a bigger embedding modelBThe gold contexts came from hand-picked chunks, not the production retriever, so offline eval measured a corpus the live system never usesCRAGAS computed faithfulness incorrectly
A teammate uses GPT-4 as an LLM judge and reports its scores as the deploy gate, with no human comparison. What’s the senior objection?
AAn unvalidated judge can be systematically wrong at scale — measure its agreement with human ground truth (Cohen’s κ) and calibrate before trusting it to gateBGPT-4 is too cheap to be a reliable judgeCJudges should always use a 1–10 quality score for nuance
Your guardrail reliably blocks jailbreaks in the user prompt, but a study shows that injecting documents into its context flips its judgments in a meaningful fraction of cases — and your retrieved chunks are exactly such documents. What’s the lesson?
AThe guardrail just needs a higher thresholdBDisable retrieval for sensitive queriesCRAG has two input channels — the user prompt AND retrieved context; you must validate retrieved context (and whitelist tool calls), not just the user message
Eval is the most-tested and most-differentiating topic in applied LLM interviews. Interviewers probe: do you split retrieval from generation, do you know your metrics have blind spots, do you validate the judge, do you build golden sets with honest provenance, and do you guard both input channels. Lead with the discipline, then the specific failure each step prevents.
01“How do you evaluate a RAG system?” → split it: retrieval (recall/precision, MRR/NDCG) vs generation (faithfulness/relevancy), never one black-box score.
02“Which retrieval metric and why?” → recall@k as the must-pass ceiling; MRR for single-answer queries; NDCG@k for graded relevance / reranker quality.
03“Answers are wrong — how do you debug?” → split eval + the 2×2 (retrieval good/bad × generation good/bad) to localize the stage, then name the lever.
04“Can you trust RAGAS faithfulness?” → not alone on numeric/multi-hop corpora (it returned null up to ~83.5% on FinanceBench); pair it with a confidence-aware detector.
05“How do you build a golden set?” → error-analysis on real traces, binary labels, and gold CONTEXTS harvested from the production retriever — not LLM-generated pairs over hand-picked chunks.
06“How do you know your LLM judge is good?” → measure agreement with human ground truth via Cohen’s κ, calibrate the prompt, prefer narrow binary questions.
07“Offline scores went up — ship it?” → no; shadow/canary on production traffic per-span, because offline sets over-sample easy queries and embedding swaps drift silently.
08“Prompt injection in RAG?” → two input channels — validate retrieved context, not just the user prompt; whitelist any tool calls the model emits.
Going deeper, the follow-ups that reward operators: “faithful but wrong — how?” (the answer is grounded in an irrelevant retrieved chunk — faithfulness passes, relevance/retrieval fails, which is why you split); “your judge and humans disagree” (fold the disagreements into the judge prompt as few-shot, re-measure κ, repeat); “what about the long tail in eval?” (offline sets miss it — canary on real traffic, and error-analyze tail failures specifically); and “a real injection exploit?” (EchoLeak / CVE-2025-32711, the 2025 zero-click Microsoft 365 Copilot data-exfiltration case, rode in through retrieved content — validate the retrieval channel and constrain tool use). Always name the silent failure the discipline prevents.
Could you stand up a split eval + calibrated judge that gates your next RAG deploy?
Not yetRoughlyYes
Takeaways
Pick the metric for the task: recall@k = ceiling; MRR = single answer; NDCG@k = graded/rerank quality.
Split eval: retrieval (precision/recall) vs generation (faithfulness/relevancy) — the 2×2 localizes the failing stage.
Metrics lie: RAGAS faithfulness breaks on numeric/multi-hop; judges have position/verbosity/self-preference bias — calibrate with κ.
Build golden sets from the production retriever (harvest real contexts), prefer binary labels, error-analyze real traces.
Operate it: per-span tracing, shadow/canary on the real distribution, and guardrails on retrieved context (not just the user prompt).
Finally: assemble everything into a production RAG you could design on a whiteboard.