The accuracy number you trust for a classifier evaporates here: same input, different output; no single ground truth; and an offline score that lies about production. Why the old metrics break, the eval maturity ladder, and the discipline that replaces them.
The metric you trusted just stopped working
As a data scientist you have a reflex: split train/test, compute accuracy or F1 or AUC, ship the model that wins. That reflex breaks completely on a generative system. The output is a free-form string sampled from a distribution, so the same input gives a different answer twice; there is rarely one correct output to compare against; and the offline number you compute on a curated set systematically disagrees with how the system behaves on live traffic. This lesson is about why the comfortable machinery fails — and the new discipline (error analysis, golden sets, calibrated judges) that replaces it. Skip it and every later metric is a number with no error bars.
Start with the root cause, because interviewers do. An LLM is autoregressive: given the tokens so far it scores every token in its vocabulary, applies a softmax, and samples one — so the output is a draw from a distribution, not a return value. temperature=0 makes that draw greedy (always the argmax) but not deterministic: batched inference, floating-point non-associativity on GPUs, mixture-of-experts routing, and silent provider model updates all make the same prompt return different text on different days. Your evaluation harness therefore cannot use assert output == expected. The unit of truth has to move from string equality to graded judgement against criteria, and that single shift is what makes GenAI eval a different sport from the precision/recall you already know.
Three properties of generative systems break the classical toolkit, and they organise this whole track. (1) Non-determinism — the same input yields different outputs, so a single run is a sample, not a measurement; you need aggregate metrics with confidence intervals, not point estimates. (2) No single ground truth — for "summarise this ticket" or "answer this question" there are many acceptable outputs and many unacceptable ones that look similar, so exact-match and lexical-overlap metrics (BLEU, ROUGE) measure the wrong thing. (3) The offline-online gap — your curated eval set is a slice, and production traffic drifts away from it, so a green offline score can sit on top of a quietly degrading product. Each property gets its own remediation later; first, feel why the old metrics genuinely cannot cope.
The instinct is to reach for an off-the-shelf text metric. Resist it. BLEU and ROUGE measure n-gram overlap against a reference string — they reward surface word matching, not meaning. A summary that paraphrases the reference perfectly scores low; a summary that copies the reference’s words while inverting a key fact scores high. Eugene Yan’s task-specific survey is blunt about this: for summarisation and translation, lexical-overlap metrics correlate weakly with human judgement on anything long-form, and the field has moved to NLI-based factual-consistency checks (BARTScore) and learned metrics (BLEURT, COMET/Kiwi) precisely because ROUGE and BLEU miss the failures that matter. Shipping a "general" overlap score on a clinical or financial system hides exactly the errors you most need to catch.
code
1Why classical text metrics fail on generative output23 Metric Measures Fails when...4 ---------- --------------------- ---------------------------------5 Exact match string equality any valid paraphrase (almost always)6 BLEU n-gram precision meaning preserved, words differ7 ROUGE n-gram recall words copied, a fact inverted8 BERTScore embedding overlap weak on long-form; misses entailment9 Accuracy/F1 label == label there is no single label to compare1011 The through-line: surface similarity is not correctness. For free-form text12 you need an entailment/criteria check, not a token-overlap score.
There is a subtler trap for data scientists specifically: class imbalance hides under aggregate accuracy. If 97% of your support answers are fine and 3% are dangerously wrong, a judge that blindly says "pass" scores 97% accuracy while missing every failure that can hurt you. This is why the research repeatedly pushes recall on the risky class over headline accuracy, and why "we get 90% agreement" is a non-answer (more on that with Cohen’s kappa in L4). Interview angle. When an interviewer offers you a single accuracy number, the senior move is to ask "on what label distribution, and what’s the recall on the failure class?" — surfacing imbalance is a reliable signal you think like an evaluator, not a leaderboard-chaser.
Non-determinism: from point estimates to distributions
Because output is a sample, one run tells you almost nothing. Run the same 200-example eval twice and the pass rate moves — not because the system changed, but because decoding is stochastic and the judge is itself an LLM. The discipline that replaces "run it once" is the discipline you already know from A/B testing: report a rate with a confidence interval, and size the sample so the interval is smaller than the effect you care about. Eugene Yan’s rule of thumb is worth memorising — at a 3% defect rate, ~200 sampled traces give roughly ±2.4% at 95% confidence; if you need to detect a 1% regression you need far more, or stratified sampling that concentrates on the failure-prone slices.
python
1# Treat an eval score like any proportion estimate: it has a confidence interval.2import math34def wilson_ci(passes, n, z=1.96):5 if n == 0:6 return (0.0, 0.0)7 p = passes / n8 denom = 1 + z*z/n9 center = (p + z*z/(2*n)) / denom10 half = (z * math.sqrt(p*(1-p)/n + z*z/(4*n*n))) / denom11 return (round(center - half, 3), round(center + half, 3))1213# A judge passes 188/200 traces -> 94% pass, but the band is what you ship on:14print(wilson_ci(188, 200)) # ~ (0.90, 0.96) -> a 2-point "regression" is NOISE15# Rule of thumb: ~200 samples ~ +/-2.4% at a 3% defect rate. Want tighter? Sample more.
A second consequence: a single eval run cannot catch a silent provider regression. Anthropic’s April 2026 Claude Code postmortem ("An update on recent Claude Code quality reports") traced weeks of "it got dumber" reports to three interacting changes shipped weeks apart — a default reasoning-effort cut (high→medium), an idle-session thinking-context bug, and a verbosity-reduction prompt tweak — none individually large, and each hitting a different slice of traffic on a different schedule. The only defense is to pin model and prompt versions and run a fixed eval on every change, then watch the score over time with its CI. For a data scientist this is the familiar move of holding the test set fixed and tracking the metric run-over-run — just applied to a non-stationary system you don’t control.
No single ground truth: the open-coding discipline
If there is no gold answer to diff against, where does the signal come from? From looking at your data — the activity every serious practitioner names as the highest-ROI move in the whole field. Hamel Husain’s method is borrowed straight from qualitative social science: sample ~50–100 real (or realistic) traces, open-code them (write a free-form note on what went wrong with each), then axial-code those notes into a small taxonomy of failure categories. You stop when you hit theoretical saturation — roughly when 20 consecutive traces surface no new category. The output is not a metric yet; it is the map of how your system fails, and every later metric, rubric, and judge is built to measure a category on that map.
This inverts the data scientist’s usual order. You are not picking a metric and then measuring; you are discovering the failure modes first, then defining metrics that detect them. A typical taxonomy for a support assistant after a day of open coding: wrong intent routed, missing a required disclaimer, hallucinated a policy, answered a different question, correct-but-too-long, refused a valid request. Each bucket becomes one binary assertion ("did the response include the required disclaimer: yes/no"). Interview angle. "Your team has no labels and no eval — what do you do in week one?" The strong answer leads with error analysis on real traces and a 5–8 bucket taxonomy, not "stand up LangSmith" or "order 500 labels from a vendor." Reaching for infrastructure before looking at data is the single most common junior tell.
The moat in GenAI eval is the labeling habit, not the metric library. Every practitioner lineage — Hamel, Eugene, Shreya — converges on the same first move: look at your data, by hand, before you automate anything.
One hard rule the research is unanimous on: never outsource the initial open coding or ground-truth labeling to an LLM or a vendor without your own eyes on the data. A strong model will happily generate a thousand "labels" that are silver, not gold — confidently wrong in exactly the ways you were trying to find. The human owns the rubric and the taxonomy; the LLM can later operate the rubric at scale, but it cannot author what "good" means for your product. That separation of roles — human defines, machine scales — is the spine of the whole track.
The offline-online gap: your test set is a slice
The third failure is the one that bites in production. Your offline golden set is a fixed slice of the input distribution, but live traffic is non-stationary: users ask new things, the retrieval corpus changes, prompts get refactored, and the provider swaps the model under you. Fiddler decomposes this into prompt drift (inputs deviate from what you tuned for), model drift (responses change), and user drift (intents shift); Braintrust adds judge drift (your human-aligned judge and real quality decouple). The result: a green offline score sitting on top of a product that is quietly getting worse — the GenAI version of training-serving skew, and just as silent.
The fix is a feedback loop, not a bigger test set. Offline evals gate releases against a held-out golden set; online evals score sampled live traffic on the same behavioural dimensions; and — the part juniors miss — production failures get promoted back into the offline set so a regression that escaped once can never escape silently again. Hamel’s cadence makes it concrete: review 100+ fresh traces every 2–4 weeks plus 10–20 weekly outliers, and write a new assertion for whatever the human finds. Offline and online are complements, not substitutes — treating them as separate worlds is a documented failure mode. We build each half in L5.
code
1The offline<->online loop (the thing that actually keeps a system honest)23 OFFLINE (release gate) ONLINE (live monitor)4 -------------------------- --------------------------------5 fixed golden set 1-5% sample of real traffic6 binary assertions in CI same dimensions, reference-free judge7 block deploy on regression alert on score outside its baseline CI8 ^ |9 | v10 +------ promote production failures into the golden set1112 A green offline score on a drifting product is the GenAI training-serving skew.
The eval maturity ladder — climb it in order
Most teams fail because they reach for the most expensive lever before the cheap ones underneath it work. Hamel’s ladder makes the ordering explicit, and it is a dependency chain, not a menu. Level 1 — assertions: cheap deterministic checks that fail the build on a hard violation (a regex for a leaked UUID or PII pattern, a must-contain refusal string, a JSON-schema validation). His canonical case: a chatbot that was 95% accurate yet occasionally leaked internal UUIDs; the fix was a one-line Level-1 assert, shipped as a CI gate in an afternoon. Level 2 — human + model grading: binary rubrics and calibrated LLM judges for the subjective residue assertions can’t catch. Level 3 — A/B testing: live experiments, only once Levels 1–2 give you a stable regression floor.
The ladder is load-bearing because skipping a level means re-paying for it later. A/B testing a non-deterministic system with no regression floor produces confidence intervals too wide to drive a shipping decision — you are measuring noise. If your team cannot articulate its Level-1 invariants, the Level-2 judge number is meaningless and the Level-3 A/B lift is uninterpretable. Eugene Yan collapses the same idea into a tight loop: label a small set → align the judge to it → re-run the harness on every config change, where "align" is itself a metric you re-validate whenever the app, model, or prompt moves. Interview angle. "Where do you start improving an AI product?" → error analysis → Level-1 assertions → calibrated Level-2 judges → only then A/B. Naming the ladder, in order, signals you’ve shipped evals, not just read about them.
A data scientist proposes evaluating a summarisation feature with ROUGE against a set of reference summaries, picking the prompt with the highest score. What’s the senior objection?
AROUGE is fine; just average it over enough references to reduce varianceBROUGE measures word overlap, not correctness — a paraphrase scores low and a near-copy with a flipped fact scores high; use criteria-based grading + a factual-consistency (NLI) checkCSwitch to BLEU, which is designed for generation tasks
You run your 200-example judge eval on Monday and get 94% pass; on Tuesday, with no code change, it reads 92%. A teammate wants to roll back yesterday’s prompt. Best response?
A2 points is within the confidence band at n=200 (~±2.4% at a 3% defect rate) — it’s noise, not a regression; don’t roll back, and size the sample to the effect you need to detectBRoll back immediately — any drop in the eval score is a regressionCRe-run until you get 94% again, then ship
You inherit a customer-support assistant with no labels and no eval rig. You have one week. What ships first?
AStand up LangSmith and a dashboard so you can start collecting tracesBOrder 500 labels from a vendor and train a quality classifierCOpen-code ~50–100 real traces into a 5–8 bucket failure taxonomy, turn each bucket into a binary assertion, and gate CI on them
Users report the bot "got noticeably worse this week," but your team deployed nothing and the offline eval is still green. Most likely cause and the right safeguard?
ARandom LLM variance — nothing to do, it’ll average outBThe offline set went stale or the provider silently updated the model; pin versions, score sampled live traffic online, and promote the new failures into the offline setCYour context window shrank, truncating inputs
A teammate says "let’s just A/B test three prompt variants in production and ship the winner" — the team has no assertions and no golden set yet. What’s the strongest objection?
AA/B testing is always wrong for LLMs; only offline evals are validBWithout a regression floor, A/B on a non-deterministic system gives confidence intervals too wide to drive a shipping decision — build Level-1 assertions and a calibrated judge first, then A/BCA/B testing needs millions of users, which you don’t have, so use a survey instead
A GenAI-eval screen for a data scientist tests whether you’ve unlearned the leaderboard reflex and adopted the evaluator’s discipline. Interviewers probe four things: do you know why classical metrics fail, do you start with error analysis on real data, do you treat scores as statistics with error bars, and can you explain the offline-online gap and the loop that closes it. Lead with the mechanism, then the production implication, and use numbers.
01“Why is evaluating an LLM harder than evaluating a classifier?” → non-determinism (sample, not return value), no single ground truth, and an offline score that drifts from production.
02“Is temperature 0 deterministic?” → no — greedy decoding, but batching/FP/MoE/provider updates still vary it; design for variation, never assert exact strings.
03“Why not BLEU/ROUGE for a chatbot?” → they score word overlap, not correctness; a paraphrase fails, a near-copy with a flipped fact passes — use criteria-based + NLI checks.
04“Where do you start with no evals?” → open-code ~50–100 real traces into a failure taxonomy, then one binary assertion per bucket — data first, not infrastructure.
05“How big should the eval sample be?” → size it to the effect; ~200 traces ≈ ±2.4% at a 3% defect rate, so report a CI and stratify for rare failures.
06“What’s the eval maturity ladder?” → L1 assertions → L2 human+judge grading → L3 A/B; it’s a dependency chain, skipping a level re-bills you later.
07“Why can a green offline score still mean a degrading product?” → the offline set is a fixed slice; prompt/model/user/judge drift — score live traffic and loop failures back.
08“Users say it got worse but you shipped nothing — why?” → silent provider update or drift; pin versions and run a fixed eval on every change to catch it.
To go deeper, expect follow-ups that separate "read a blog" from "shipped one": "how do you keep your golden set from going stale?" (scheduled refresh from production plus a quarterly "do these scenarios still match what we ship?" review — L2/L5); "your judge and your humans drift apart over time — what do you do?" (re-align on a rolling window, monitor Cohen’s kappa — L3/L4); and "how do you convince a skeptical PM the metric reflects what users care about?" (show the chain: humans agree on the rubric → the judge matches humans → the judge tracks a business signal like re-open rate — L4). In every case, name the artifact and the number, not a vibe.
Could you explain to a PM why your classifier-eval playbook fails on an LLM, and name the discipline that replaces it?
New to itGetting thereConfident
Takeaways
Output is a sample, not a return value — temperature 0 is greedy, not deterministic; never assert exact strings.
Classical metrics (BLEU/ROUGE/exact-match/accuracy) measure surface similarity, not correctness — they hide the failures that matter.
No single ground truth → start with error analysis: open-code real traces into a failure taxonomy, by hand.
An eval score is a statistic — report a sample size and a confidence interval (~200 traces ≈ ±2.4% at 3% defect).
The offline set is a slice; the offline-online gap is GenAI training-serving skew — close it with a feedback loop.
Climb the maturity ladder in order: L1 assertions → L2 human+judge → L3 A/B. Skipping a level re-bills you later.
Next: building the golden dataset and the rubric — the reference data every metric is computed against.