The three evaluation modalities and the right allocation between them, the offline regression / sampled-production / human-calibration split, LLM-as-judge as the only thing that scales — and its systematic traps (position bias, verbosity bias, self-preference) plus how elite teams calibrate it against human golden sets.
Three modalities, not one activity
The second-biggest AI-PM eval mistake is treating “evaluation” as a single thing. It is three modalities, each catching a different class of failure: offline (a held-out set, run pre-release, catches regressions), online (sampled live traffic, catches drift you never imagined), and human (expert labels, the ground truth everything else is calibrated against). Get the allocation wrong — try to human-review production traffic, or trust a single offline number — and your eval program is either unaffordable or blind. This lesson designs the mix, then goes deep on the judge that makes it scale and the biases that make the judge lie.
Hamel Husain and Shreya Shankar give the cleanest allocation rule, and it’s a function of cost, latency, and whether a failure reproduces. CI / offline should “favor assertions or other deterministic checks” — they’re cheap and reproducible, so they’re the right tool for regression. Production / online should “sample live traces” and score them with “reference-free evaluators like LLM-as-judge,” because that’s the only way to see distribution drift. Human labels are reserved for “validating those judges against human annotations” — i.e. calibration, not ongoing scoring. The shape that falls out: roughly ~80% deterministic CI assertions, ~15% LLM-as-judge on sampled traffic, ~5% human review, with the human layer serving as the benchmark the judge is measured against.
code
1THE EVAL MODALITY MIX (allocate by cost / reproducibility)23 MODALITY ~SHARE TOOL CATCHES4 -------- ------ -------------------------- ---------------------------------5 Offline ~80% deterministic assertions regressions (cheap, reproducible)6 /CI on schema, format, facts7 Online ~15% LLM-as-judge on SAMPLED drift you never wrote a test for8 live traces (reference-free)9 Human ~5% expert labels / golden set GROUND TRUTH; calibrates the judge1011 Galileo: human-only scoring is "mathematically impossible for production12 systems processing millions of requests." Humans set the bar; judges scale it.
Why not just use humans? Galileo is blunt: human-only evaluation is “mathematically impossible for production systems processing millions of requests.” And why not just trust the offline set? Because a frozen CI set decays — Hamel & Shankar recommend a fresh look at “100+ traces” every 2–4 weeks precisely because user behaviour shifts in ways your fixed set never captured. The two failure modes are symmetric: offline-only goes stale and misses drift; human-only doesn’t scale. The mix exists to cover both, with the human layer as the small, expensive anchor that keeps the cheap, scalable layers honest.
Designing the eval set: roles, dimensions, and failure hypotheses
A grader is only as good as its instructions, and Aman Khan’s eval formula (running in Lenny’s Newsletter) is the most actionable PM template: a judge prompt sets the role, the context, the goal, and the terminology/label definitions. That maps directly onto OpenAI’s testing_criteria model, where graders use explicit textual criteria to decide if an output is correct. The PM writes this rubric; engineering implements and runs it. If you can’t articulate the difference between an eval (a scored rubric against ground truth) and a vibe check (a subjective thumbs-up) in one sentence, the interviewer marks you down.
For the set itself, Hamel & Shankar’s “Lucy” example (a real-estate assistant) is the canonical pattern: build coverage with dimensions (cuisine type → property type, neighbourhood, price band) and write explicit failure hypotheses up front. Dimensions let a small set (100+ examples) cover a large behaviour space; failure hypotheses force you to write down the negative cases (“should not do X”), not just the positives. This is the same discipline as Notion’s split between regression evals (broad CI) and frontier evals (narrow tasks designed to differentiate competing models). Interview angle. “How would you build the eval set?” → dimensions for coverage, failure hypotheses for negatives, harvested production traces for the true distribution — never “collect more data.”
One non-negotiable that PMs get wrong: harvest the gold contexts from your production retriever’s runs, not from passages you hand-picked, or your offline scores will read far higher than reality. The eval has to see what production sees. The categories of grader to name in an interview: human evaluation (ground truth on subjective quality), LLM-as-judge (cost-effective scale on open-ended quality), grounded evaluation (output checked against a DB / source of truth — the no-hallucination check), and automated assertions (length, format, JSON-schema, SQL-validity — cheap regression on every release).
LLM-as-judge: the only thing that scales — and its four biases
LLM-as-judge is what makes online evaluation affordable: a separate model scores outputs against your rubric, so you can grade sampled live traffic continuously. But it is not trustworthy in isolation, and naming its systematic biases is a high-signal interview move. Galileo catalogues four: verbosity bias (longer answers score higher regardless of quality), self-preference (a judge favours outputs from its own model family), position bias (in pairwise comparison the option presented first wins more often), and reference-answer bias (the judge over-weights surface similarity to a provided reference). Each is a way the judge can be confidently wrong at scale.
code
1LLM-AS-JUDGE: the four systematic biases (Galileo)23 BIAS WHAT HAPPENS MITIGATION4 ----------------- ------------------------------- ----------------------------5 Position bias option shown FIRST wins more in randomize order; score BOTH6 pairwise comparisons orderings and average7 Verbosity bias longer answer scores higher length-normalize; penalize8 regardless of correctness padding in the rubric9 Self-preference judge favours its own model use a DIFFERENT family as10 family's outputs judge; cross-check11 Reference-answer over-weights surface similarity rubric scores reasoning, not12 to the provided reference string overlap1314 + human counter-bias: reviewers rate confident-but-WRONG outputs 15-20% higher.
The mitigations are concrete and the kind of thing a strong candidate rattles off. Position bias → randomise option order and, for pairwise judging, score both orderings and average (or count a win only if it’s consistent across swaps). Verbosity bias → length-normalise and write the rubric to penalise padding. Self-preference → use a judge from a different model family than the one being evaluated, and cross-check. Reference-answer bias → score the reasoning against criteria, not string overlap with a reference. And the human side has its own counter-bias Galileo quantifies: reviewers “rate confident-but-incorrect outputs 15–20% higher,” which is exactly why neither humans nor judges are trustworthy alone.
Calibrating the judge against a human golden set
The architecture that resolves the standoff is hybrid and calibrated: humans build a golden dataset, the LLM judge runs continuously against production, and you periodically measure the judge’s agreement with the humans on the golden set — Galileo describes elite teams running “quarterly review sessions to identify systematic errors and update scoring rubrics.” The division of labour is explicit: human evaluation is best for defining the initial quality rubric, handling edge cases that need domain expertise, and judging subjective outputs; LLM judges are for high-volume regression testing and continuous CI/CD evaluation. Humans set and recalibrate the bar; the judge scales it between recalibrations.
Concretely, calibration is a number you track: the judge-vs-human agreement rate on the golden set. If it drops, you don’t trust the judge’s production scores until you’ve re-tuned the rubric or the judge model. Notion runs exactly this combination — “LLM-as-judge for quality assessment and human review for nuanced judgment,” with “narrow, well-defined evaluations” backed by “heuristic scorers” for the objective checks — which is how a 70-engineer org keeps a consistent quality bar without humans grading every trace. Interview angle. “When the LLM judge disagrees with the human spot-check, what do you do?” → treat the human as ground truth, measure the agreement rate, and re-calibrate the rubric/judge before trusting its scores again — don’t average a biased judge with a human and call it settled.
A PM proposes having the support team manually review every production conversation to score quality. Why does a senior reviewer reject this as the primary eval?
AManual review is too subjective to ever be usefulBHuman-only scoring doesn’t scale to production volume — reserve humans for a golden set, run LLM-as-judge on sampled traffic, and deterministic assertions in CICThe support team isn’t qualified to judge AI output
Your LLM judge does pairwise “A vs B” comparisons and a new variant keeps winning. An engineer notices the new variant is always shown as option A. Most likely issue?
APosition bias — the first option wins more often; randomize order and score both orderings, counting a win only if it survives the swapBThe new variant is genuinely better and you should ship itCVerbosity bias — the new variant must be longer
You evaluate a GPT-family assistant using a judge from the same GPT family, and scores look great. What’s the senior concern?
ANothing — same-family judging is the most accurate optionBLatency — same-family judges are slowerCSelf-preference bias — the judge favours its own family’s outputs; use a different-family judge and cross-check against a human golden set
Where should an AI PM spend the most human-labelling effort?
ABuilding and maintaining a high-quality golden dataset that defines the rubric and calibrates the LLM judgeBManually scoring as many production traces as possible each weekCRe-labelling the offline regression set every single build
A frozen offline eval set has shown green for three months, but users report new failure types. What does this reveal about the eval design?
AThe offline set is wrong and should be deletedBA fixed CI set decays — add sampled-production evaluation (LLM-as-judge on 100+ fresh traces every 2–4 weeks) to catch drift the frozen set misses, and fold new failures back inCThe model regressed and the provider should be contacted
The eval-harness question is the most-tested AI-PM “system design” probe, and KORE1’s strong/weak contrast is the cleanest teaching artifact: strong answers separate offline (golden datasets, precision/recall, faithfulness) from online (task-completion, escalation rate, correction rate) and name a fallback when the model is wrong; weak answers say “accuracy” and “user feedback.” Lead with the three modalities and the allocation, then the judge, then the calibration.
01“Walk me through your eval harness.” → offline (golden/regression, precision/recall/faithfulness) + online (task-completion, escalation, correction rate) + human calibration; name the ship gate.
02“Offline vs online eval?” → offline = held-out set pre-release, catches regressions; online = sampled live traces, catches drift; they answer different questions.
03“How do you scale evaluation to millions of requests?” → LLM-as-judge on sampled traffic for volume; humans only for the golden set and calibration.
04“What are the traps of LLM-as-judge?” → position, verbosity, self-preference, reference-answer bias — mitigate with order-randomization, length-norm, different-family judge, reasoning rubric.
05“When the judge disagrees with the human spot-check?” → human is ground truth; track the agreement rate and recalibrate the rubric/judge before trusting it.
06“How do you build the eval set?” → dimensions for coverage + failure hypotheses for negatives + harvested production traces for the true distribution.
07“Eval vs vibe check?” → an eval is a scored rubric against ground truth with a gate; a vibe check is a subjective thumbs-up that doesn’t gate anything.
08“How do you evaluate a RAG system?” → evaluate retrieval (recall@k) BEFORE generation (faithfulness); a generation metric can’t fix a retrieval miss.
Going deeper, expect: “how often do you refresh the golden set, and who decides ground truth?” (on a cadence, with a named domain owner — ground truth is a role, not a vote); “what’s your eval coverage of minority-language or under-represented slices?” (Notion’s multilingual needle test — coverage by dimension is the answer); and “how do you keep the judge from drifting?” (track judge-vs-human agreement on the golden set and recalibrate quarterly). The thread is always: humans anchor, judges scale, a tracked number keeps them aligned.
Could you design the offline/online/human mix for an AI feature and name the LLM-as-judge biases plus how you’d calibrate the judge?
New to itGetting thereConfident
Takeaways
Evaluation is three modalities: offline (regressions), online (drift), human (ground truth) — allocate ~80/15/5.
Build the set with dimensions (coverage) + failure hypotheses (negatives) + harvested production traces (true distribution).
LLM-as-judge is the only thing that scales online — but it carries position, verbosity, self-preference, and reference-answer bias.
Mitigate mechanically: randomize order + score both, length-normalize, use a different-family judge, score reasoning not string overlap.
Humans over-reward confident-but-wrong answers ~15–20% — so neither humans nor judges are trustworthy alone.
Calibrate: humans build the golden set, the judge scales it, and a tracked agreement rate (recalibrated quarterly) keeps it honest.
Next: responsible AI & risk — the NIST GenAI profile, guardrails, and red-teaming a probabilistic product.