The reference data is the eval — and most teams mis-build both halves. Sampling that finds real failures, rubrics that two annotators actually agree on (binary over Likert), the dimensions-and-tuples trick for coverage, and reference-based vs reference-free metrics as a cost-modeling choice.
The artifact every metric depends on
A "golden dataset" is not a CSV of model outputs — it is a labeled set of examples with criteria, where the criteria are stable enough that two annotators (human or model) agree. Both halves — the examples and the rubric — must be designed deliberately, and mis-designing either invalidates every metric downstream. This is the unglamorous foundation: get the reference data right and your judge has ground truth; get it wrong and you ship a confident number measuring nothing.
Why this is the highest-leverage lesson in the track: the dataset’s error distribution is dominated by ambiguity at the rubric boundary, not by coverage gaps. A 100-example set with stable, well-anchored criteria produces a more reliable judge than a 1,000-example set with vague criteria — because the judge’s mistakes cluster on the cases where the rubric itself is unclear. So the senior priority order is counterintuitive: invest in rubric anchors before raising sample size. Doubling the data rarely doubles the signal; sharpening the rubric almost always does.
Interview angle. "Build a golden dataset for a new assistant from scratch — walk me through six weeks." The strong answer sequences error-analysis-first sampling, a binary rubric with written anchors, a coverage plan (dimensions-and-tuples), and an explicit fail/pass balance — and names sample sizes. The weak answer is "use GPT-4 to generate 10,000 questions," which produces silver data mistaken for gold and skips the place you actually learn how the system fails.
Sampling: coverage of failure categories, not raw size
Hamel’s sizing rule is empirical, not theoretical: review at least 100 traces, and stop labeling once ~20 consecutive traces expose no new failure category. The implication is that category coverage, not row count, is the target. Three sampling modes work in combination. (1) Random pulls the long tail and surfaces silent categories you didn’t anticipate. (2) Stressed-query deliberately probes constraints — adversarial inputs, multilingual, edge-case formats. (3) Active sampling from production ranks traces by failure signal: thumbs-down, regenerations, low embedding similarity to known-good answers, refusals. The mix is what gives you both the head and the tail.
The crucial design choice juniors get wrong: the golden set must include both the pass and fail regimes. Eugene Yan’s rule is ~200+ labeled examples with 50–100 deliberately-sampled failures, so the judge can be calibrated against what "bad" looks like, not just what "good" looks like. Failure-only sampling leaves you blind to regressions on the long tail of acceptable behaviour; pass-only sampling means your judge has never seen the errors it’s supposed to catch. Interview angle. "Why not just label the failures you found in production and skip the must-pass scenarios?" → because a judge tuned only on failures can’t recognise a healthy answer, and you lose the ability to detect regressions on the cases that currently work.
code
1Sizing the golden set (the numbers to say out loud)23 Purpose Size Stop rule / note4 ---------------------- ----------------- --------------------------------5 Failure-mode discovery review >=100 stop ~20 traces w/ no new category6 CI regression suite ~100 labeled category coverage > raw count7 Judge calibration >=200, 50-100 fail needs BOTH pass and fail regimes8 Per-scenario power ~246 / scenario 80% pass rate, +/-5%, 95% conf910 Coverage, not volume. A 100-row set with stable anchors beats a 1,000-row11 set with vague criteria — judge error lives at the rubric boundary.
For coverage without a thousand random questions, use the dimensions-and-tuples pattern: define the axes that matter (e.g. dietary restriction, multi-doc reasoning, ambiguity, out-of-scope, language) and sample tuples of axis values, rather than asking an LLM for "1,000 diverse questions." This guarantees the hard combinations appear instead of trusting a generator’s diversity. And when you do synthesise inputs, the research has a sharp counter-intuition: synthetic failures from a strong model are often out-of-distribution relative to real failures — a frontier model fails in different, cleaner ways than your system does. Use smaller, weaker models to generate organic-looking failures, and always SME-spot-check a sample of any synthetic data before trusting it.
Rubric design: binary beats Likert
The single most-cited anti-pattern in the entire field is the Likert 1–5 rubric, and it fails for two mechanical reasons. (1) Central-tendency collapse: annotators without deep specialisation default to 3, so the middle rating becomes a noise anchor that obscures the binary that actually matters — it worked or it didn’t. (2) Cross-annotator agreement craters: a Likert requires each annotator to internalise an undocumented calibration curve (what is a 3 vs a 4?), so two people independently rating the same item 3 and 5 land at only ~0.2–0.3 Cohen’s kappa. The replacement is a one-sentence binary pass/fail with an explicit "I don’t know" path:
Pass if the response cites a primary source for the prevalence claim; otherwise Fail. Mark "I don’t know" if the response makes no prevalence claim. — the "I don’t know" branch is excluded from the score but counted in failure-mode review, so ambiguity doesn’t pollute the metric.
Two more rubric rules the research is firm on. One criterion per item — never compound criteria like "cites a source AND is concise AND avoids jargon," because a failure on any one collapses three signals into a single uninterpretable Fail. Split them into three independent binary checks. And supply 2–3 written anchor examples per label — these are the calibration target the judge is built against and the single highest-leverage artifact in the system; a rubric with no anchors ships a metric with no ground truth. Binary judgments paired with these anchors are what give you the high inter-rater reliability that a Likert scale structurally cannot.
code
1Rubric: do / don't (this table is the whole lesson in miniature)23 Property Do Anti-pattern4 -------------- ------------------------------- ----------------------------5 Granularity binary pass/fail (+ "unknown") Likert 1-5; 0-10 numeric6 Definition one sentence + 2-3 anchors multi-paragraph, no examples7 Independence one criterion per item "cites AND concise AND ..."8 Edge cases "I don't know" excluded from score coerced into pass/fail9 Per dimension one evaluator per dimension a single "God Evaluator"1011 Binary + anchors gives inter-rater kappa a Likert structurally can't reach.
Store more than the row. A production-grade golden set carries a schema per case: the input plus context/sources (URLs, doc ids, timestamps) for reproducibility, scenario tags (intent, persona, difficulty, language, safety bucket) for stratified sampling and drift detection, expected elements (required entities, steps, citations) that drive a groundedness judge, and governance (reviewer, audit trail, retention, risk tag) for compliance. Interview angle. "What does a golden-set record contain?" Naming schema fields — not just "question and answer" — signals you’ve built one that survives drift and audit, not a throwaway spreadsheet.
python
1# A golden-set record is examples + criteria + provenance -- not a Q/A pair.2record = {3 "id": "supt-0431",4 "input": "Can I get a refund on a gift-card purchase?",5 "context_sources": ["policy/refunds#gift-cards"], # reproducibility6 "scenario_tags": {"intent": "refund", "persona": "consumer",7 "difficulty": "edge", "lang": "en", "safety": "none"},8 "rubric": [ # one criterion each9 {"name": "on_policy", "ask": "Does it state the gift-card refund rule?"},10 {"name": "cited", "ask": "Does it cite the refunds policy?"},11 ],12 "expected_elements": ["non-refundable", "refunds#gift-cards"],13 "label": {"on_policy": True, "cited": True}, # human gold14 "governance": {"reviewer": "sme_amir", "risk": "low", "reviewed_at": "2026-06"},15}16# Stratify by scenario_tags; calibrate the judge against rubric+label; audit via governance.
The labeling protocol: agreement before scale
Labels are only "gold" if humans agree on them, so the protocol matters as much as the rubric. The discipline that survives a room of annotators: run a pilot round on a shared overlap set (≈30 items two people both label), compute inter-annotator agreement with Cohen’s kappa, and only lock the instructions once kappa clears a bar (Eugene Yan’s bands: 0.4–0.6 substantial, >0.7 excellent). Disagreements in the pilot are the most valuable output — they expose ambiguous rubric boundaries you fix by sharpening the criterion or adding an anchor. Then write down the rationales so they survive a new annotator joining: a protocol that lives only in one labeler’s head is a single point of failure.
The order is load-bearing: lock the rubric before reviewing any outputs. Shreya Shankar names the failure when you don’t — criteria drift, where annotators quietly refine their standard after seeing model outputs, so the bar moves under you and yesterday’s labels stop meaning what today’s do. The same caution scopes synthetic data: an LLM can operate the locked rubric to pre-label at scale, but a human spot-checks a sample to estimate the synthetic-label accuracy (15–20% is a common audit fraction) before any of it counts as gold. Interview angle. "Walk me through your labeling protocol with three annotators." → pilot round, kappa on an overlap set, lock instructions + rationales, then scale — naming kappa and criteria-drift here is a strong senior signal.
python
1# Gate the labeling protocol on inter-annotator agreement BEFORE you scale.2from collections import Counter34def cohens_kappa(a, b): # a, b: lists of binary labels5 n = len(a); po = sum(x == y for x, y in zip(a, b)) / n6 ca, cb = Counter(a), Counter(b)7 pe = sum((ca[k]/n) * (cb[k]/n) for k in set(a) | set(b)) # chance agreement8 return (po - pe) / (1 - pe) if pe < 1 else 1.0910pilot_a = [1,1,0,1,0,0,1,1,0,1] # annotator A on the 30-item overlap11pilot_b = [1,1,0,1,1,0,1,0,0,1] # annotator B12k = cohens_kappa(pilot_a, pilot_b)13assert k >= 0.6, f"kappa {k:.2f} too low -- sharpen the rubric, re-pilot, don't scale yet"14# Only after the pilot clears the bar: lock instructions + rationales, then label at scale.
Reference-based vs reference-free: a cost-modeling choice
Treating these as interchangeable hides the real tradeoff. Reference-based metrics are cheap to build and brittle to drift; reference-free metrics are expensive to build and durable to drift. The choice is not which is "better" — it’s which cost curve you can sustain for the life of the application. A reference-based metric compares output to a known-good answer (exact/contains match, lexical overlap, embedding similarity, or learned metrics like BLEURT/COMET); every reference must be hand-written or curated, so the marginal cost is linear — you label every new example. A reference-free metric judges the output against the question and any context, typically via an LLM judge; the build cost is high (judge calibration) but the marginal cost is ~free because the judge generalises to novel inputs.
The deeper distinction is the shape of the signal. Reference-based metrics are instance-level — they tell you this specific example regressed. Reference-free metrics are distribution-level — they tell you this prompt category is failing more often. Both signals are needed, and the trap is treating one as a substitute for the other. The senior strategy is banded: a reference-based regression suite (100–200 hard, stable queries) for instance-level drift detection in CI, and a reference-free judge on stratified production samples for distribution-level generalisation. Run only reference-based on production and you hide drift on novel queries; run only reference-free on a CI set and you over-weight judge noise.
code
1Reference-based vs reference-free -- pick by cost curve, then band them23 Dimension Reference-based Reference-free (LLM judge)4 ------------------ -------------------- ---------------------------5 Build cost low (if refs exist) high (judge calibration)6 Marginal / example linear (label each) ~free (judge generalises)7 Drift sensitivity high (missing refs) low (judge generalises)8 Signal shape instance-level distribution-level9 Best for closed answer keys open / freeform / RAG10 Latency sub-ms (lexical) 1-10 s/call, parallelisable1112 Senior strategy: reference-based regression suite (CI) + reference-free13 judge on production samples. Each catches what the other hides.
Match the metric to the task, not to the popular framework. Eugene Yan’s task-specific guidance: for translation, prefer COMET/BLEURT (and chrF over BLEU where tokenisation distorts character counts); for summarisation, prefer NLI-based factual-consistency and reward models over ROUGE/BERTScore, which correlate weakly on long-form. For open-ended assistant answers and RAG, reference-free LLM judges and the Ragas family (L4) dominate. Interview angle. "What metric for summarisation quality?" → not ROUGE — a factual-consistency (NLI) check plus a binary helpfulness judge, because ROUGE rewards overlap while the failure you care about is an unsupported or inverted claim.
A teammate proposes a 1–5 quality rating for each response, averaged across two annotators, as the eval rubric. What’s the senior objection and fix?
ALikert collapses to the middle and tanks inter-annotator agreement (~0.2–0.3 kappa); replace with one-sentence binary pass/fail per dimension + 2–3 written anchorsBLikert is fine; just add a third annotator and average to reduce varianceCSwitch to a 0–10 scale for finer resolution
Your golden set is built entirely from thumbs-down production traces (all failures). A reviewer flags a risk. What’s the problem?
ANothing — failures are exactly what you want to catch, so a failure-only set is idealBIt needs both regimes: include 50–100 pass cases too, so the judge can recognise "good" and you can detect regressions on the long tail of acceptable behaviourCIt’s too small; just collect 10,000 more thumbs-down traces
You’re evaluating a closed-domain tax-calculation assistant where each query has exactly one correct numeric answer, and the question bank is stable. Which metric family fits, and why?
AReference-free LLM judge only — it generalises and needs no answer keyBReference-based (exact/contains match against the known answers) for the stable bank, since you have a trusted answer key and want instance-level regression signalCROUGE against reference explanations
You need coverage of hard input combinations (multilingual + multi-doc + ambiguous) but random sampling rarely produces them. Best approach?
AAsk a frontier LLM for "1,000 diverse, challenging questions" and trust its diversityBDefine the axes that matter and sample tuples of axis values (dimensions-and-tuples), then SME-spot-check; use weaker models to generate organic-looking failuresCJust collect more random production traffic until the combinations show up
A rubric item reads: "Pass if the answer cites a source AND is under 100 words AND avoids jargon." A response cites a source, is concise, but uses one jargon term. How should the rubric have been designed?
AKeep the compound criterion; mark it Fail since one condition failedBLower the bar so jargon doesn’t count against itCSplit into three independent binary criteria (cited / concise / jargon-free) scored separately, so each signal is interpretable and trackable
Golden-set rounds test whether you can design reference data a judge can be trusted against. Interviewers probe four things: do you sample for failure-category coverage with a pass/fail balance, do you choose binary over Likert and know why, can you name schema fields beyond Q/A, and do you treat reference-based vs reference-free as a cost-modeling decision rather than a fashion choice. Lead with the mechanism and a number.
01“Build a golden set from scratch.” → error-analysis sampling (random + stressed + active), ~200 labeled with 50–100 failures, binary rubric + anchors, dimensions-and-tuples coverage.
02“How big should it be?” → coverage over volume: stop ~20 traces with no new category; ~246/scenario for 80% pass at ±5%, 95% conf.
03“Binary or Likert?” → binary — Likert collapses to the middle and gives ~0.2–0.3 kappa; binary + 2–3 anchors gives high inter-rater reliability.
04“Why include passing cases?” → the judge must recognise "good" and you must detect regressions on currently-working behaviour; failure-only blinds you.
05“Synthetic data — safe?” → SME-spot-check it; frontier-model failures are often OOD vs real, so use weaker models for organic failures.
06“Reference-based or reference-free?” → a cost curve: reference-based is cheap/brittle/instance-level (closed keys), reference-free is costly/durable/distribution-level (open) — band them.
07“Metric for summarisation?” → not ROUGE — NLI factual-consistency + a binary helpfulness judge; overlap misses inverted/unsupported claims.
08“What’s in a golden record?” → input + context/sources + scenario tags + expected elements + label + governance — not just a Q/A pair.
Going deeper, the follow-ups that separate offers: "three annotators disagree — what’s your protocol?" (pilot round, compute kappa on an overlap set, lock written instructions and rationales that survive a new annotator joining); "how do you keep the set from going stale?" (scheduled refresh from production plus a quarterly scenario-fit review); and "how do you get ground-truth relevant docs over a 100k-doc corpus with no labels?" (synthetic generation with a strong model, then SME spot-check 15–20% to estimate the synthetic-label accuracy). Tie every answer to an artifact and a number.
Could you design a golden set + rubric for a new assistant — sampling, balance, binary criteria, schema — and defend the choices?
New to itGetting thereConfident
Takeaways
A golden set is examples + stable criteria — coverage of failure categories beats raw row count.
Sample random + stressed + active, and include both pass and fail (~200 labeled, 50–100 failures).
Binary pass/fail with 2–3 written anchors beats Likert; one criterion per item, "I don’t know" excluded from the score.
Use dimensions-and-tuples for coverage; SME-review synthetic data — frontier-model failures are often out-of-distribution.
Store a schema (context, scenario tags, expected elements, governance), not a Q/A pair.
Reference-based (cheap, brittle, instance-level) vs reference-free (costly, durable, distribution-level) is a cost choice — band them.
Next: LLM-as-judge — the dominant reference-free metric, and the position, verbosity, and self-preference biases that distort it.