Capability vs regression evals, grading the trajectory not just the answer, pass^k over pass@k, LLM-as-judge done right, and how teams build eval sets from real failures — the discipline that turns an agent demo into a shippable, non-regressing product, plus the interview round that probes all of it.
You can’t change what you can’t measure
A probabilistic, multi-step agent changes behaviour when you touch a prompt, a tool, or when the provider silently updates the model. Without evals you can’t tell “it got better” from “it quietly got worse at something it used to do” — and grading only the final answer hides the multi-step bugs entirely. Evals are not a QA afterthought for agents; they are the development loop itself. The teams that ship reliable agents are the ones who treat the eval set as the most valuable artifact they own — more valuable than the prompt, because the prompt is downstream of what the evals tell you.
Anthropic’s framing splits evals into two roles. Capability evals ask “what can it do?” and should start low — a mountain to climb (Claude went 40% → 80%+ on SWE-bench Verified; Opus 4.5 went 42% → 95% on CORE-Bench after a bug fix). Regression evals ask “does it still handle everything it used to?” and should sit near 100%, so any change that breaks existing behaviour fails CI. The lifecycle: ship a small capability eval (20–50 tasks drawn from real failures); once it saturates, graduate it into the regression suite. Saturation is a signal, not success.
Where eval sets come from: error analysis, not vibes
The hardest part of evals is not the harness; it’s knowing what to measure. Hamel Husain’s widely-taught method is error analysis: take 50-100 real (or realistic) traces, read them, write a free-text note on what went wrong in each, then cluster those notes into failure modes (axial coding). The clusters become your eval categories, and their frequencies tell you what to fix first. This is the antidote to the two most common eval mistakes: inventing metrics in the abstract (which measure things that never fail) and adding a generic “helpfulness 1-5” judge (which is too coarse to drive any decision). Interview angle. “How would you start evaluating an agent that has none?” → “Read 50 real traces, label the failure modes, build a small eval per dominant mode” — not “write an LLM-judge for overall quality.”
code
1ERROR-ANALYSIS -> EVAL pipeline (Hamel Husain's method)23 1. collect 50-100 real/realistic traces4 2. OPEN-CODE: one free-text note per trace -- what went wrong, specifically5 3. AXIAL-CODE: cluster notes into failure modes6 e.g. "skipped the auth check" 42%7 "wrong tool for date math" 18%8 "looped on a 404 (treated as 503)" 15%9 "hallucinated an order id" 12%10 4. build a small, targeted eval per dominant mode (start with the 42%)11 5. fixes lower that mode's frequency -> RE-READ traces -> new modes surface1213 measure what actually breaks; don't invent metrics that never fail.
Trajectory eval beats output eval (for agents)
Grade the multi-step transcript, not just the final answer — it catches “wrong but lucky” and “right but wildly inefficient.” Notion went from 3 to 30 fixes/day by switching to trajectory grading: was the right tool selected, with the right arguments, in the right order? Braintrust’s step efficiency makes it concrete — 7 tool calls where 3 would do is 43%, a regression an output-only eval marks “success.” The nuance (Anthropic): also read transcripts, because some “failures” are valid solutions the grader didn’t anticipate — so layer deterministic checks (tool/argument correctness) with model-based rubrics (plan quality), and don’t over-punish creative-but-correct paths.
There are two ways to grade a trajectory, and mature suites use both. Reference-based (final-state) checks compare the end state to a golden answer — did the file get the right contents, did the ticket reach the right status — and are cheap and unambiguous, but they miss how the agent got there. Reference-free (process) checks grade the path itself — tool choice, argument correctness, ordering, step efficiency, whether a required safety check ran — and are what catch the wrong-but-lucky and right-but-wasteful runs. The trap is grading only the final state: an output-only eval blesses an agent that took 7 redundant tool calls, skipped a permission check, and got lucky. Interview angle. If asked “why isn’t final-answer accuracy enough for an agent?”, the answer is: it can’t see efficiency, can’t see a skipped safety step, and can’t distinguish a robust solution from a lucky one — all of which are the difference between a demo and production.
pass^k, not pass@k
The metric you pick encodes your reliability bar. pass@k (“at least one of k attempts worked”) flatters an agent; pass^k (“all k attempts worked”) is the production bar — “a support agent that fails every third request isn’t production-ready.” A useful diagnostic: a 0% pass@100 almost always means a broken task, not an incapable agent.
code
1Reliability metrics23 pass@1 single-shot success cheap smoke test; hides variance4 pass@k >=1 of k attempts succeed best-case ceiling; dev-time screening5 pass^k ALL k attempts succeed production reliability / SLA bar67 why the gap is huge: if one attempt is 90% reliable,8 pass@5 = 1 - 0.10^5 = 99.999% (looks amazing)9 pass^5 = 0.90^5 = 59% (the user-facing truth on 5 turns)1011 Ship customer-facing agents on pass^k. pass@k makes a coin-flip look reliable.
The reason pass^k is the honest metric is the same multiplicative compounding from the patterns lesson, applied to independent attempts instead of sequential steps: a per-attempt success of 90% gives a pass^5 of just 59%. A real user doesn’t get five tries — they get one, and then another, and another, and they experience the conjunction. The diagnostic value of 0% pass@100 is worth memorising for an interview: if an agent fails all 100 attempts, it is almost never that the model is incapable — it is a broken task (impossible grader, missing tool, wrong setup) or, for a security eval, the model correctly refusing a malicious request. A 0% that’s actually a correct refusal is a grader bug masquerading as a capability gap.
LLM-as-judge: powerful, biased, calibrate it
LLM-as-judge scales eval, but it’s a noisy classifier with documented biases: position (robustness drops below 0.5 with 3–4 options), self-preference (self-enhancement error 16.1 for one model, 8.91 for another), verbosity, and authority bias. The mitigation playbook: randomise candidate order and average; never use the same model family to generate and judge; give the judge an explicit “Unknown” option so it doesn’t hallucinate a score; and calibrate against a small human-labelled set (track Cohen’s κ). Beware rubric-induced preference drift — a rubric edit that passes validation but shifts judgments elsewhere.
The single biggest upgrade to an LLM judge is to stop asking it for a number. A “rate this 1-5” judge is unreliable and uncalibrated; a binary, criterion-specific judge with a written pass/fail definition and a required justification is dramatically more stable, because you have decomposed a fuzzy quality score into concrete yes/no questions that map to your error-analysis failure modes. Then calibrate: label a few dozen examples yourself, measure agreement (Cohen’s κ; aim for ≥0.6-0.8 before you trust the judge), and re-calibrate whenever you edit the rubric. Interview angle. “How do you make an LLM judge trustworthy?” → binary per-criterion prompts with justifications, cross-family (judge ≠ generator), randomised order, an “Unknown” escape hatch, and human-calibrated agreement you actually measure.
python
1# A judge you can trust: binary, criterion-specific, justified, calibrated.2JUDGE_PROMPT = (3 "You are grading whether the agent COMPLETED THE REFUND CORRECTLY.\n"4 "PASS only if ALL hold: (1) it verified the order exists, "5 "(2) it checked the refund window, (3) the refund amount matches the order.\n"6 "If you cannot tell from the transcript, answer UNKNOWN -- do not guess.\n"7 "Answer JSON: {verdict: PASS|FAIL|UNKNOWN, reason: }"8)9# calibrate against human labels before trusting it:10# kappa = cohen_kappa(judge_labels, human_labels) # want >= 0.6-0.811# if kappa < 0.6: the rubric is ambiguous -> tighten the criteria, re-label.12# never let the same model family generate AND judge (self-preference bias).
Eval-gated CI: the safeguard against silent regression
The payoff of all this is a gate: every prompt edit, tool change, model upgrade, and provider checkpoint runs the regression suite in CI, and a drop below the bar blocks the change. This is the only defense against the failure mode the foundations track named — a provider silently updates the model and quality sags weeks later, with nothing in your git log to blame. Pin the model and prompt versions, run the eval on every change, and the regression eval turns “users told us it got dumber” into “CI told us, before we shipped.” Notion’s 3→30 fixes/day jump was not just better grading; it was a faster loop — trajectory grading told them precisely which step broke, so a fix took minutes instead of a debugging session.
Interview prep
Eval interviews are where AI-engineer roles separate “I prompted an agent” from “I shipped one.” Interviewers probe whether you can start an eval program from nothing (error analysis), pick the right metric for the bar (pass^k vs pass@k), grade the process not just the answer (trajectory eval), and make an LLM judge trustworthy (binary, cross-family, calibrated). Lead with the failure modes you’d measure, then the harness — never the other way around.
01“How do you start evaluating an agent with no evals?” → Error analysis: read 50-100 traces, open-code what went wrong, cluster into failure modes, build a small targeted eval per dominant mode — not a generic quality judge.
02“Capability vs regression evals?” → Capability starts low and you climb it (the mountain); regression sits near 100% and gates CI (the guardrail). Graduate saturated capability evals into the regression suite.
03“Why grade the trajectory, not just the final answer?” → Output-only blesses wrong-but-lucky, right-but-wasteful, and skipped-safety-step runs; process grading (tool choice, args, order, step efficiency) catches them. Notion: 3→30 fixes/day.
04“pass@k or pass^k for a customer-facing agent?” → pass^k — the user experiences the conjunction of attempts (90% per-attempt → 59% pass^5); pass@k makes a coin flip look reliable.
05“What does a 0% pass@100 usually mean?” → A broken task (impossible grader / bad setup) or a correct refusal on a security eval — almost never an incapable agent.
06“How do you make an LLM judge trustworthy?” → Binary criterion-specific prompts with justifications, judge ≠ generator family, randomised order, an Unknown option, and human-calibrated agreement (Cohen’s κ ≥ 0.6-0.8).
07“How do evals prevent silent regressions?” → Pin model+prompt versions and gate every change on the regression suite in CI; a provider checkpoint that degrades quality fails the gate before users notice.
08“What’s the most valuable artifact in an agent project?” → The eval set built from real failures — the prompt and architecture are downstream of what it tells you.
Going deeper. Volunteer the nuances that show real experience: reference-based (final-state) vs reference-free (process) checks and why you need both; rubric-induced preference drift (editing a rubric can silently shift judgments elsewhere, so re-calibrate after edits); self-preference bias as the concrete reason judge and generator must differ (self-enhancement error 16.1 vs 8.91 for two models); and that an eval suite is itself a maintained dataset that rots — you re-run error analysis as new failure modes emerge. If asked about cost, note you run cheap deterministic checks on every request and sample the expensive judge.
You inherit an agent with no evals and a vague “make it better” mandate. What’s the right first move?
ARead 50-100 real traces, open-code what went wrong in each, cluster into failure modes, and build a small targeted eval for the most frequent modeBAdd an LLM-as-judge that rates overall quality 1-5 on every responseCUpgrade to the newest model and re-test by hand
A security eval for prompt injection reports 0% pass@100 on a task where the agent should NOT comply with an injected instruction. Most likely explanation?
AThe agent is completely incapable and needs a bigger modelBThe provider silently downgraded the modelCThe grader is scoring a correct refusal as a failure — the agent is doing the right thing and the eval is inverted/broken
Your LLM judge disagrees with your spot-checks often, and you used the same model family to generate and to judge. Best fix?
ASwitch the judge to a 1-5 quality scale for more granularityBUse a different model family to judge, rewrite the rubric as binary criterion-specific pass/fail with justifications, randomise order, and calibrate against human labels (Cohen’s κ)CRaise the judge’s temperature so it considers more options
Could you start an eval program from error analysis, stand up capability + regression evals, grade trajectories, calibrate an LLM judge, and defend it all in an interview?
Not yetMostlyConfident
Takeaways
Build eval sets from error analysis on real traces (open-code → cluster failure modes), not invented metrics.
Capability evals start low and climb; regression evals sit near 100% and gate CI.
Grade the trajectory (tool choice, args, order, step efficiency, skipped safety steps), not just the final answer.
Report pass^k for customer-facing reliability (90% per attempt → 59% pass^5); 0% pass@100 usually means a broken task or a correct refusal.
LLM-as-judge is biased — binary criterion-specific prompts, cross-family, randomise order, give it “Unknown,” calibrate vs humans (Cohen’s κ).
Eval-gated CI on pinned versions is the only defense against a silent provider regression.
Next: observability — span-level traces so you can debug “works in staging, fails in prod.”