Evals are the only artifact that tells a customer you know what “good” means — the five-class metric taxonomy, LLM-as-judge calibration to >80% human agreement, offline gates vs online sampling, capability-vs-regression suites, and the grader-is-production-code lesson behind Anthropic’s 42%→95% jump.
The differentiator round, not a checkbox
The round AI-native solutions engineers get that traditional ones do not is “how do you know your AI is actually working?” — and the answer is evals. Evals are the only artifact that lets you tell a customer, an auditor, or a panel that you know what “good” means and can detect when it slips. A thumbs-up button is a one-bit sentiment signal; an eval suite is a measurement instrument. This lesson is that measurement instrument, and the eval-the-AI interview round built on it.
Start with the taxonomy, because every customer-facing system composes 2–3 classes and naming them is the senior signal. Evidently’s LLM-evaluation framework breaks metrics into five families: LLM-as-a-judge (a model grades outputs against a rubric or reference), semantic/statistical (embedding similarity, NDCG, retrieval hit-rate), deterministic/exact-match (string equality, JSON-schema validation, tool-call assertions), behavioral (did the agent achieve the multi-step task end-to-end), and safety (hallucination, jailbreak, PII leakage). No single class measures a real product — you compose at least three.
code
1THE FIVE METRIC CLASSES (compose 2-3 per customer system)23 Class Measures Example4 -------------------- ----------------------- -------------------------5 Deterministic / exact did it follow the rules JSON parses? tool called?6 Semantic / statistical retrieval & similarity recall@k, NDCG, hit-rate7 LLM-as-a-judge open-ended quality rubric score, pairwise win8 Behavioral end-to-end task success did the agent close the task9 Safety harm & leakage hallucination, jailbreak, PII1011 Rule: prefer the CHEAPEST sufficient grader per task. Deterministic where you12 can; judge where you must; never reach for a judge to check JSON validity.
Anthropic’s “Demystifying evals for AI agents” gives the structural pattern that turns this taxonomy into a harness: organize tasks into evaluation suites that target specific capabilities, write graders choosing the cheapest sufficient method per task, run them inside a harness that records every step, and analyze transcripts while watching for eval saturation — the point where the agent passes every solvable task and the suite stops adding signal. The non-obvious move: split suites into capability evals (low pass rate, find what works) and regression evals (~100% pass rate, prevent backsliding), and once a capability eval saturates, promote it into the regression suite.
LLM-as-judge: calibrate it, don’t trust it on vibes
For open-ended outputs you need a model to grade, and Evidently’s LLM-as-a-judge guide prescribes the build: a few-shot prompt, chain-of-thought (have the judge explain its verdict before scoring), low temperature, and the most capable judge model available. Choose the format by what you’re grading: pairwise comparison for subjective taste (“is A better than B”), direct rubric scoring for graded quality, and reference-based scoring when a gold answer exists. Evidently reports well-built judges reach >80% agreement with human evaluators — comparable to inter-human agreement.
That >80% is an upper bound that requires work, not a free lunch. Before you trust a judge in production, calibrate it against a 30–50 case expert-labelled set and measure judge-vs-human agreement; only then is it good enough for continuous eval — and even then it is a first pass, not a go/no-go authority. The interview probe here is behavioral and pointed: “tell me about the last time your LLM-as-judge scores diverged from human review.” A strong answer describes catching the divergence on a labelled calibration set and treating the judge as a screen that escalates disagreements to humans; a weak answer treats the judge as ground truth.
python
1# Calibrate a judge before trusting it: agreement on an expert-labelled set.2def judge_human_agreement(labelled, judge):3 agree = 04 for case in labelled: # 30-50 expert-scored examples5 verdict = judge(case["input"], case["output"]) # CoT, low temp, strong model6 agree += 1 if verdict == case["human_label"] else 07 return agree / len(labelled) # want >= ~0.8, comparable to inter-human89# < 0.8 ? fix the RUBRIC (ambiguous criteria) before blaming the judge model,10# then route judge-vs-human DISAGREEMENTS to a human queue -- judge is a first pass.
The 42%→95% trap: your grader is production software
The most important eval lesson for a solutions engineer is that the grader, not the model, is often what’s wrong. Anthropic documented their own eval scores jumping from 42% to 95% after fixing grading bugs and ambiguous task specifications — the model never changed; the measurement did. The implication is sharp: every grader in your harness is production software. Put graders under version control, write tests for graders, and after every prompt change verify that known-good outputs still score highly.
The reuse pattern that falls out: ship every customer eval suite with a meta-eval — ~50 fixed cases that exercise each grader — that gates any future change to the graders themselves. This is the “eval of the eval,” and it is the silent killer when missing: if you only test the model and never test the grader, a subtly broken grader makes a broken model look fine (or a fine model look broken) and you ship the wrong conclusion. Interview angle. “Your eval pass-rate suddenly jumped 30 points after a refactor — do you celebrate?” → no; that’s the 42%→95% signature of a grader or task-spec bug, and you bisect the grader before you trust the number.
Offline and online are both required, not either/or
Two eval modes, and a customer launch needs both. Offline evals run a frozen golden dataset in CI and gate deployments — no merge if a metric regresses past threshold. Online evals sample production traffic (target >5%) and run LLM judges asynchronously to detect drift and surface new failure modes the golden set never imagined. Evidently is explicit on this split — offline CI automation plus production monitoring — and so is Anthropic’s framework, which treats drift-catching regression suites as a permanent fixture, not a one-time gate. Mirror that split in the customer’s harness; an offline-only eval goes stale the moment real users behave unexpectedly.
code
1OFFLINE vs ONLINE (a launch needs BOTH)23 Offline (gate) Online (detect)4 --------------------------------- ---------------------------------5 Frozen golden dataset in CI Sample > 5% of production traffic6 Blocks merge on regression Async LLM-judge + drift dashboard7 Deterministic + judge graders Surfaces NEW failure modes8 Answers "did we break a known case" Answers "is reality drifting"910 Microsoft: a thumbs button is one bit; instrument SPECIFIC failure modes11 ("agent refused the tool", "citation was wrong") and collect multi-bit feedback.
Microsoft’s “beyond thumbs-up and thumbs-down” framing is the operator’s warning against the most common eval shortcut: a thumbs button is a single-bit sentiment signal that correlates poorly with task success. The remediation is to instrument specific failure modes — “the agent refused to use the customer’s tool,” “the citation pointed at the wrong document,” “it answered from stale data” — and collect multi-bit feedback on each, which then feeds the online eval and the next golden-set additions. This is the bridge to Lesson 3’s feedback loops.
The launch checklist a solutions engineer brings to a customer: (i) ship a 50-case golden dataset before any customer code; (ii) add a deterministic pass that gates merge; (iii) add an LLM-as-judge harness for open-ended outputs, calibrated to ≥80% human agreement; (iv) add an online pipeline sampling >5% of traffic; (v) wire a feedback-capture path so user signals feed the analysis layer; (vi) write a meta-eval that detects grader drift. Saying this list out loud in a discovery call is what de-risks the deal in the buyer’s eyes.
The SE difference is in how you present the eval plan, not just whether you can build one. The eval suite is the artifact that converts the buyer’s biggest unspoken objection — “how do I know this won’t hallucinate on my customers and how do I prove it to my risk committee?” — into a controlled, demonstrable answer, so you lead with it as a de-risking instrument, not an engineering deliverable. Three moves make it sell: (1) co-author the golden set with the champion from their real, hard cases, so the acceptance bar is theirs and the eval doubles as the success criteria in the SOW; (2) show the failing cases, not just the passing ones — a vendor who surfaces where the system breaks and how the gate catches it earns more trust than one who demos only green; (3) hand the champion the pass-rate trend line and the meta-eval story as the evidence they carry into the steering committee. Interview angle. “How does an eval plan help you win the deal, not just ship quality?” → it is the proof-of-control a regulated buyer needs to say yes; you present it as the acceptance contract, co-built from their cases, with the failure modes shown openly — the eval is a selling artifact, the same one Lesson 3’s trend line renews on.
Behavioral evals: grade the trajectory, not just the answer
The hardest customer-facing AI is an agent, and an agent can produce the right final answer through a broken path — or the wrong answer because one step failed — so the behavioral metric class grades the trajectory, not only the output. Two questions every agent eval must answer separately: did it achieve the end-to-end task? (task success) and did it take a sane path to get there? (no redundant tool calls, no unnecessary escalation, no looping). A purely outcome-based eval will pass an agent that stumbled into the right answer after five wasted tool calls — and that agent is a latency and cost incident waiting to happen, even though its accuracy looks perfect.
Concretely, you grade an agent on three layers: final-answer correctness (deterministic or judge), tool-call correctness (did it call the right tool with the right args — deterministic assertions on the trace), and trajectory quality (step count, recovery from a failed call, whether it asked for human approval when it should have). This is where evals and observability (Lesson 3) fuse: the trajectory you grade offline is the same span tree you trace in production. Interview angle. “How do you eval an agent, not just an LLM call?” → score task success AND the trajectory (tool-call assertions + step efficiency + recovery), because an agent that gets the right answer the wrong way is a production risk your outcome metric will hide.
A customer asks, “How will we know the AI is actually working in production?” You have 90 seconds. Strongest answer?
AA golden dataset gated in CI plus online sampling of >5% of traffic with a calibrated LLM-judge and a drift dashboard; deterministic graders where possible, judge where neededBWe collect thumbs-up / thumbs-down from users and review the negativesCWe benchmark the model on MMLU and HumanEval before launch
After a refactor, your eval pass-rate jumps from 61% to 92% with no change to the model or prompt. What’s the right reaction?
ACelebrate and ship — the system clearly got betterBBisect the grader and task spec first — a pass-rate that jumps without a model change is a grader bug until proven otherwiseCLower the threshold so the result is more conservative
You need to grade whether the agent returned valid JSON with the right fields, and separately whether its written summary is high-quality. Which grading split is correct?
AUse the LLM-as-judge for both, since it understands the task bestBDeterministic grader (schema + field assertions) for the JSON; calibrated LLM-as-judge with a rubric for the summary qualityCUse exact string match for both against a reference answer
Your offline golden-set eval has passed 100% for weeks, but users report new wrong answers. What’s the gap and the fix?
AThe model degraded; swap to a newer modelBRaise the sampling temperature so the eval is harderCThe golden set saturated and went stale; add online evals that sample production traffic to surface new failure modes, then fold them back into the golden set
A capability eval you wrote to probe a hard multi-step task has climbed to a ~100% pass rate. What should you do with it?
ADelete it — it no longer fails, so it’s providing no valueBPromote it into the regression suite so it permanently guards against backsliding, and write a harder capability eval to find the new frontierCLower its difficulty so it produces a spread of scores again
The eval-the-AI round is the differentiator between an AI solutions engineer and a traditional one. It probes whether you can name an eval set (golden + adversarial), pair automatic metrics with human spot-checks, distinguish online from offline, calibrate a judge, and detect regressions in latency, cost, and quality. Strong answers name metrics and the cheapest-sufficient grader; weak answers say “we have an eval suite” without naming a single metric, or treat LLM-as-judge as ground truth.
01“How do you know your AI is actually working?” → golden set in CI + >5% online sampling + calibrated judge as a first pass + human review on disagreements + drift dashboard.
02“Name your eval metric classes.” → deterministic/exact, semantic/statistical, LLM-as-judge, behavioral (end-to-end), safety — compose 2–3 per system.
03“How do you trust an LLM judge?” → calibrate to ≥80% human agreement on a 30–50 case labelled set, CoT + low temp + strong model, escalate disagreements.
04“When did a judge diverge from humans?” → caught it on the calibration set; treated the judge as a screen, fixed the rubric (ambiguity), never as ground truth.
05“Your pass-rate jumped 30 points after a refactor — celebrate?” → no; 42%→95% grader-bug signature; bisect the grader and task spec first.
06“Offline vs online evals?” → offline gates merges on a frozen set; online samples production to catch drift and new failure modes; you need both.
07“What’s the eval-of-the-eval?” → a meta-eval of ~50 fixed cases that gates changes to the graders, because the grader is production code.
08“A capability eval hit 100% — now what?” → promote it to the regression suite and write a harder capability eval to find the new frontier.
Push it deeper with the follow-ups that separate “read a blog” from “shipped one”: “how do you know your golden set represents production?” (harvest queries and gold contexts from the production retriever’s real runs, not hand-picked chunks, or offline scores read far higher than reality); “what’s your cost-per-request budget at 10× traffic, and does the eval track it?” (evals measure cost and latency regressions, not just quality); and “what’s the longest a hallucinated answer stayed in production before you caught it?” (a detection-latency question that only the online-eval + observability stack answers). Each maps to one of the seven architecture-rubric axes.
Could you stand up a customer eval harness end to end — taxonomy, calibrated judge, offline gate, online sampling, meta-eval — and defend it in the eval round?
New to itGetting thereConfident
Takeaways
Evals are the differentiator — “how do you know it works?” is answered with a stack, not a thumbs button.
Compose 2–3 of the five metric classes; use the cheapest sufficient grader per task.
Calibrate the LLM judge to ≥80% human agreement; use it as a first pass, escalate disagreements.
The grader is production code — a pass-rate that jumps without a model change is a grader bug.
You need both offline gates (CI golden set) and online detection (>5% sampling + drift dashboard).
Saturated capability evals graduate into the regression suite; keep pushing the frontier.
Next: once it’s live, how do you see what it’s doing — observability, span-level telemetry, and feedback loops that close the system.