Lesson 6 of 6 · 46 min

Capstone: eval pipeline for a chatbot

Assemble the whole discipline into one trustworthy eval pipeline for a support chatbot — error analysis → golden set → calibrated judge → regression gate → production monitor — and rehearse proving to a skeptical PM that the metric reflects what users actually care about.

Capstone

Build a trustworthy eval pipeline

Time to assemble it: a complete evaluation pipeline for a customer-support chatbot that deflects tickets — a golden set built from real traces, a judge calibrated against human labels, a regression suite that gates CI, a monitor on live traffic, and a feedback loop that keeps it all honest. This is the canonical GenAI-eval interview ("evaluate this chatbot before and after launch, and convince a skeptical PM your metric is real") and a real deliverable. Everything below is a callback to a specific lesson.
Treat this as the interview itself. The round opens with "we’re launching a support chatbot — how do you evaluate it, what do you monitor, and how do you prove the metric reflects what users care about?" The gap between a mid and a senior signal is structure and the refusal to start with metrics: you start with error analysis on real traces, define binary criteria per failure mode, calibrate the judge against humans, gate the deploy, monitor production, and loop failures back. The single best opening move is to name the chain of trust before drawing any pipeline — it signals you know an eval is a practice, not a tool you bolt on.
  1. 01Scope? — what does "good" mean for this bot (resolved, faithful, on-policy, polite)? Define before measuring.
  2. 02Failure modes? — open-code ~50–100 real traces into a taxonomy before any metric.
  3. 03Ground truth? — golden set with both pass and fail regimes, stratified by scenario tags.
  4. 04Judge? — pointwise per dimension, calibrated to human labels (TPR/TNR + Cohen’s kappa).
  5. 05Trust? — can you show the judge tracks a business signal (re-open, CSAT, escalation)?

The reference pipeline

End to end: 1) error analysis — open/axial-code real traces into a 5–8 bucket failure taxonomy → 2) golden set — hand-label ~200 examples (50–100 failures), stratified, with a schema and one binary rubric per bucket → 3) judge — build one pointwise judge per dimension, calibrate TPR/TNR and Cohen’s kappa against the human labels, with bias mitigations → 4) regression gate — deterministic assertions + golden-set comparison + judge, blocking the build on regression → 5) production monitor — assertions on 100%, judge on a 1–5% sample with CIs, drift detection, weekly human spot-review → 6) loop — promote production failures back into the golden set.
code
1Trustworthy chatbot eval pipeline -- the lesson behind each stage23  BUILD (offline, once + iterated)4    error analysis: open/axial-code ~50-100 traces ......... L1 (look at data first)5    golden set: ~200 labeled, 50-100 fails, stratified ..... L2 (both regimes + schema)6    binary rubric, one criterion per dimension ............. L2 (anchors, not Likert)7    judge: pointwise/dimension, swap+cross-family .......... L3 (bias mitigations)8    calibrate: TPR/TNR + Cohen's kappa vs human labels ..... L3/L4 (trust the judge)9  GATE (every change / deploy)10    deterministic assertions (regex/schema/PII) ............ L5 (Level-1 floor)11    golden-set regression + judge on subjective tiers ...... L5 (layered gate)12    block build on regression; gate model upgrades ......... L5 (replay-set delta)13  OPERATE (live)14    assertions 100% + reference-free judge on 1-5% sample .. L5 (CI-band alerts)15    drift monitors + 10-20 weekly outlier review ........... L5 (four drift vectors)16       |  promote production failures -> golden set (the loop)1718  Build sets the bar; gate stops regressions; operate catches drift; the loop keeps it living.
Narrate it as three concerns, not one pipeline: the build path defines what "good" means and proves the judge matches humans; the gate stops a known regression from shipping; the operate path catches the unknown drift on live traffic; and the loop ties operate back to build. Candidates who only describe "run a judge on outputs" miss the calibration that makes the judge trustworthy and the loop that keeps it current. Interview angle. Open with these swim-lanes — build / gate / operate / loop — and the interviewer immediately sees you treat eval as a system with separate concerns, not a single script that prints a number.
python
1# The calibrated judge at the heart of the pipeline -- one dimension shown.2def evaluate_dimension(traces, judge, gold, dimension):3    # 1) score production traces with a single-purpose pointwise judge4    preds = [judge_pointwise(t.input, t.response, dimension, judge) for t in traces]5    # 2) you may ONLY trust preds because the judge was calibrated on gold:6    tpr, tnr, kappa = calibrate(judge, gold, dimension)   # vs human labels, held-out7    assert kappa >= 0.6, "judge not aligned enough to ship on this dimension"8    rate = sum(preds) / len(preds)9    lo, hi = wilson_ci(sum(preds), len(preds))            # report a band, not a point10    return {"dimension": dimension, "pass_rate": rate, "ci": (lo, hi),11            "judge_tpr": tpr, "judge_tnr": tnr, "kappa": kappa}12# Trust chain: humans agree on the rubric -> judge matches humans (kappa) ->13# judge tracks a business signal (re-open rate). Break any link, the number is theater.

The chain of trust: proving it to a skeptical PM

The hardest part of the interview is the PM question: "how do you know your pipeline measures what we care about?" Frame trustworthiness as a three-link chain, each with its own evidence. (1) The golden labels are right — show inter-annotator agreement (Cohen’s kappa or Krippendorff’s alpha) on a ~100-case pilot, so two humans agree on what "good" means. (2) The judge matches humans — show TPR/TNR on a held-out labeled set, calibrated to ~75–80% agreement (kappa in the substantial-to-excellent band) with a separate failure analysis on the false positives and negatives. (3) The judge tracks the user signal — side-by-side sampling on production traffic, computing the judge score and a behavioural metric (re-open, escalation, CSAT) and reporting their correlation. Break any link and the number is theater.
Two pitfalls to pre-empt out loud. Percent agreement is not the proof — "we have 90% agreement" is uninformative on imbalanced labels; lead with Cohen’s kappa, and flag the high-rho/low-kappa pattern as majority-class defaulting. And criteria drift — Shreya Shankar’s documented failure where annotators refine their standards after seeing outputs — means you must lock the rubric before reviewing any outputs, or the bar moves under you. Interview angle. The differentiator is acknowledging a false-failure-rate budget as a tunable knob (typically 5–10% false positives on safety-critical dimensions, looser on stylistic ones) — it shows you treat the judge as a probabilistic sensor with a cost function, not an oracle.

What breaks — the failure modes to pre-empt

A senior design names how the eval itself fails before the interviewer asks. The five that recur: (1) judge biases — position/verbosity/self-preference quietly skew the metric (defense: swap-and-resolve, length-budget, cross-family judge — L3); (2) percent-agreement theater — a green agreement number hiding near-zero kappa on imbalanced labels (defense: report kappa, re-stratify — L4); (3) golden-set staleness — the offline set drifts from production and the gate passes a degrading product (defense: scheduled refresh + the online loop — L2/L5); (4) judge drift — the human-aligned judge and real quality decouple over time (defense: rolling-window re-calibration, monitor kappa — L3/L5); and (5) cost blow-up — a full-coverage judge ensemble hitting the $1M/month tier (defense: sampling rate as a knob, cascade judges — L5).

The decisions an interviewer will push on

  1. 01Where do you start? → error analysis on ~50–100 real traces, not infrastructure or vendor labels (L1).
  2. 02Golden set design? → ~200 labeled, 50–100 failures, stratified, schema + binary rubric per dimension (L2).
  3. 03Judge — pointwise or pairwise? → pointwise for objective dims (resolved, faithful, on-policy), pairwise for subjective comparisons (L3).
  4. 04How do you trust the judge? → TPR/TNR + Cohen’s kappa vs human labels, bias mitigations baked into the eval library (L3/L4).
  5. 05RAG sub-metrics? → faithfulness + answer relevance + context precision/recall, each an uncalibrated judge (L4).
  6. 06CI gate? → deterministic-first, judge last, block on regression, gate model upgrades with a replay set (L5).
  7. 07Production drift? → four vectors, alert outside a baseline CI, weekly outlier review, promote failures back offline (L5).
  8. 08Eval cost? → sampling rate is architectural; cascade judges, cache deterministic sub-scores ($1M/month is real) (L5).
Going deeper, the curveballs that separate offers: "the judge disagrees with humans on 12% of cases" (error-analyse the FPs/FNs, fold disagreements in as few-shot anchors, re-validate — not "retrain the judge"); "you have budget for exactly one production judge — which dimension?" (faithfulness/safety, because those failures are reputational while relevance degrades into CSAT); "offline passes but users complain" (the offline-online gap — score live traffic and promote the new failures offline); and "prove a 2-point score move is real" (it’s inside the CI at n=200 — show the band before claiming a regression). Tie every answer to a number, an artifact, and the failure it prevents.

Rollout: ship the eval without fooling yourself

The last mile is operating the eval itself. Four practices the interviewer will respect: (1) version everything — pin the model, the prompt, the judge prompt, and the golden-set version, and log every score with those ids so a regression is attributable; (2) gate in CI — run the layered suite on every change, blocking a deploy that regresses past a margin; (3) re-calibrate on a cadence — re-measure judge-vs-human kappa on a rolling window so judge drift can’t silently erode the metric; and (4) close the loop — every production failure becomes a golden-set case and a CI assertion. The throughline of the whole track: an eval is a living practice — calibrated, gated, monitored, and fed by its own failures — not a number you compute once and trust forever.
A trustworthy eval isn’t the pipeline that prints a good score once — it’s the one where humans agree on the rubric, the judge is calibrated against them, the gate blocks regressions, the monitor catches drift, and every escaped failure becomes a permanent test. The number is only as honest as the chain behind it.
videoLLM Evals: Common Mistakes (the senior failure modes)Hamel HusainarticleA Field Guide to Rapidly Improving AI ProductsHamel Husainrepoopenai/evals — framework + the "detecting prompt regressions" recipeOpenAIreporagas — reference-free RAG evaluation libraryexplodinggradients/ragas

Checkpoint

A PM asks "how do you know your eval pipeline measures what users care about?" What’s the strongest structure for your answer?

AWalk the three-link chain of trust: humans agree on the rubric (kappa on a pilot) → the judge matches humans (TPR/TNR on a holdout) → the judge correlates with a business signal (re-open/CSAT)BShow the judge’s pass rate is high and stable over the last monthCCite that you use Ragas and an LLM judge, which are industry standard
Sign up free to answer and see why

Checkpoint

You’re standing up the pipeline from scratch in 30 days. What’s the correct order of operations?

AStand up the dashboard and judge first, then collect labels later if there’s timeBA/B test prompt variants immediately to find the best one, then build evals around the winnerCError-analyse real traces → build a stratified golden set with binary rubrics → calibrate one judge per dimension (TPR/TNR, kappa) → wire CI gates → ship the monitor + feedback loop
Sign up free to answer and see why

Checkpoint

For the support bot you can afford exactly one calibrated judge dimension to run on production traffic. Which do you pick, and why?

AConciseness — shorter answers are cheaper to serve and readBFaithfulness / on-policy — confidently wrong or off-policy answers are the reputational, high-risk failures, while relevance degrades gracefully into CSAT/re-open rateCPoliteness — it’s what users notice first
Sign up free to answer and see why

Checkpoint

Three months after launch, the offline gate is green but support escalations are creeping up. What does a trustworthy pipeline already have in place?

AOnline scoring on sampled live traffic with CI-band alerts, drift monitors, a weekly outlier review, and a loop that promotes the new failures into the golden setBA bigger judge model swapped in last week to improve accuracyCAuto-scaling to handle the increased escalation volume
Sign up free to answer and see why

Checkpoint

A stakeholder points out your judge agrees with humans 90% of the time and asks why you keep insisting on more validation. Best response?

A90% percent agreement can hide a near-zero Cohen’s kappa on imbalanced labels — the judge may barely beat chance on the rare failure class; I report kappa and a false-failure budget, and re-stratify to confirm real skillBYou’re right — 90% is plenty, I’ll stop validating and shipC90% is too low; I’ll keep iterating the prompt until it hits 99%
Sign up free to answer and see why

Interview prep

The capstone is the interview. A GenAI-eval round tests whether you start with error analysis, build reference data a judge can be trusted against, calibrate the judge with the right statistic, gate and monitor it, and prove the chain of trust to a non-technical stakeholder. The structure that wins: scope "good" → error-analyse → golden set → calibrated judge → gate → monitor → loop — and lead with the chain of trust.
  1. 01“Evaluate this chatbot before and after launch.” → error analysis → golden set (both regimes) → calibrated judge → CI gate → sampled online monitor → loop.
  2. 02“Prove the metric reflects user value.” → three-link chain: humans agree (kappa) → judge matches humans (TPR/TNR) → judge tracks re-open/CSAT.
  3. 03“Where do you start with 30 days?” → open-code real traces, not LangSmith-first or vendor labels.
  4. 04“Pointwise or pairwise for the bot?” → pointwise per objective dimension (resolved/faithful/on-policy), pairwise for subjective comparisons.
  5. 05“One production judge — which dimension?” → faithfulness/safety; reputational failures over gracefully-degrading relevance.
  6. 06“Offline green but users complain — why?” → the offline-online gap; score live traffic and promote failures back offline.
  7. 07“Is a 2-point drop a regression?” → not at n=200 (inside the CI ~±2.4% at 3% defect); show the band first.
  8. 08“90% agreement — enough?” → no — report Cohen’s kappa; imbalance inflates percent agreement, kappa can be near zero.
A final synthesis worth carrying into the room: every practitioner lineage — Hamel’s "look at your data," Eugene’s "label → align → run," Shreya’s "validate the validators" — converges on one principle: every metric is a model of an ideal reader, and every model must be tested. The human owns the rubric; the machine scales it; the discipline is keeping those roles separate and proving, link by link, that the number means what you claim. Teams that ship metric numbers without that validation are protecting themselves from noticing the failure, not from the failure itself.

Could you build this whole eval pipeline — and prove its trustworthiness to a skeptical PM — in a 45-minute interview?

Not yetMostlyYes

You can now

  • Assemble an eval pipeline: error analysis → golden set → calibrated judge → CI gate → production monitor → loop.
  • Calibrate the judge as a sensor — TPR/TNR and Cohen’s kappa against human labels, with bias mitigations.
  • Prove the chain of trust: humans agree on the rubric → the judge matches humans → the judge tracks a business signal.
  • Gate deploys deterministic-first, monitor four drift vectors with CIs, and promote production failures back offline.
  • Treat eval cost and sampling rate as architecture, and defend a 2-point move as noise inside the confidence interval.

You’ve completed GenAI Evaluation for Data Scientists — go measure a non-deterministic system you can actually trust.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.