Lesson 5 of 6 · 48 min

Regression suites & monitoring

Wire evals into CI as deployment gates, monitor four drift vectors in production with confidence intervals, and treat eval cost as an architecture decision — because at scale the eval bill can exceed $1M/month.

Evals as CI, not a notebook

A golden set and a calibrated judge are inert until they gate a deploy and watch production. This lesson operationalises everything prior: turn the eval into a regression suite that blocks a bad release in CI, a monitor that catches drift on live traffic with confidence intervals, and a cost model that keeps the whole thing affordable — because for some teams the LLM-eval bill exceeds $1M/month, which makes eval an architecture decision, not an afterthought.
An LLM regression suite is not a classic unit-test suite, and the difference is the whole reason teams get it wrong. Classic unit tests assert exact outputs; LLM outputs have many acceptable forms, so exact-match tests are unreliable and you need a layered comparator. Three documented antipatterns from the Evidently regression tutorial: (i) exact-match unit tests are ineffective because multiple answers are acceptable; (ii) an LLM-judge in CI introduces speed, cost, and external-dependency risk; (iii) standard accuracy metrics mislead on imbalanced classes, so recall on the risky class is often the better gate. The fix is a layering — deterministic checks where you can, judges only where you must.
LLM & RAG Evaluation Playbook for Production AppsPaul Iusztin / ODSC

The CI gate: layered, deterministic-first

A defensive CI configuration layers the suite so a single config change can’t invalidate every test at once (Latitude’s unit/contract/regression tiering). (1) Deterministic assertions run on 100% of cases — regex, exact/contains match, JSON-schema validation, must-contain disclaimer, must-not-leak PII — cheap and reliable, the Level-1 floor. (2) A reference-based regression set compares new outputs against a frozen golden baseline with semantic similarity. (3) An LLM judge only for the subjective tiers, where nothing deterministic suffices. The gate fails the build on regression — Braintrust ships exactly this as a native GitHub Action that declines the merge on a threshold breach. The principle: push as much as possible to deterministic checks, and spend the expensive, slower, externally-dependent judge calls only where they earn their keep.
code
1Layered CI eval gate (cheap+deterministic first, judge last)23  Layer              Coverage   Comparator                 Gate4  ----------------   --------   ------------------------   ----------------------5  Unit assertions    100%       regex / exact / schema     block on ANY flip6  Golden regression  ~100-200   semantic sim vs frozen     block on score drop > m7  LLM judge (subj.)  sample     calibrated pointwise       block outside CI band8  Latency / cost     100%       p99 + $/req budget         block on budget breach910  Sizing a gate: ~246 cases / behaviour -> 80% pass at +/-5%, 95% conf.11  Every production incident becomes a NEW permanent assertion (eval-driven dev).
The cultural practice underneath the tooling is eval-driven development: every production failure is converted into a permanent test case, so a bug fixed once can never silently reappear. This is the LLM analogue of regression-test-on-every-bug, and it’s what makes the golden set a living artifact rather than a frozen snapshot. One concrete deadline on OpenAI’s public deprecations page (announced June 2026): the Evals platform becomes read-only on 2026-10-31 and shuts down 2026-11-30 — a hard migration date for any team anchored to it, and a reminder to treat any single vendor as a complement with an integration plan, not a permanent foundation. Interview angle. "How do you stop a fixed bug from regressing?" → turn the incident into a golden-set case and a CI assertion; "how do you keep the suite from sprawling?" → one assertion per concern, single owner, retire obsolete ones in a weekly review.

Production drift: four vectors, one feedback loop

Offline gates catch known regressions; drift is the unknown that escapes them. Production distributions are non-stationary, and the research names four-plus vectors: prompt drift (inputs deviate from what you tuned for), model drift (responses change — including a silent provider update), user drift (new topics and intents), and judge drift (your human-aligned judge and real quality decouple). Each silently moves the eval score. The monitoring stack composes four moving parts: cheap behavioural assertions on 100% of traffic (a disclaimer that should fire 99% of the time dropping to 95% over 24h is a regression); a reference-free LLM judge on a 1–5% sample scored daily with confidence intervals; distribution-shift monitors (embed inputs/outputs, alert on topic-mix or KL drift) to catch novel failure types no assertion was written for; and a human spot-review cadence (10–20 weekly outliers) that writes new assertions for what it finds.
The non-negotiable: alert on a metric moving outside its baseline confidence interval, not on it merely moving. "We alert on the judge score changing" is noise; "we alert when the judge score falls outside its 28-day moving-baseline CI, refresh the slice weekly, and trace each alert to a golden-set example" is monitoring. And Shreya Shankar’s SPADE work reframes the goal: production monitoring shouldn’t only alert, it should sample, label, and re-incorporate anomalies into the next calibration cycle — data-quality assertions at pipeline boundaries that feed the loop. For a model upgrade specifically, freeze a ~1,000-trace replay set, run the new model against it, score with the judge, and hold the upgrade if the delta exceeds the validated judge noise floor.
code
1Production monitoring -- the four-part stack (size to the regression you care about)23  Signal                  Cadence   Catches                    Action4  ---------------------   -------   ------------------------   ---------------------5  Behavioural assertions  live      policy/format violations   alert if 99%->95% /24h6  Reference-free judge    daily     quality drift (sampled)    alert outside CI band7  Distribution shift      live      novel failure types        embed + KL / topic mix8  Human spot review       weekly    the unknown unknowns       write a NEW assertion910  200 samples ~ +/-2.4% at a 3% defect rate -> size the sample so the CI is11  SMALLER than the regression you must detect. Vendor model upgrade? Freeze a12  ~1,000-trace replay set, score the new model, hold if delta > judge noise floor.

Online scoring in practice: sample, score, alert on a CI

The mechanics of the online judge are where teams either save money or page themselves to death. Sample, don’t score everything — 1–5% of live traffic through an expensive judge, 100% through cheap regex assertions — and size the sample to the regression you must catch: at a 3% defect rate, 200 samples gives ~±2.4%, so to detect a 1% move you need a larger or stratified sample. Then alert on a moving-baseline confidence interval, not on raw movement: compare today’s windowed score against a 28-day baseline and fire only when it falls outside the band. Stratify the sample by scenario tag so a regression confined to one slice (say, multilingual queries) isn’t washed out by the healthy majority — an aggregate score can stay green while a sub-population quietly breaks.
python
1# Online monitor: sample, judge, and alert only on a STATISTICALLY real drop.2def online_check(today_traces, judge, baseline_rate, baseline_ci_halfwidth,3                 sample_rate=0.02):4    sample = reservoir_sample(today_traces, rate=sample_rate)   # 1-5% of live traffic5    passes = sum(judge_pointwise(t.input, t.response, "on_policy", judge)6                 for t in sample)7    rate = passes / max(len(sample), 1)8    lo, hi = wilson_ci(passes, len(sample))                     # today's band9    # regression = today's UPPER bound is below the baseline's LOWER bound10    if hi < baseline_rate - baseline_ci_halfwidth:11        page("on_policy drift", rate=rate, ci=(lo, hi), n=len(sample))12    return {"rate": rate, "ci": (lo, hi), "n": len(sample)}13# Stratify by scenario tag so a slice-specific regression isn't hidden by the majority.

The cost of eval: an architecture decision

Eval cost is not an ops footnote — it belongs in the architecture review. Galileo reports some customers’ LLM-eval bills exceeding $1M/month at production scale; human evaluation runs $20–100/hour versus $0.03–15 per 1M tokens for an LLM judge. Three drivers explain the spread: (1) input/output length — judges on long-context RAG spend most of their tokens re-reading context; (2) judge choice — a fine-tuned 8B judge aligned to your domain rubric can be 30–50× cheaper than calling a frontier model every time; (3) frequency — running the judge on every prompt every hour wildly inflates cost versus sampling. The senior levers: pick the cheapest judge that clears the bias scorecard, sample don’t score every trace (reserve full coverage for regression suites and incident debugs), cache deterministic sub-scores and only run subjective sub-scores on samples, and wire eval spend onto the same dashboard as metric movement so finance and engineering see one number.
Interview angle. "Your eval bill is the largest line item — what do you cut?" The strong answer treats sampling rate as a tunable architectural parameter, not an accident of how often you remembered to invoke the judge: a 4-judge ensemble on every trace hits the $1M tier fast, while the same judges on a stratified 1% sample cost a fraction. Then cascade — cheap regex/schema on 100%, a small fine-tuned judge on the next tier, a frontier judge only on the ambiguous residual — and cache the deterministic claim-checks so you only pay the judge for the genuinely subjective calls. Naming the cascade and the sampling knob signals you’ve operated eval at scale, not just run it in a notebook.
code
1Eval cost levers (numbers to reason in, not memorize)23  Lever                    Effect                         Tradeoff4  ----------------------   ----------------------------   ----------------------5  Human eval               $20-100 / hour                 gold quality, doesn't scale6  Frontier LLM judge       $0.03-15 / 1M tokens           scales, has biases7  Fine-tuned 8B judge      ~30-50x cheaper than frontier  needs domain alignment8  Sample 1-5% (not 100%)   ~20-100x fewer judge calls     wider CI -> size to effect9  Cache deterministic      pay judge only for subjective  cache-invalidation care1011  $1M/month is real at full-coverage ensembles. Sampling rate is an12  ARCHITECTURAL parameter. Cascade: regex 100% -> small judge -> frontier residual.
articleWhat is eval-driven development (regression gates + GitHub Action)BraintrustdocsOpenAI API deprecations (Evals platform: read-only 2026-10-31, shutdown 2026-11-30)OpenAIarticleAddressing GenAI Evaluation Challenges: Cost & AccuracyGalileoarticleHow to Monitor LLMOps Performance with Drift MonitoringFiddlervideoBeyond the Hype: Monitoring LLMs in ProductionClaire Longo / MLOps.community

Checkpoint

A teammate writes the CI eval suite as exact-string-match assertions against reference answers. The build goes red on a valid paraphrase. What’s the fix?

ALock the model to temperature 0 so outputs are identical and exact-match worksBLayer the suite: deterministic checks (regex/schema) where outputs are constrained, semantic similarity vs a frozen golden baseline, and a judge only for subjective tiers — gate on recall of the risky classCRemove the failing tests so the build passes
Sign up free to answer and see why

Checkpoint

Your production monitor pages the team every time the daily judge score changes by any amount. The team is drowning in false alarms. Best fix?

AAlert only when the score falls outside its baseline confidence interval (e.g. a 28-day moving baseline), sized so the CI is smaller than the regression you must catchBRaise the alert threshold to a fixed 5-point dropCStop monitoring the judge score daily; check it monthly instead
Sign up free to answer and see why

Checkpoint

Your LLM-eval bill is now the single largest line item. Leadership wants it cut without losing the ability to catch regressions. What’s the highest-leverage move?

ASwitch the judge to the cheapest available model across the boardBRun the judge less often by lowering the eval cadence to weeklyCTreat sampling rate as a tunable parameter: cascade (regex on 100% → small fine-tuned judge → frontier judge on the residual), sample 1–5% for the expensive judge, and cache deterministic sub-scores
Sign up free to answer and see why

Checkpoint

A vendor model upgrade ships Friday. Your offline suite is green on the current model. What changes in your process before you let the upgrade reach users?

ANothing — the offline suite is green, so the upgrade is safe to roll outBFreeze a ~1,000-trace replay set, run the new model against it, score with the judge, and hold the upgrade if the delta exceeds the validated judge noise floorCIncrease the production sample rate to 100% for a week and watch
Sign up free to answer and see why

Checkpoint

Monitoring catches a failure mode in production that your CI gate never tested. After mitigating the live incident, what’s the durable follow-up?

APromote the failure into the golden set and add a CI assertion for it, so it can never regress silently againBNote it in the incident doc and move onCAdd a one-off manual check to the weekly review
Sign up free to answer and see why

Interview prep

Ops rounds test whether you can run eval as a production system: a layered CI gate, drift monitoring with confidence intervals, the offline-online feedback loop, and cost as a first-class constraint. Interviewers probe four things: do you layer deterministic-first and gate the build, do you alert on a CI not on movement, do you name the four drift vectors and the replay-set gate for model upgrades, and do you treat sampling rate and judge choice as cost levers. Use numbers.
  1. 01“LLM regression suite vs unit tests?” → not exact-match (many valid outputs); layer deterministic + semantic-similarity + judge, gate on risky-class recall.
  2. 02“How do you stop a fixed bug regressing?” → eval-driven development: promote the incident into the golden set + a CI assertion.
  3. 03“Name the drift vectors.” → prompt, model, user, and judge drift; each silently moves the eval score.
  4. 04“How do you detect a 5% regression in 24h?” → size the sample to the effect (~200 ≈ ±2.4% at 3% defect), alert outside the moving-baseline CI.
  5. 05“Gate a vendor model upgrade?” → frozen ~1,000-trace replay set, score the new model, hold if delta > judge noise floor.
  6. 06“Offline passes, online fails — what do you do?” → route the failure into offline as a new assertion; that loop is the point.
  7. 07“Your eval bill is the top line item — cut it?” → sampling rate is architectural; cascade (regex → small judge → frontier residual), cache deterministic sub-scores.
  8. 08“Human vs LLM judge cost?” → $20–100/hr vs $0.03–15/1M tokens; a fine-tuned 8B judge is ~30–50× cheaper than a frontier one.
Going deeper, the follow-ups: "distribution-shift on input embeddings moves before the judge does — false alarm?" (it can be, but it’s the earliest indicator and the only way to catch novel failure types; pair it with a weekly judge on stratified samples to resolve ambiguity); "how do you size the offline set per behaviour?" (a power calc — ~246 cases for an 80% pass rate at ±5%, 95% confidence); and "OpenAI Evals is shutting down — what’s your plan?" (treat any vendor as a complement with an integration plan; the platform goes read-only 2026-10-31, shuts 2026-11-30, so migrate the harness off it). Tie every answer to a number and the failure it prevents.

Could you stand up a layered CI eval gate, a drift monitor with confidence intervals, and a defensible eval-cost plan?

New to itGetting thereConfident

Takeaways

  • An LLM regression suite is layered, not exact-match: deterministic checks → semantic similarity → judge, gated on risky-class recall.
  • Eval-driven development: every production failure becomes a permanent golden-set case + CI assertion.
  • Monitor four drift vectors (prompt, model, user, judge); alert outside a baseline CI, never on mere movement.
  • Gate a vendor model upgrade with a frozen ~1,000-trace replay set and the judge noise floor.
  • The feedback loop — online failure → offline assertion — makes offline and online one system.
  • Eval cost is architectural: sampling rate is a knob, cascade judges, cache deterministic sub-scores ($1M/month is real).

Next: the capstone — assemble a golden set + judge + regression suite for a chatbot and prove it’s trustworthy.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.