Lesson 5 of 6 · 48 min

Launch gates & incident handling

Go/no-go gates with named owners and binary pass criteria, shadow → A/B → canary → full as the rollout order, the rehearsed rollback that is a P0 launch deliverable, AI-specific incident response (Gemini, Air Canada), and the four-family post-launch monitoring that catches drift.

A launch is a stress test on your eval plan

For a deterministic feature, “launch” is a deploy. For an AI feature it’s a stress test on your eval plan against the real input distribution — and the things that break are the ones your offline set never imagined. The senior discipline has three parts most teams under-build: a go/no-go gate where every row has a named owner and a binary pass criterion, a rehearsed rollback treated as a P0 launch sub-task (not a checkbox), and a post-launch monitoring stack that pages on a shift in the failure rate, not a single bad output. This lesson assembles all three, anchored on real incidents.
The launch order that minimises risk is shadow → A/B → canary → full, and it exists because LLM A/B tests need new statistics. Statsig frames shadow testing as the lowest-risk default: run the new model alongside production on the same traffic but discard its outputs, so you see the real distribution with zero customer exposure. Then A/B (split traffic) when two variants are both shippable and ranking matters — but Traceloop warns you must “wait until you have a large enough sample size to ensure your results are statistically significant” because “LLM outputs can vary,” and many LLM evaluations need very large samples. Hamel & Shankar make it precise: ship online metrics with confidence intervals and launch only when the lower bound crosses your threshold.
code
1ROLLOUT ORDER (lowest risk first)23  STAGE     USER IMPACT          WHEN                          STAT RISK4  -------   ------------------   ---------------------------   ------------------------5  Shadow    none (output         early model swaps / prompt    lowest — no exposure6            discarded)           changes, uncertain upside7  A/B       random subset see    two variants both shippable,  higher — many LLM evals8            variant              ranking matters               need large N for sig.9  Canary    first 1-5% see       regression evals green, want  medium — small N is slow10            variant              real-distribution signal       to converge11  Full      everyone             lower CI bound crosses the    -12                                 threshold1314  Notion ships frontier models in <24hrs by running regression+frontier15  evals BEFORE traffic — speed comes from trustworthy offline gates.
Notion’s pattern shows the velocity payoff of doing gates well: a <24-hour target from a frontier-model release to customer availability, achievable precisely because regression evals (broad CI) plus frontier evals (model-differentiating) run before the new model reaches live traffic. The gate is the launch infrastructure: any new model must pass regression AND show positive deltas on frontier before it ships. Interview angle. “How do you ship a model change safely and fast?” → trustworthy offline gates first (so you don’t need a long, risky A/B), then shadow → canary; speed and safety are not opposites if the eval suite is the gate.
Shipping AI That Works: An Evaluation Framework for PMsAI Engineer (Aman Khan, Arize)

The go/no-go gate: named owners, binary criteria

A launch-gate checklist only matters if every row has a named owner and a binary pass criterion — “the responsible-AI team will review” is not a gate. Distilling NIST, Anthropic, Microsoft, and Meta into something a PM applies on demo day: a context-of-use memo (PM); an Impact Assessment (PM + Legal); a capability evaluation below your defined thresholds (Research); a red-team report where all Critical findings have a mitigation or accepted-residual status (Safety); guardrails active with ≥2 layers per path and a tested kill switch (Engineering); a signed release plan (PM + Release Manager); a published Risk Report; a rehearsed rollback (staging rollback completed in under 15 minutes); post-launch monitoring live on day one; and a Model/System Card published. Each is a binary, owned object.
The row teams most often fake is the rollback. A kill-switch that has not been rehearsed on staging is not a kill switch — it is a hypothesis. Treat the rollback as a P0 sub-task of the launch with a rehearsal drill and a measured time-to-revert, and make it a feature-flag or model-revert you’ve actually executed, not a runbook nobody has run. This is the lesson Google’s Gemini incident taught publicly: when the image generator mis-rendered historical figures (Feb 2024), Google had to pause the people-image feature within days — CEO Sundar Pichai called the outputs “completely unacceptable” and co-founder Sergey Brin admitted “we definitely messed up” with “not thorough testing.” Experts attributed the failure to shipping ahead of evaluation under competitive pressure.
AI incidents have a property normal outages don’t: a single confident hallucination can be legally binding speech. In Moffatt v. Air Canada (2024), the airline’s chatbot misquoted bereavement-fare policy; Air Canada argued the chatbot was “responsible for its own actions,” and the British Columbia Civil Resolution Tribunal ruled against the airline and ordered it to pay damages, treating the bot’s statements as attributable to the company. Separately, a French court flagged “untraceable” AI-hallucinated case law in a legal filing. The mechanism: probabilistic outputs without retrieved grounding become actionable utterances. The PM consequence — release customer-facing GenAI only with retrieval grounding and a confidence threshold below which it escalates to a human, and make citation-existence a launch-gate eval (any system that can emit URLs/citations needs an automated existence check).
A complete AI incident-handling stack has six layers, each a concrete artifact: detection (severity classifier off the Microsoft AI Bug Bar, paged with a severity tag), triage (categorise against a public taxonomy — NIST 600-1 / MLCommons), mitigation (single-rail disable via the guardrail runbook), communication (a templated external statement, mirroring Google’s Feb-2024 pattern), rollback (the rehearsed kill-switch drill), and postmortem (a blameless write-up that feeds the next eval set). The detection layer is where the distributional mindset matters most: a single bad output is not an incident — a shift in the failure rate is. You page on hallucination rate rising >10pp week-over-week, not on one user’s screenshot.

Post-launch monitoring: quality, cost, latency, drift

Post-launch monitoring is the production-side mirror of the eval suite, and Braintrust divides the live signal into four metric families a PM must instrument simultaneously on day one: quality (groundedness, hallucination rate, factuality, harm rates), cost (token spend per task, model-routing waste, retries), latency (p50/p95 time-to-first-token, tail behaviour), and drift (semantic drift in inputs and outputs, embedding-space drift, label-distribution drift). Drift is the new, most-often-missing leg of the stool — it’s how you discover that the input distribution shifted out from under a model that still looks healthy on every individual request.
code
1POST-LAUNCH MONITORING — four families, day one (Braintrust)23  FAMILY     METRICS                              PAGE WHEN4  --------   ----------------------------------   -------------------------------5  Quality    groundedness, hallucination rate,    weekly aggregate regresses past6             factuality, harm rate                baseline + control limit7  Cost       $/task, routing waste, retries       daily spend > budget - buffer8  Latency    p50 / p95 TTFT, tail                 p95 crosses the user-facing SLO9  Drift      semantic + embedding + label drift   statistical-significance10                                                  threshold crossed1112  RULE: every live-dashboard metric must trace to a row in the pre-launch13  eval plan. If it's not in CI, it's unmeasurable in production.14  WRITE-BACK: every incident sample becomes a permanent eval case.
Two senior rules close the loop. First, every metric on the live dashboard must trace back to a row in the pre-launch eval plan — if a metric is missing in CI, it’s unmeasurable in production, which is why monitoring and the eval suite are the same artifact viewed from two sides. Second, write-back: every incident sample is folded back into the eval set as a permanent test case, so the failure can never silently regress in again (the discipline from L1, L3, and L4). Page on statistical-control rules — a quality lower bound, a cost spike, a drift-significance threshold — not on individual outputs, because a probabilistic product fails on distributions. Interview angle. “What do you monitor after launch?” → quality / cost / latency / drift simultaneously, each tracing to a CI row, paging on a rate shift, with incident write-back into the eval set.
articleBraintrust — What is LLM monitoring? (quality, cost, latency, drift)BraintrustarticleAir Canada held liable for its chatbot giving a passenger bad adviceBBCarticleGoogle apologizes after Gemini generated inaccurate historical imagesThe VergearticleThe Definitive Guide to A/B Testing LLM Models in ProductionTraceloop

Checkpoint

You’re about to roll out a new model that improves offline scores. You have low confidence about real-world behaviour and want zero customer risk while you learn. First step?

AShadow test — run the new model on real traffic but discard its outputs, comparing against production with no customer exposureBRun a 50/50 A/B test immediately to get the fastest readCShip to 100% and watch the dashboards
Sign up free to answer and see why

Checkpoint

A launch checklist row reads “Rollback: documented in the runbook.” Why does a senior PM mark this gate as not-passed?

ARunbooks are obsolete — rollback should be fully automated with no docsBA documented-but-unrehearsed kill switch is a hypothesis, not a kill switch — require a staging rollback drill with a measured time-to-revert (target <15 min)CThe runbook should be owned by the PM, not engineering
Sign up free to answer and see why

Checkpoint

Your customer-facing support bot occasionally states a refund policy that doesn’t exist. Beyond fixing the prompt, what’s the launch-gate-level concern?

ANone — occasional wrong answers are inherent to LLMs and acceptableBThe bot’s confident misstatements are legally attributable to the company (Air Canada) — gate on retrieval grounding, a confidence-threshold human escalation, and citation/policy-existence eval checksCSwitch to a larger model so it hallucinates less
Sign up free to answer and see why

Checkpoint

You instrument quality, cost, and p95 latency after launch. A reviewer says the monitoring is still incomplete for an AI product. What’s missing?

ADrift — semantic/embedding/label drift detection, which catches the input distribution shifting under a model that still looks healthy per requestBNothing — quality, cost, and latency fully cover an AI productCTotal request count
Sign up free to answer and see why

Checkpoint

On-call gets one screenshot of a wrong AI answer from a single user. How should the incident process treat it?

AImmediately roll back the feature — any wrong output is a SEV-1BLog it as a sample and check the failure-rate trend; escalate only if the rate has shifted (e.g. hallucination rate up >10pp WoW), and fold the case into the eval setCIgnore it — single outputs never matter
Sign up free to answer and see why

Interview prep

The launch/incident round tests whether you reason about rollout as a risk-staged process with owned gates and a rehearsed rollback — and whether you treat AI incidents as distributional and legally consequential. The behavioural version (“tell me about an AI feature that backfired”) is graded on system-design literacy, not empathy: name the detector, the kill switch, the comms, and the eval write-back, not “I escalated quickly and reassured customers.”
  1. 01“How do you roll out a model change safely?” → shadow → A/B → canary → full; gate full rollout on the lower CI bound crossing the threshold.
  2. 02“How can you ship fast AND safely?” → trustworthy offline gates (regression + frontier) before traffic, like Notion’s <24hr frontier deploys.
  3. 03“What’s your rollback plan?” → a rehearsed, timed kill switch (feature flag / model revert), staging drill <15 min — a P0 launch deliverable, not a runbook.
  4. 04“The model is confidently wrong in front of a customer — fallback?” → confidence-gated escalation to a human, grounding, logged incident, retraining trigger.
  5. 05“Is a chatbot’s wrong answer the company’s liability?” → yes (Air Canada) — gate on grounding + escalation + citation-existence checks.
  6. 06“What do you monitor post-launch?” → quality / cost / latency / drift simultaneously, each tracing to a CI row, paging on a rate shift.
  7. 07“When is a bad output an incident?” → when the failure RATE shifts (e.g. hallucination up >10pp WoW), not on a single sample.
  8. 08“Walk me through handling an AI incident.” → detect (severity tag) → triage (taxonomy) → mitigate (rail disable) → comms (template) → rollback (drill) → postmortem → eval write-back.
Going deeper: “what’s your SLA for acknowledging an AI incident, and does it differ from a normal outage?” (yes — tie it to the Bug-Bar severity: Critical content-safety = 15 min); “what did you add to your golden eval set after the incident?” (the failing cases as permanent regression tests — write-back is the close-the-loop signal); and “how did you communicate to non-technical leadership?” (a blameless postmortem with the system fix and explicit re-launch criteria, mirroring Google’s public pattern). Always: detector, kill switch, comms template, eval write-back.

Could you define a go/no-go gate with owners, sequence a shadow→canary rollout, rehearse a rollback, and stand up four-family monitoring with incident write-back?

New to itGetting thereConfident

Takeaways

  • Roll out lowest-risk-first: shadow → A/B → canary → full; gate full rollout on the lower CI bound crossing the threshold.
  • A go/no-go gate is only real if every row has a named owner and a binary pass criterion.
  • Rollback is a P0 deliverable — rehearse the kill switch on staging and measure time-to-revert (<15 min).
  • AI outputs are legally attributable (Air Canada) — gate customer-facing GenAI on grounding + confidence-escalation + citation-existence checks.
  • Monitor four families simultaneously — quality, cost, latency, drift — each tracing to a CI row.
  • An incident is a shift in the failure rate, not one bad output; fold every incident sample back into the eval set.

Next: the capstone — write the eval suite and launch plan for an AI feature, exactly as the PM case round expects.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.