Go/no-go gates with named owners and binary pass criteria, shadow → A/B → canary → full as the rollout order, the rehearsed rollback that is a P0 launch deliverable, AI-specific incident response (Gemini, Air Canada), and the four-family post-launch monitoring that catches drift.
A launch is a stress test on your eval plan
For a deterministic feature, “launch” is a deploy. For an AI feature it’s a stress test on your eval plan against the real input distribution — and the things that break are the ones your offline set never imagined. The senior discipline has three parts most teams under-build: a go/no-go gate where every row has a named owner and a binary pass criterion, a rehearsed rollback treated as a P0 launch sub-task (not a checkbox), and a post-launch monitoring stack that pages on a shift in the failure rate, not a single bad output. This lesson assembles all three, anchored on real incidents.
The launch order that minimises risk is shadow → A/B → canary → full, and it exists because LLM A/B tests need new statistics. Statsig frames shadow testing as the lowest-risk default: run the new model alongside production on the same traffic but discard its outputs, so you see the real distribution with zero customer exposure. Then A/B (split traffic) when two variants are both shippable and ranking matters — but Traceloop warns you must “wait until you have a large enough sample size to ensure your results are statistically significant” because “LLM outputs can vary,” and many LLM evaluations need very large samples. Hamel & Shankar make it precise: ship online metrics with confidence intervals and launch only when the lower bound crosses your threshold.
code
1ROLLOUT ORDER (lowest risk first)23 STAGE USER IMPACT WHEN STAT RISK4 ------- ------------------ --------------------------- ------------------------5 Shadow none (output early model swaps / prompt lowest — no exposure6 discarded) changes, uncertain upside7 A/B random subset see two variants both shippable, higher — many LLM evals8 variant ranking matters need large N for sig.9 Canary first 1-5% see regression evals green, want medium — small N is slow10 variant real-distribution signal to converge11 Full everyone lower CI bound crosses the -12 threshold1314 Notion ships frontier models in <24hrs by running regression+frontier15 evals BEFORE traffic — speed comes from trustworthy offline gates.
Notion’s pattern shows the velocity payoff of doing gates well: a <24-hour target from a frontier-model release to customer availability, achievable precisely because regression evals (broad CI) plus frontier evals (model-differentiating) run before the new model reaches live traffic. The gate is the launch infrastructure: any new model must pass regression AND show positive deltas on frontier before it ships. Interview angle. “How do you ship a model change safely and fast?” → trustworthy offline gates first (so you don’t need a long, risky A/B), then shadow → canary; speed and safety are not opposites if the eval suite is the gate.
A launch-gate checklist only matters if every row has a named owner and a binary pass criterion — “the responsible-AI team will review” is not a gate. Distilling NIST, Anthropic, Microsoft, and Meta into something a PM applies on demo day: a context-of-use memo (PM); an Impact Assessment (PM + Legal); a capability evaluation below your defined thresholds (Research); a red-team report where all Critical findings have a mitigation or accepted-residual status (Safety); guardrails active with ≥2 layers per path and a tested kill switch (Engineering); a signed release plan (PM + Release Manager); a published Risk Report; a rehearsed rollback (staging rollback completed in under 15 minutes); post-launch monitoring live on day one; and a Model/System Card published. Each is a binary, owned object.
The row teams most often fake is the rollback. A kill-switch that has not been rehearsed on staging is not a kill switch — it is a hypothesis. Treat the rollback as a P0 sub-task of the launch with a rehearsal drill and a measured time-to-revert, and make it a feature-flag or model-revert you’ve actually executed, not a runbook nobody has run. This is the lesson Google’s Gemini incident taught publicly: when the image generator mis-rendered historical figures (Feb 2024), Google had to pause the people-image feature within days — CEO Sundar Pichai called the outputs “completely unacceptable” and co-founder Sergey Brin admitted “we definitely messed up” with “not thorough testing.” Experts attributed the failure to shipping ahead of evaluation under competitive pressure.
AI incident response: probabilistic products, legal teeth
AI incidents have a property normal outages don’t: a single confident hallucination can be legally binding speech. In Moffatt v. Air Canada (2024), the airline’s chatbot misquoted bereavement-fare policy; Air Canada argued the chatbot was “responsible for its own actions,” and the British Columbia Civil Resolution Tribunal ruled against the airline and ordered it to pay damages, treating the bot’s statements as attributable to the company. Separately, a French court flagged “untraceable” AI-hallucinated case law in a legal filing. The mechanism: probabilistic outputs without retrieved grounding become actionable utterances. The PM consequence — release customer-facing GenAI only with retrieval grounding and a confidence threshold below which it escalates to a human, and make citation-existence a launch-gate eval (any system that can emit URLs/citations needs an automated existence check).
A complete AI incident-handling stack has six layers, each a concrete artifact: detection (severity classifier off the Microsoft AI Bug Bar, paged with a severity tag), triage (categorise against a public taxonomy — NIST 600-1 / MLCommons), mitigation (single-rail disable via the guardrail runbook), communication (a templated external statement, mirroring Google’s Feb-2024 pattern), rollback (the rehearsed kill-switch drill), and postmortem (a blameless write-up that feeds the next eval set). The detection layer is where the distributional mindset matters most: a single bad output is not an incident — a shift in the failure rate is. You page on hallucination rate rising >10pp week-over-week, not on one user’s screenshot.
Post-launch monitoring is the production-side mirror of the eval suite, and Braintrust divides the live signal into four metric families a PM must instrument simultaneously on day one: quality (groundedness, hallucination rate, factuality, harm rates), cost (token spend per task, model-routing waste, retries), latency (p50/p95 time-to-first-token, tail behaviour), and drift (semantic drift in inputs and outputs, embedding-space drift, label-distribution drift). Drift is the new, most-often-missing leg of the stool — it’s how you discover that the input distribution shifted out from under a model that still looks healthy on every individual request.
code
1POST-LAUNCH MONITORING — four families, day one (Braintrust)23 FAMILY METRICS PAGE WHEN4 -------- ---------------------------------- -------------------------------5 Quality groundedness, hallucination rate, weekly aggregate regresses past6 factuality, harm rate baseline + control limit7 Cost $/task, routing waste, retries daily spend > budget - buffer8 Latency p50 / p95 TTFT, tail p95 crosses the user-facing SLO9 Drift semantic + embedding + label drift statistical-significance10 threshold crossed1112 RULE: every live-dashboard metric must trace to a row in the pre-launch13 eval plan. If it's not in CI, it's unmeasurable in production.14 WRITE-BACK: every incident sample becomes a permanent eval case.
Two senior rules close the loop. First, every metric on the live dashboard must trace back to a row in the pre-launch eval plan — if a metric is missing in CI, it’s unmeasurable in production, which is why monitoring and the eval suite are the same artifact viewed from two sides. Second, write-back: every incident sample is folded back into the eval set as a permanent test case, so the failure can never silently regress in again (the discipline from L1, L3, and L4). Page on statistical-control rules — a quality lower bound, a cost spike, a drift-significance threshold — not on individual outputs, because a probabilistic product fails on distributions. Interview angle. “What do you monitor after launch?” → quality / cost / latency / drift simultaneously, each tracing to a CI row, paging on a rate shift, with incident write-back into the eval set.
You’re about to roll out a new model that improves offline scores. You have low confidence about real-world behaviour and want zero customer risk while you learn. First step?
AShadow test — run the new model on real traffic but discard its outputs, comparing against production with no customer exposureBRun a 50/50 A/B test immediately to get the fastest readCShip to 100% and watch the dashboards
A launch checklist row reads “Rollback: documented in the runbook.” Why does a senior PM mark this gate as not-passed?
ARunbooks are obsolete — rollback should be fully automated with no docsBA documented-but-unrehearsed kill switch is a hypothesis, not a kill switch — require a staging rollback drill with a measured time-to-revert (target <15 min)CThe runbook should be owned by the PM, not engineering
Your customer-facing support bot occasionally states a refund policy that doesn’t exist. Beyond fixing the prompt, what’s the launch-gate-level concern?
ANone — occasional wrong answers are inherent to LLMs and acceptableBThe bot’s confident misstatements are legally attributable to the company (Air Canada) — gate on retrieval grounding, a confidence-threshold human escalation, and citation/policy-existence eval checksCSwitch to a larger model so it hallucinates less
You instrument quality, cost, and p95 latency after launch. A reviewer says the monitoring is still incomplete for an AI product. What’s missing?
ADrift — semantic/embedding/label drift detection, which catches the input distribution shifting under a model that still looks healthy per requestBNothing — quality, cost, and latency fully cover an AI productCTotal request count
On-call gets one screenshot of a wrong AI answer from a single user. How should the incident process treat it?
AImmediately roll back the feature — any wrong output is a SEV-1BLog it as a sample and check the failure-rate trend; escalate only if the rate has shifted (e.g. hallucination rate up >10pp WoW), and fold the case into the eval setCIgnore it — single outputs never matter
The launch/incident round tests whether you reason about rollout as a risk-staged process with owned gates and a rehearsed rollback — and whether you treat AI incidents as distributional and legally consequential. The behavioural version (“tell me about an AI feature that backfired”) is graded on system-design literacy, not empathy: name the detector, the kill switch, the comms, and the eval write-back, not “I escalated quickly and reassured customers.”
01“How do you roll out a model change safely?” → shadow → A/B → canary → full; gate full rollout on the lower CI bound crossing the threshold.
02“How can you ship fast AND safely?” → trustworthy offline gates (regression + frontier) before traffic, like Notion’s <24hr frontier deploys.
03“What’s your rollback plan?” → a rehearsed, timed kill switch (feature flag / model revert), staging drill <15 min — a P0 launch deliverable, not a runbook.
04“The model is confidently wrong in front of a customer — fallback?” → confidence-gated escalation to a human, grounding, logged incident, retraining trigger.
05“Is a chatbot’s wrong answer the company’s liability?” → yes (Air Canada) — gate on grounding + escalation + citation-existence checks.
06“What do you monitor post-launch?” → quality / cost / latency / drift simultaneously, each tracing to a CI row, paging on a rate shift.
07“When is a bad output an incident?” → when the failure RATE shifts (e.g. hallucination up >10pp WoW), not on a single sample.
08“Walk me through handling an AI incident.” → detect (severity tag) → triage (taxonomy) → mitigate (rail disable) → comms (template) → rollback (drill) → postmortem → eval write-back.
Going deeper: “what’s your SLA for acknowledging an AI incident, and does it differ from a normal outage?” (yes — tie it to the Bug-Bar severity: Critical content-safety = 15 min); “what did you add to your golden eval set after the incident?” (the failing cases as permanent regression tests — write-back is the close-the-loop signal); and “how did you communicate to non-technical leadership?” (a blameless postmortem with the system fix and explicit re-launch criteria, mirroring Google’s public pattern). Always: detector, kill switch, comms template, eval write-back.
Could you define a go/no-go gate with owners, sequence a shadow→canary rollout, rehearse a rollback, and stand up four-family monitoring with incident write-back?
New to itGetting thereConfident
Takeaways
Roll out lowest-risk-first: shadow → A/B → canary → full; gate full rollout on the lower CI bound crossing the threshold.
A go/no-go gate is only real if every row has a named owner and a binary pass criterion.
Rollback is a P0 deliverable — rehearse the kill switch on staging and measure time-to-revert (<15 min).
AI outputs are legally attributable (Air Canada) — gate customer-facing GenAI on grounding + confidence-escalation + citation-existence checks.
Monitor four families simultaneously — quality, cost, latency, drift — each tracing to a CI row.
An incident is a shift in the failure rate, not one bad output; fold every incident sample back into the eval set.
Next: the capstone — write the eval suite and launch plan for an AI feature, exactly as the PM case round expects.