Lesson 6 of 6 · 50 min

Capstone: eval suite + launch plan

The whole track becomes one deliverable: given an AI feature, write the eval suite AND the launch plan exactly as a PM case round expects — success criteria, the metric stack, the offline/online/human eval mix, the responsible-AI gates, and the rollout/rollback/incident plan, assembled into one defensible artifact you can whiteboard end to end.

Capstone

Eval suite + launch plan

Where the five lessons become one artifact

Each lesson so far solved one piece in isolation. The capstone is the integration test: given a real AI feature, produce two documents a senior AI PM is expected to write before launch — an eval suite (how you measure success) and a launch plan (how you ship it without ending up in the news). This is the exact shape of the AI-PM case round: graders reward a candidate who refuses the surface question, names a layered system, names the trade-off concretely, and pre-declares a rollback trigger — the identical pattern across every question family. We’ll build both documents on one worked feature, then rehearse defending every decision under follow-up pressure.
The worked feature, because a case round always grounds you in one: a customer-facing support assistant for a travel company — it answers policy questions (refunds, bereavement fares, baggage) by retrieving from a help-center corpus and, where allowed, takes an action (rebook, issue a credit). It is deliberately the Air Canada shape from L5: customer-facing, retrieval-grounded, action-taking, and legally consequential. Everything that follows — the success criteria, the metric stack, the eval portfolio, the gates, the rollout, the incident plan — is built for this feature, because the senior signal is a plan tuned to the task, not a generic checklist recited from memory.
How To ACE AI Product Sense Interviews (OpenAI PM Mock Interview)Aakash Gupta

Step 1 — Success criteria: the distributional definition of done

Open every case the same way: refuse the binary and write a distributional definition of done (L1). For the support assistant: “For 90% of in-scope policy questions, the assistant returns a grounded, correct answer with a citation; for the remaining 10%, it escalates to a human rather than guessing — and the unsupported-claim rate stays below a hard floor.” Note the three parts: a target rate on a scoped input class (in-scope policy questions, not “anything a user types”), a graceful tail (escalate, don’t confabulate), and an explicit floor on the failure that hurts (ungrounded claims). Then earn the AI: state the rule-based baseline — an FAQ-lookup / decision-tree — and commit to shipping the assistant only if it beats that baseline on the same labelled set (PAIR’s unique-value test). And pick the mode: this is augmentation with a human in the loop, not full automation, which dictates the metric stack next.
code
1DETERMINISTIC "done"  vs  DISTRIBUTIONAL "done" (support assistant)23  DETERMINISTIC (wrong here)         DISTRIBUTIONAL (what you write)4  --------------------------------   ------------------------------------------5  "answers policy questions          ">=90% of IN-SCOPE policy Qs: grounded +6   correctly"                         correct + cited"7  "no P0 bugs"                        "the other <=10%: ESCALATE, never guess"8  (no failure floor)                  "unsupported-claim rate < hard floor X%"9  (AI assumed)                        "beats the FAQ/decision-tree baseline on10                                       the same labelled set, or ship the rule"1112  Three new parts a binary never had: a RATE, a SCOPED class, a graceful TAIL13  -- plus a floor on the dangerous failure and a baseline the AI must beat.

Step 2 — The metric stack and the divergence alarm

Now lay the metric stack (L2) — one dominant metric per layer, never a single blended score. For the support assistant: business / NSM — tickets deflected (resolved without a human); product engagement — repeat-use and self-service rate; model quality — grounded-faithfulness and unsupported-claim rate; system / reliability — p95 latency, $/resolution, and escalation-to-human rate. Then the instrument a non-AI PM never builds: the divergence alarm. Deflection (the NSM) can climb while the model quietly rots on a slice — a new fare type, a new language — because the 90% who are fine keep deflecting. So a held-out, per-slice faithfulness gate runs independently and blocks promotion if faithfulness drops more than a set delta, regardless of what deflection is doing. Bind each quality metric to a cost-of-error: a wrong baggage answer is an annoyance; a wrong bereavement-fare answer is the Air Canada lawsuit.
code
1SUPPORT ASSISTANT — metric stack (one per layer) + divergence alarm23  LAYER            DOMINANT METRIC               COST-OF-ERROR BINDING4  --------------   ---------------------------   ----------------------------5  Business / NSM   tickets deflected             low deflection = cost, not harm6  Product engage.  self-service / repeat use     proxy for trust7  Model quality    grounded faithfulness;        wrong policy answer = legal8                   unsupported-claim rate        liability (Air Canada)9  System / reliab. p95 latency; $/resolution;    high escalation = AI not earning10                   escalation-to-human rate      its keep1112  DIVERGENCE ALARM: weekly held-out, PER-SLICE faithfulness gate, independent13  of deflection; blocks model promotion on a faithfulness drop > set delta.
The case-round move here is sequencing: scope the task, name one metric per layer, then name the divergence alarm — and refuse to recite a metric before you’ve scoped what “success” means for this feature. Interview angle. The probe “your deflection rate is up — declare victory?” is the L2 trap; the strong answer is “not yet — I’d check whether model quality diverged from the rising NSM via the held-out per-slice faithfulness gate, because deflection can climb while the model regresses on a cohort.”

Step 3 — The eval suite: portfolio, modalities, calibrated judge

The eval suite is a portfolio, not a dataset (L1, L3). Three sets: a tight regression set (~100 purpose-built CI examples with deterministic assertions — citation present, policy-id valid, JSON schema, refusal on out-of-scope), a sampled-production set (100+ fresh traces every 2–4 weeks scored by LLM-as-judge to catch drift), and a small frontier set (~20 hand-curated cases to differentiate competing models, Notion’s pattern). Build coverage by dimension (policy type × language × question intent × adversarial / “should-refuse”), not by row count — a missing category is a silent confidence error. Harvest the gold contexts from the production retriever’s runs, not hand-picked passages, or offline scores read higher than reality. And for a RAG feature, evaluate retrieval (recall@k) before generation (faithfulness) — a generation metric can’t fix a retrieval miss.
code
1EVAL SUITE = three sets + the modality mix23  SET              SIZE / CADENCE            GRADER              CATCHES4  --------------   -----------------------   -----------------   ------------------5  Regression (CI)  ~100, frozen, per build   deterministic       regressions6                                             assertions7  Sampled prod     100+ / 2-4 weeks          LLM-as-judge on     drift you never8                                             sampled traces      wrote a test for9  Frontier         ~20, hand-curated         judge + human       model differences1011  MODALITY MIX  ~80% deterministic CI  /  ~15% judge on samples  /  ~5% human12  Coverage by DIMENSION: policy type x language x intent x adversarial.13  RAG: score retrieval (recall@k) BEFORE generation (faithfulness).
The judge that makes the online layer affordable is a biased instrument, not an oracle (L3). Name its traps and the mechanical mitigations: position bias → randomise order, score both orderings; verbosity bias → length-normalise, penalise padding; self-preference → judge with a different model family than the assistant; reference-answer bias → score reasoning against criteria, not string overlap. Then calibrate: humans build a golden set, the judge runs continuously, and you track the judge-vs-human agreement rate and recalibrate the rubric quarterly. The division of labour: humans set and recalibrate the bar; the judge scales it; deterministic assertions carry the high-volume regression load. Interview angle. “Walk me through your eval harness” wants exactly this separation — offline golden/regression with precision/recall/faithfulness, online task-completion/escalation/correction-rate, a fallback when the model is wrong, and a named ship gate.
Two distinctions the grader is specifically listening for. First, an eval is not a vibe check: an eval is a scored rubric against ground truth that gates something; a vibe check is a subjective thumbs-up that gates nothing. If you can’t state that difference in one sentence, the round marks you down. Second, build the set from explicit failure hypotheses, not just happy paths — write down, up front, the ways this assistant will fail (cites the wrong policy version, answers an out-of-scope legal question, mixes languages, invents a refund window) and make each a negative test case. Anthropic’s “balanced sets” discipline is the same point: test where the behaviour should and should not occur, so the assistant must refuse, not only answer.
An eval suite is not “a few prompts in a spreadsheet.” It is a portfolio (regression + sampled-prod + frontier), measured across three modalities (deterministic / judge / human) with a deliberate ~80/15/5 split, covered by dimension, and gated by a judge that is calibrated to a human golden set. That sentence, said unprompted, is the senior signal.

Step 4 — Responsible-AI gates + the go/no-go table

Bolt on the responsible-AI gates (L4) and assemble the go/no-go table (L5) — the deliverable that turns the plan into a launch decision. Every row has a named owner and a binary pass criterion; “the responsible-AI team will review” is not a gate. Anchor it to the NIST functions in shape (GOVERN/MAP/MEASURE/MANAGE), map the relevant NIST GenAI-profile risks (confabulation, data privacy, IP, human-AI configuration) to eval cases in the PRD — that’s building responsible AI in, not auditing for it — and attach a one-page Risk Report (categories × levels × mitigations × residual uncertainty × rollback), the artifact every frontier lab converges on. For the action-taking part, ship execution-rail guardrails: least-privilege scoped permissions, human confirmation for irreversible actions (issuing a credit), spend caps and rate limits with numbers, and a documented kill switch.
code
1GO/NO-GO TABLE — support assistant (every row: owner + binary criterion)23  GATE                          OWNER            PASS CRITERION4  ---------------------------   --------------   ---------------------------------5  Context-of-use memo (MAP)     PM               user/task/failure-mode named6  Distributional done + base-   PM               90% grounded+cited; tail escalates;7    line beat                                    beats FAQ baseline on labelled set8  Eval suite green (MEASURE)    PM + Eng         regression + frontier pass; judge9                                                 calibrated to golden set10  Risk Report published         PM               categories x levels x mitigations11                                                 x residual x rollback12  Guardrails active             Eng              >=2 rails/path; scoped perms +13                                                 confirm + caps; kill switch TESTED14  Rollback rehearsed            SRE + PM         staging revert measured < 15 min15  Monitoring live day one       SRE              quality/cost/latency/DRIFT + alerts16  Release plan signed (MANAGE)  PM + Rel Mgr     all rows met or residual accepted
The single most-faked row is the rollback. A kill switch that has not been rehearsed on staging is a hypothesis, not a kill switch — treat it as a P0 launch sub-task with a measured time-to-revert, because incidents are exactly when un-rehearsed plans fail (Google’s Gemini pause is the public lesson). And the legal row is non-negotiable for this feature: an ungrounded confident misstatement is attributable, legally-actionable speech the company owns (Air Canada), so grounding + a confidence-threshold human escalation + citation/policy-existence eval checks are launch-gate criteria, not nice-to-haves.
Responsible-AI questions in the case round test two literacies at once — ethics (can you hold a launch when a subgroup is harmed) and architecture (can you build constraint into the design, not the audit). The probe “your model works for 90% of users but poorly for one demographic — what do you do, and on what timeline?” has a known strong shape: segment the harm (stratify the eval across the affected slice), quantify it (what harm, how frequent, what severity), ship the threshold not the apology (“we don’t auto-launch for group X until faithfulness on that slice hits Y; until then we shadow-serve / human-review that slice”), and name the owner of remediation with a deadline. The weak answer is the generic platitude — “we’d test for bias and check edge cases” — which names no taxonomy, no threshold, no slice, no gate, and no owner.

Step 5 — Rollout, monitoring, and the incident plan

Sequence the rollout lowest-risk-first: shadow → A/B → canary → full (L5). Shadow the assistant on real tickets with outputs discarded to see the true input distribution at zero customer risk; A/B only when two variants are both shippable, waiting for a sample large enough for significance (LLM outputs vary); canary to 1–5%; and gate full rollout on the lower confidence-interval bound crossing the threshold, not the point estimate. Stand up four-family monitoring on day one — quality, cost, latency, and drift (the most-often-missed leg) — with every live metric tracing back to a row in the eval plan. Then the incident plan: detect on a rate shift (a single bad answer is not an incident; hallucination rate up >10pp WoW is), triage against a taxonomy, mitigate by disabling a single rail, communicate from a template, roll back via the rehearsed drill, and write every incident sample back into the eval set as a permanent regression case.
code
1LAUNCH PLAN — rollout + monitoring + incident loop23  ROLLOUT   shadow -> A/B -> canary -> full4            gate FULL on lower CI bound crossing threshold (not point estimate)56  MONITOR   quality | cost | latency | DRIFT   (day one; each traces to a CI row)7            page on a RATE SHIFT, not a single output89  INCIDENT  detect (severity tag) -> triage (taxonomy) -> mitigate (rail disable)10            -> comms (template) -> rollback (rehearsed drill) -> postmortem11            -> WRITE-BACK: incident sample becomes a permanent eval case
Step back and read the whole artifact as one loop: the eval suite defines and measures success, the gates decide go/no-go, the rollout exposes the plan to the real distribution gradually, monitoring catches what the offline set never imagined, and incident write-back folds every surprise back into the eval suite — so the same failure can never silently return. That loop is the deliverable. A demo that worked three times is not it; a distribution measured against a written bar, gated, staged, monitored, and self-correcting, is.
articleBeyond vibe checks: A PM’s complete guide to evalsLenny’s Newsletter (Aman Khan)articleDemystifying evals for AI agents (pass@k, balanced sets, graders)Anthropic EngineeringdocsPeople + AI Guidebook — User Needs + Defining SuccessGoogle PAIRarticleAI Product Manager Interview Questions 2026 (strong vs weak answers)KORE1

Checkpoint

You’re given the support-assistant case and asked “the model is 85% accurate — do we ship?” What’s the strongest first move?

ARefuse the binary and write a distributional done: a target rate on a scoped input class, a graceful tail (escalate, don’t guess), a hard floor on unsupported claims, and a baseline the AI must beatBAsk which model and provider it is, then pick the highest-accuracy optionCSay yes — 85% clears the bar for a support bot
Sign up free to answer and see why

Checkpoint

Building the eval suite for the assistant, a teammate says “let’s just collect 20,000 labelled tickets and score accuracy on all of them.” Better design?

AAgree — the larger the labelled set, the more trustworthy the evalBA portfolio: ~100 purpose-built regression examples (deterministic assertions) + 100+ sampled prod traces every 2–4 weeks (LLM-judge) + ~20 frontier cases, covered by dimension, with retrieval scored before generationCSkip offline evals and rely on production thumbs-up at scale
Sign up free to answer and see why

Checkpoint

In the case round, deflection (your NSM) is up 9% and leadership wants to declare success. What does a senior AI PM say?

AConfirm the 9% is statistically significant, then declare successBDeclare success — a rising NSM is the definition of a working featureCNot yet — check whether model quality diverged from the rising NSM via the held-out, per-slice faithfulness gate that runs independently of deflection
Sign up free to answer and see why

Checkpoint

Your launch-gate table includes the row “Rollback: documented runbook (owner: SRE).” Why does a senior PM mark this row not-passed?

AThe row should be owned by the PM, not SREBA documented-but-unrehearsed kill switch is a hypothesis, not a kill switch — require a staging rollback drill with a measured time-to-revert (target <15 min)CRunbooks are obsolete — rollback must be fully automated with zero documentation
Sign up free to answer and see why

Checkpoint

Live, the assistant confidently states a bereavement-fare refund policy that doesn’t exist, and one user screenshots it. How should the launch plan have handled this, and how does the incident process treat the screenshot?

AImmediately roll back the whole feature — any wrong customer-facing output is a SEV-1BGate-side: grounding + confidence-threshold human escalation + citation/policy-existence eval checks (ungrounded output is legally attributable — Air Canada). Incident-side: log the sample, check the failure-RATE trend, escalate only on a rate shift, and write the case back into the eval setCTreat it as inherent LLM behaviour and take no action — occasional wrong answers are acceptable
Sign up free to answer and see why

Interview prep

The capstone maps to the full AI-PM case round — “design the eval and launch plan for this AI feature,” the synthesis question that pulls on every earlier lesson. Graders reward one pattern across all of it: refuse the surface question, name a layered system, name the trade-off in concrete terms (cost, time, subgroup, owner), and pre-declare a divergence alarm or rollback trigger. Drive the conversation as the five steps above — success criteria → metric stack → eval suite → gates → rollout/incident — naming the artifact each step produces.
  1. 01“Design the eval and launch plan for this AI feature.” → distributional done → metric stack + divergence alarm → eval portfolio (3 sets / 3 modalities) → go/no-go table with owners → shadow→canary rollout + day-one drift monitoring + incident write-back.
  2. 02“The model is 85% accurate — do you ship?” → refuse the binary: scoped target rate, graceful tail, dangerous-failure floor, and a baseline the AI must beat.
  3. 03“How do you measure success?” → one metric per layer (business/product/model/system) plus a held-out, per-slice divergence alarm; never a single blended score.
  4. 04“Walk me through your eval harness.” → offline golden/regression (precision/recall/faithfulness) + online (task-completion/escalation/correction) + human calibration; name the ship gate.
  5. 05“How do you scale evaluation?” → ~80% deterministic CI / ~15% judge on sampled traffic / ~5% human; judge calibrated to a human golden set, recalibrated quarterly.
  6. 06“Build fairness/safety into the spec, not audit later.” → map NIST GenAI-profile risks to eval cases in the PRD; attach a one-page Risk Report with owners per row.
  7. 07“What’s your rollout and rollback?” → shadow→A/B→canary→full, gate full on the lower CI bound; rollback is a rehearsed, timed kill switch (<15 min), a P0 deliverable.
  8. 08“When is a bad output an incident?” → when the failure RATE shifts (e.g. hallucination up >10pp WoW), not one screenshot; log the sample and write it back into the eval set.
Push it deeper. Expect integration follow-ups that cross lessons: “deflection is up but a slice is being harmed — how do you know within 24 hours?” (per-slice quality dashboards + alerting and the held-out faithfulness gate, not a quarterly review). “The judge passes changes users dislike — what broke?” (an uncalibrated, drifted judge with position/verbosity bias — re-anchor to the human golden set and monitor agreement). “Legal says it’s fine, advocacy says it’s harmful — what do you do?” (escalate to the review board with halt power, document the residual risk in the Risk Report, make the call explicit not silent). “Ship fast AND safe — contradiction?” (no — trustworthy offline gates first, like Notion’s <24hr frontier deploys, so you don’t need a long risky A/B). In every case: name the metric, the slice, the gate, and the owner before proposing a fix.

Given any AI feature, could you whiteboard the full eval suite AND launch plan end to end — success criteria, metric stack, eval portfolio, go/no-go gates, and rollout/rollback/incident loop — and defend every decision under follow-up?

Not yetMostlyConfident

You can now

  • Open any AI-PM case by refusing the binary and writing a distributional definition of done with a baseline the AI must beat.
  • Lay a metric stack — business / model-quality / system — with a held-out, per-slice divergence alarm, never a single blended score.
  • Build the eval suite as a portfolio (regression + sampled-prod + frontier) across three modalities with a judge calibrated to a human golden set.
  • Assemble a go/no-go table where every row has a named owner and a binary criterion, with a rehearsed rollback and a one-page Risk Report.
  • Sequence shadow→A/B→canary→full, monitor quality/cost/latency/drift on day one, and fold every incident back into the eval set.
  • Defend the whole plan as one self-correcting loop — define, measure, gate, stage, monitor, write-back — tuned to the specific feature.

Track complete. You can now define success for a non-deterministic product, build the eval suite, and gate the launch like a senior AI PM — and defend all of it in the case round.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.