Lesson 6 of 8 · 58 min
Case type: AI / non-deterministic feature
Full AI agent-assist case: eval-first Eval Spec, HITL, cost/latency, launch gates; plus eval pipeline taxonomy beat.
Evals are the spec, not the QA phase
How to solve AI feature cases
1AI FEATURE SPINE23 1. JOB + TAIL user job; cost of wrong answer (annoy/cost/danger)4 2. WHY AI why non-deterministic capability beats deterministic5 3. EVAL FIRST golden set · rubric · slice hard cases · thresholds6 4. SYSTEM SHAPE deterministic tools + model + validators (not "chat")7 5. UX FOR UNCERTAINTY confidence, sources, regenerate, edit, refuse8 6. HITL which actions need human approve / review / audit9 7. COST / LATENCY budget per call; model routing; cache; streaming10 8. LAUNCH GATES offline bar → limited online → GA criteria11 9. ABUSE / SAFETY prompt injection, data leakage, misuse12 10. METRICS product north star + model quality + guardrails1314 Anti-patterns: eval as QA after ship; ignore $/call; "just use the biggest model."1COST-OF-ERROR TAIL → DESIGN CONSEQUENCES23 Tail Autonomy default Eval rigor HITL4 ----------------- ------------------- ---------------- ---------------5 Annoying more auto OK lighter golden spot checks6 Costly (money) assist-first hard slices approve sends7 Dangerous/legal refuse / escalate red-team suite human owns act89 Name the row in minute 1. Everything else hangs from it.Key idea
Primary case prompt
Worked strong answer
1WHY AI + SYSTEM SHAPE23 Why AI: language generation over heterogeneous tickets beats rigid macros4 for long-tail intents; retrieval can ground in approved content.56 System shape (deterministic shell):7 1) Retrieve: help center chunks + similar resolved tickets + account fields8 (ACL-filtered)9 2) Policy tags: intent classifier (refund, auth, outage, general)10 3) Generate: draft with citations to retrieved sources11 4) Validate: regex/policy checkers (never invent order IDs; block banned12 phrases; PII scrub)13 5) HITL: agent edits + explicit Send; auto-send only later on low-risk intents14 that pass confidence + eval gates1516 Non-goal v1: fully autonomous agent emailing customers on all intents.Eval Spec (one-pager you should speak aloud)
1EVAL SPEC — Agent Assist Drafts v123 Task: given ticket + retrieved context, produce draft reply + citations45 Golden set: 300 tickets stratified by intent (refund, login, billing,6 outage, abusive, multilingual), including adversarial (prompt injection in7 ticket body) and missing-context cases.89 Rubric dimensions (1–5 each, human + LLM-as-judge with calibration):10 · Factual groundedness (supported by sources?)11 · Policy compliance (refund/security rules)12 · Tone / empathy fit for brand13 · Actionability (clear next step)14 · Safety (no leakage of other customers' data)1516 Hard slices must pass, not just average:17 · Refund eligibility edge cases18 · "Ignore previous instructions" in ticket text19 · Empty retrieval / conflicting sources → must refuse or ask2021 Offline launch bar (example):22 · Groundedness ≥4.5 on 90% of golden; 0 critical policy fails23 · Injection suite: 100% block or safe refuse24 · p95 draft latency ≤ 4s with streaming first tokens ≤ 1s25 · Cost ≤ $0.03 / draft at projected volume (or routed model mix)2627 Online gates:28 · Agent edit distance / accept rate29 · Reopen rate and CSAT on assisted vs control30 · Escape hatches used (refuse, escalate)31 · $ and latency budgets hold in production traffic3233 Drift: weekly golden regression; alert on groundedness or cost spikes1COST / LATENCY / ROUTING23 Budget: p95 ≤4s end-to-end; ≤$0.03/draft blended4 Levers:5 · Cache retrieval for repeated KB queries6 · Small model for intent + large model only for hard intents7 · Truncate context with reranked top-k, not whole ticket history8 · Stream tokens to agents (perceived latency)9 If cost spikes: degrade to retrieval-only suggested snippets (no freeform)1011 UX FOR UNCERTAINTY:12 · Citations inline; click-to-source13 · "Low confidence" badge when retrieval weak → force agent attention14 · Regenerate with constraints (shorter, more formal)15 · One-click "I don't know / escalate" templateKey idea
Companion beat: evaluation pipeline question
1EVAL PIPELINE TAXONOMY23 Offline: golden sets, rubric scoring, regression CI on prompts/models4 LLM-as-judge: cheap scale; must calibrate to humans on a panel5 Human: hard slices, safety, periodic recalibration6 Online: A/B on product metrics; shadow mode; feedback (thumbs, edits)7 Ops: drift monitors, cost dashboards, incident replay into golden set89 Anti-pattern: 'We'll just use GPT-4 to grade everything' with no human10 calibration and no hard-case coverage.Weak vs strong
1WEAK2 'Chatbot in the corner, use GPT-4, fine-tune later, measure CSAT. Evals3 before launch if we have time.'45STRONG6 Tail-priced job → eval spec with hard slices → deterministic shell →7 HITL by action class → cost/latency budgets → staged launch gates →8 kill switch → product + model metrics.Common mistake
If the model is frontier-quality, we can skip a custom eval set.
Interview ways (AI feature)
- 01"Design an AI feature for X." → Job + cost-of-error tail, eval first, system shape, UX for uncertainty, HITL, cost/latency, launch gates.
- 02"How do you know it's good enough to ship?" → Offline thresholds on golden hard slices + online product gates + kill switch.
- 03"What about hallucinations?" → Grounding, validators, refuse when retrieval empty, citations, audit sampling — not "temperature 0" alone.
- 04"Why not auto-send everything?" → Action-class HITL; money/security intents stay human in v1.
- 05"Cost seems high." → Routing, cache, top-k context, degrade path to non-generative assist.
- 06"Build the eval pipeline." → Offline / judge / human / online / drift — with calibration and hard cases.
Full spoken senior answer (~13 minutes)
1SPOKEN SENIOR ANSWER — Support agent-assist drafts (~13 min)23 [0:00–1:30 JOB + TAIL]4 "AI feature case. Job: when an agent handles a ticket, draft a reply accurate5 to policy and account context so handle time drops without raising wrong-6 answer rate. Cost-of-error is medium-high: wrong refund or security guidance7 hits margin, trust, and compliance — not merely annoying. Therefore v1 is8 assist, not autonomous send, for money and security intents. Before UI, I9 define eval and cost-of-error because those decide architecture."1011 [1:30–4:00 WHY AI + SYSTEM SHAPE]12 "Why AI: long-tail language over heterogeneous tickets beats rigid macros;13 retrieval can ground in approved content. System shape is a deterministic14 shell: retrieve help-center chunks, similar resolved tickets, account fields15 with ACL filters; classify intent; generate draft with citations; validate16 with policy checkers — never invent order IDs, block banned phrases, scrub17 PII; human edits and explicit Send. Non-goal v1: fully autonomous email on18 all intents."1920 [4:00–8:00 EVAL SPEC — SPEAK THE ONE-PAGER]21 "Eval Spec for Agent Assist Drafts v1. Task: given ticket and retrieved22 context, produce draft plus citations. Golden set: three hundred tickets23 stratified by intent — refund, login, billing, outage, abusive, multilingual24 — plus adversarial prompt injection in ticket body and missing-context cases.25 Rubric one to five: factual groundedness, policy compliance, tone, actionability,26 safety against cross-customer leakage. Hard slices must pass, not just averages:27 refund edge cases, ignore-previous-instructions attacks, empty retrieval must28 refuse or ask. Offline bar examples: groundedness at least four point five on29 ninety percent of golden; zero critical policy fails; injection suite one30 hundred percent block or safe refuse; p95 draft latency under four seconds with31 first tokens under one second streaming; cost at or under three cents per draft32 blended. Online gates: agent edit distance and accept rate; reopen and CSAT33 assisted versus control; refuse and escalate rates; dollar and latency budgets34 in production. Drift: weekly golden regression; alert on groundedness or cost35 spikes. I will not ship on demo vibes or mean score alone."3637 [8:00–11:00 HITL + COST + LAUNCH]38 "HITL by action class: draft-only default; auto-send only later on low-risk FAQ39 intents with high groundedness and no account mutation after online gates;40 never auto-send refunds, legal, security, vulnerability reports in v1. Cost and41 latency: cache retrieval, small model for intent, large model for hard intents,42 top-k context, stream tokens; degrade to retrieval snippets if generative path43 blows budget. Launch: offline bar; dogfood; ten percent agents assist-only;44 expand; auto-send one low-risk intent with kill switch; GA when CSAT, reopen,45 and cost hold two weeks. Kill switch disables generation globally; macros46 remain."4748 [11:00–13:00 METRICS + CLOSE]49 "North star: median handle time on assisted tickets without CSAT drop. Drivers:50 draft accept rate, retrieval hit rate. Guardrails: reopen, policy audit fails,51 cost per assisted ticket, agent trust pulse. XFN: support ops, security, ML52 platform eval harness, finance for inference COGS. Close: eval is the spec;53 hard slices decide; autonomy is earned by action class, not slogans."Eval Spec depth — what interviewers probe
1EVAL SPEC DEEP DIVE (have these ready under pushback)23 GOLDEN SET DESIGN4 · Stratify by intent frequency AND harm severity (not only popular intents)5 · Include: empty retrieval, conflicting sources, multilingual, adversarial6 · Version the set; every production incident can graduate a case into golden7 · Size: start ~100 high-quality labeled; grow to 300+ as slices appear89 RUBRIC + JUDGES10 · Human labels on a panel for calibration; LLM-as-judge only after agreement11 · Separate critical fails (policy/safety) from quality scores (tone)12 · Critical fails are ship blockers even if mean quality is high1314 HEURISTIC GRADERS (cheap CI)15 · Schema: citations present; JSON fields valid if structured mode16 · Regex: banned phrases; fabricated ID patterns; PII leakage patterns17 · Retrieval: citation IDs must exist in retrieved set1819 MODEL-JUDGE20 · Use for open-ended quality; re-calibrate monthly vs humans on 5% sample21 · Never sole gate for safety or refund policy2223 ONLINE + DRIFT24 · Shadow mode before enforce; edit-distance as proxy for draft quality25 · Cost and latency SLOs with pages — AI features die quietly on COGS26 · Incident replay: bad ticket → golden set within 48h2728 ANTI-PATTERNS29 · 'We'll use GPT-4 to grade everything' with no human calibration30 · Average score only; no hard slices31 · Eval after launch as QA32 · No kill switch / no degrade pathFollow-up pressure (AI feature)
1PRESSURE Q → SENIOR REPLY23 Q: "Just use the biggest model."4 A: Biggest model raises cost and sometimes latency without fixing retrieval5 holes or policy. Route by intent; eval decides, not brand names.67 Q: "When do we auto-send?"8 A: After offline and online gates, one low-risk intent, kill switch live.9 Refunds and security stay human in v1.1011 Q: "Hallucinations will happen — so what?"12 A: Cost-of-error row decides. For money/security, we refuse empty retrieval,13 validate, cite, audit sample — not shrug.1415 Q: "Eval set will take months."16 A: Start with one hundred labeled hard cases and heuristic graders this17 sprint. Perfect is the enemy; zero is disqualifying.1819 Q: "Agents will game accept rate."20 A: Pair accept rate with reopen, CSAT, and audit fails. Accept alone is21 gameable.2223 Q: "Prompt injection in ticket body — realistic?"24 A: Yes. Treat untrusted text as hostile; retrieval and system prompts must25 not obey ticket instructions; suite must be one hundred percent safe.
How to ACE AI Product Sense Interviews (OpenAI PM Mock)ExponentarticleAnthropic — Demystifying evals for AI agentsAnthropic EngineeringarticleOpenAI — Working with evalsOpenAIarticleAnthropic — Building effective agentsAnthropic EngineeringarticleAI Product Manager interview questions (2026)ExponentarticleAnthropic PM interview (process & prep)IGotAnOfferarticleAI product sense is not the same as product senseAakash GuptaarticleHow to build AI product senseTal Raviv & Aman KhanCheckpoint
You are asked to design AI drafts for support agents. First structural move?
Checkpoint
Which offline result should fail a launch gate even if average rubric scores look fine?
Checkpoint
Best guardrail metric for hallucinations in agent-assist?
Checkpoint
Cost per draft doubles after a prompt change. Senior response?
Checkpoint
When is auto-send appropriate in this product?
How ready are you to run an AI feature case with an Eval Spec and launch gates?
AI feature locked
- Eval set first; averages are not enough — hard slices decide.
- Deterministic shell + model + validators; UX for uncertainty.
- HITL by action class; cost/latency budgets with degrade paths.
- Launch gates offline → dogfood → limited online → GA; kill switch.
Next: Strategy / market entry — beachhead, asymmetric bet, kill criteria.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.