Lesson 6 of 8 · 58 min

Case type: AI / non-deterministic feature

Full AI agent-assist case: eval-first Eval Spec, HITL, cost/latency, launch gates; plus eval pipeline taxonomy beat.

Evals are the spec, not the QA phase

"Design an AI feature for [product]" is the fastest-growing case type at AI-native companies and increasingly at Big Tech. It is structurally different: success criteria for non-deterministic output, hallucination policy, cost/latency budgets, human-in-the-loop thresholds, and launch gates. If you design the UX first and mention evals as an afterthought, you fail the senior AI bar. Deep PRD templates live in AI Product Sense & PRDs and eval craft in related AI tracks — here we run a full case with a one-page Eval Spec.
Spot-it cue: AI, LLM, copilot, agent, autocomplete, smart reply, semantic search, "use GPT." First move: name the job and cost-of-error tail, define how you will know the model is good enough (eval set first), then UX, HITL, cost/latency, and gates. Companion track vocabulary welcome; do not re-derive CIRCLES.

How to solve AI feature cases

code
1AI FEATURE SPINE23  1. JOB + TAIL        user job; cost of wrong answer (annoy/cost/danger)4  2. WHY AI             why non-deterministic capability beats deterministic5  3. EVAL FIRST        golden set · rubric · slice hard cases · thresholds6  4. SYSTEM SHAPE      deterministic tools + model + validators (not "chat")7  5. UX FOR UNCERTAINTY confidence, sources, regenerate, edit, refuse8  6. HITL              which actions need human approve / review / audit9  7. COST / LATENCY    budget per call; model routing; cache; streaming10  8. LAUNCH GATES      offline bar → limited online → GA criteria11  9. ABUSE / SAFETY    prompt injection, data leakage, misuse12 10. METRICS           product north star + model quality + guardrails1314  Anti-patterns: eval as QA after ship; ignore $/call; "just use the biggest model."
Five ways AI cases differ from classic product cases: (1) success is a distribution, not a binary screenshot; (2) hallucination and calibration are first-class; (3) latency and token cost gate designs; (4) HITL thresholds differ by action class; (5) evals are the spec.
code
1COST-OF-ERROR TAIL → DESIGN CONSEQUENCES23  Tail               Autonomy default     Eval rigor        HITL4  -----------------  -------------------  ----------------  ---------------5  Annoying           more auto OK         lighter golden    spot checks6  Costly (money)     assist-first         hard slices       approve sends7  Dangerous/legal    refuse / escalate    red-team suite    human owns act89  Name the row in minute 1. Everything else hangs from it.
Do not re-teach full PRD templates here — if you need acceptance-criteria depth, use the companion AI Product Sense & PRDs track. This lesson forces the case shape: eval → system → gates, spoken under time pressure.

Primary case prompt

Prompt: Design an AI feature for a B2B customer support product (think Intercom/Zendesk-class): an agent-assist that drafts replies for human support agents using help center + past tickets. Cover quality, safety, cost, and how you'd launch.

Worked strong answer

Job + tail. Job: when an agent handles a ticket, draft a reply that is accurate to policy and account context so handle time drops without raising wrong-answer rate. Cost-of-error: medium-high — wrong refund advice or security guidance is costly (margin, trust, compliance), not merely annoying. Therefore: assist, not autonomous send, in v1 for money/security intents.
code
1WHY AI + SYSTEM SHAPE23  Why AI: language generation over heterogeneous tickets beats rigid macros4  for long-tail intents; retrieval can ground in approved content.56  System shape (deterministic shell):7    1) Retrieve: help center chunks + similar resolved tickets + account fields8       (ACL-filtered)9    2) Policy tags: intent classifier (refund, auth, outage, general)10    3) Generate: draft with citations to retrieved sources11    4) Validate: regex/policy checkers (never invent order IDs; block banned12       phrases; PII scrub)13    5) HITL: agent edits + explicit Send; auto-send only later on low-risk intents14       that pass confidence + eval gates1516  Non-goal v1: fully autonomous agent emailing customers on all intents.

Eval Spec (one-pager you should speak aloud)

code
1EVAL SPEC — Agent Assist Drafts v123  Task: given ticket + retrieved context, produce draft reply + citations45  Golden set: 300 tickets stratified by intent (refund, login, billing,6  outage, abusive, multilingual), including adversarial (prompt injection in7  ticket body) and missing-context cases.89  Rubric dimensions (1–5 each, human + LLM-as-judge with calibration):10    · Factual groundedness (supported by sources?)11    · Policy compliance (refund/security rules)12    · Tone / empathy fit for brand13    · Actionability (clear next step)14    · Safety (no leakage of other customers' data)1516  Hard slices must pass, not just average:17    · Refund eligibility edge cases18    · "Ignore previous instructions" in ticket text19    · Empty retrieval / conflicting sources → must refuse or ask2021  Offline launch bar (example):22    · Groundedness ≥4.5 on 90% of golden; 0 critical policy fails23    · Injection suite: 100% block or safe refuse24    · p95 draft latency ≤ 4s with streaming first tokens ≤ 1s25    · Cost ≤ $0.03 / draft at projected volume (or routed model mix)2627  Online gates:28    · Agent edit distance / accept rate29    · Reopen rate and CSAT on assisted vs control30    · Escape hatches used (refuse, escalate)31    · $ and latency budgets hold in production traffic3233  Drift: weekly golden regression; alert on groundedness or cost spikes
HITL thresholds by action class. Draft-only: default. Suggest macro: medium risk. Auto-send: only "general FAQ" intents with high groundedness, no account mutation, and after online gates clear. Never auto-send on refunds, legal, security, or vulnerability reports in v1.
code
1COST / LATENCY / ROUTING23  Budget: p95 ≤4s end-to-end; ≤$0.03/draft blended4  Levers:5    · Cache retrieval for repeated KB queries6    · Small model for intent + large model only for hard intents7    · Truncate context with reranked top-k, not whole ticket history8    · Stream tokens to agents (perceived latency)9  If cost spikes: degrade to retrieval-only suggested snippets (no freeform)1011  UX FOR UNCERTAINTY:12    · Citations inline; click-to-source13    · "Low confidence" badge when retrieval weak → force agent attention14    · Regenerate with constraints (shorter, more formal)15    · One-click "I don't know / escalate" template
Launch sequence. (1) Offline bar on golden + injection suite. (2) Dogfood on internal support. (3) 10% agents, assist-only, no auto-send. (4) Expand agents; enable auto-send only on one low-risk intent with kill switch. (5) GA when online CSAT/reopen/cost gates hold two weeks. Kill switch: disable generation globally; fall back to macros.
Product metrics. North star: median handle time on assisted tickets without CSAT drop. Drivers: draft accept rate, retrieval hit rate. Guardrails: reopen rate, policy violations caught in audit, cost per assisted ticket, agent trust survey. XFN: support ops (policy), security (injection/PII), ML/platform (eval harness), finance (inference COGS).

Companion beat: evaluation pipeline question

Sometimes the whole prompt is: "Build an evaluation pipeline for this LLM feature." Answer with taxonomy, not tools trivia.
code
1EVAL PIPELINE TAXONOMY23  Offline: golden sets, rubric scoring, regression CI on prompts/models4  LLM-as-judge: cheap scale; must calibrate to humans on a panel5  Human: hard slices, safety, periodic recalibration6  Online: A/B on product metrics; shadow mode; feedback (thumbs, edits)7  Ops: drift monitors, cost dashboards, incident replay into golden set89  Anti-pattern: 'We'll just use GPT-4 to grade everything' with no human10  calibration and no hard-case coverage.

Weak vs strong

code
1WEAK2  'Chatbot in the corner, use GPT-4, fine-tune later, measure CSAT. Evals3   before launch if we have time.'45STRONG6  Tail-priced job → eval spec with hard slices → deterministic shell →7  HITL by action class → cost/latency budgets → staged launch gates →8  kill switch → product + model metrics.

Interview ways (AI feature)

  1. 01"Design an AI feature for X." → Job + cost-of-error tail, eval first, system shape, UX for uncertainty, HITL, cost/latency, launch gates.
  2. 02"How do you know it's good enough to ship?" → Offline thresholds on golden hard slices + online product gates + kill switch.
  3. 03"What about hallucinations?" → Grounding, validators, refuse when retrieval empty, citations, audit sampling — not "temperature 0" alone.
  4. 04"Why not auto-send everything?" → Action-class HITL; money/security intents stay human in v1.
  5. 05"Cost seems high." → Routing, cache, top-k context, degrade path to non-generative assist.
  6. 06"Build the eval pipeline." → Offline / judge / human / online / drift — with calibration and hard cases.

Full spoken senior answer (~13 minutes)

AI cases need slightly more airtime for the Eval Spec. Speak the spec as if it were the PRD — because at senior bar, it is.
code
1SPOKEN SENIOR ANSWER — Support agent-assist drafts (~13 min)23  [0:00–1:30 JOB + TAIL]4  "AI feature case. Job: when an agent handles a ticket, draft a reply accurate5  to policy and account context so handle time drops without raising wrong-6  answer rate. Cost-of-error is medium-high: wrong refund or security guidance7  hits margin, trust, and compliance — not merely annoying. Therefore v1 is8  assist, not autonomous send, for money and security intents. Before UI, I9  define eval and cost-of-error because those decide architecture."1011  [1:30–4:00 WHY AI + SYSTEM SHAPE]12  "Why AI: long-tail language over heterogeneous tickets beats rigid macros;13  retrieval can ground in approved content. System shape is a deterministic14  shell: retrieve help-center chunks, similar resolved tickets, account fields15  with ACL filters; classify intent; generate draft with citations; validate16  with policy checkers — never invent order IDs, block banned phrases, scrub17  PII; human edits and explicit Send. Non-goal v1: fully autonomous email on18  all intents."1920  [4:00–8:00 EVAL SPEC — SPEAK THE ONE-PAGER]21  "Eval Spec for Agent Assist Drafts v1. Task: given ticket and retrieved22  context, produce draft plus citations. Golden set: three hundred tickets23  stratified by intent — refund, login, billing, outage, abusive, multilingual24  — plus adversarial prompt injection in ticket body and missing-context cases.25  Rubric one to five: factual groundedness, policy compliance, tone, actionability,26  safety against cross-customer leakage. Hard slices must pass, not just averages:27  refund edge cases, ignore-previous-instructions attacks, empty retrieval must28  refuse or ask. Offline bar examples: groundedness at least four point five on29  ninety percent of golden; zero critical policy fails; injection suite one30  hundred percent block or safe refuse; p95 draft latency under four seconds with31  first tokens under one second streaming; cost at or under three cents per draft32  blended. Online gates: agent edit distance and accept rate; reopen and CSAT33  assisted versus control; refuse and escalate rates; dollar and latency budgets34  in production. Drift: weekly golden regression; alert on groundedness or cost35  spikes. I will not ship on demo vibes or mean score alone."3637  [8:00–11:00 HITL + COST + LAUNCH]38  "HITL by action class: draft-only default; auto-send only later on low-risk FAQ39  intents with high groundedness and no account mutation after online gates;40  never auto-send refunds, legal, security, vulnerability reports in v1. Cost and41  latency: cache retrieval, small model for intent, large model for hard intents,42  top-k context, stream tokens; degrade to retrieval snippets if generative path43  blows budget. Launch: offline bar; dogfood; ten percent agents assist-only;44  expand; auto-send one low-risk intent with kill switch; GA when CSAT, reopen,45  and cost hold two weeks. Kill switch disables generation globally; macros46  remain."4748  [11:00–13:00 METRICS + CLOSE]49  "North star: median handle time on assisted tickets without CSAT drop. Drivers:50  draft accept rate, retrieval hit rate. Guardrails: reopen, policy audit fails,51  cost per assisted ticket, agent trust pulse. XFN: support ops, security, ML52  platform eval harness, finance for inference COGS. Close: eval is the spec;53  hard slices decide; autonomy is earned by action class, not slogans."

Eval Spec depth — what interviewers probe

code
1EVAL SPEC DEEP DIVE (have these ready under pushback)23  GOLDEN SET DESIGN4  · Stratify by intent frequency AND harm severity (not only popular intents)5  · Include: empty retrieval, conflicting sources, multilingual, adversarial6  · Version the set; every production incident can graduate a case into golden7  · Size: start ~100 high-quality labeled; grow to 300+ as slices appear89  RUBRIC + JUDGES10  · Human labels on a panel for calibration; LLM-as-judge only after agreement11  · Separate critical fails (policy/safety) from quality scores (tone)12  · Critical fails are ship blockers even if mean quality is high1314  HEURISTIC GRADERS (cheap CI)15  · Schema: citations present; JSON fields valid if structured mode16  · Regex: banned phrases; fabricated ID patterns; PII leakage patterns17  · Retrieval: citation IDs must exist in retrieved set1819  MODEL-JUDGE20  · Use for open-ended quality; re-calibrate monthly vs humans on 5% sample21  · Never sole gate for safety or refund policy2223  ONLINE + DRIFT24  · Shadow mode before enforce; edit-distance as proxy for draft quality25  · Cost and latency SLOs with pages — AI features die quietly on COGS26  · Incident replay: bad ticket → golden set within 48h2728  ANTI-PATTERNS29  · 'We'll use GPT-4 to grade everything' with no human calibration30  · Average score only; no hard slices31  · Eval after launch as QA32  · No kill switch / no degrade path
Primary sources worth skimming before AI-native loops: Anthropic's demystifying evals for agents, OpenAI working with evals, and Hamel Husain's evals writing. You do not need to recite papers — you need to sound like someone who has shipped under a bar.

Follow-up pressure (AI feature)

code
1PRESSURE Q → SENIOR REPLY23  Q: "Just use the biggest model."4  A: Biggest model raises cost and sometimes latency without fixing retrieval5     holes or policy. Route by intent; eval decides, not brand names.67  Q: "When do we auto-send?"8  A: After offline and online gates, one low-risk intent, kill switch live.9     Refunds and security stay human in v1.1011  Q: "Hallucinations will happen — so what?"12  A: Cost-of-error row decides. For money/security, we refuse empty retrieval,13     validate, cite, audit sample — not shrug.1415  Q: "Eval set will take months."16  A: Start with one hundred labeled hard cases and heuristic graders this17     sprint. Perfect is the enemy; zero is disqualifying.1819  Q: "Agents will game accept rate."20  A: Pair accept rate with reopen, CSAT, and audit fails. Accept alone is21     gameable.2223  Q: "Prompt injection in ticket body — realistic?"24  A: Yes. Treat untrusted text as hostile; retrieval and system prompts must25     not obey ticket instructions; suite must be one hundred percent safe.
How to ACE AI Product Sense Interviews (OpenAI PM Mock)ExponentarticleAnthropic — Demystifying evals for AI agentsAnthropic EngineeringarticleOpenAI — Working with evalsOpenAIarticleAnthropic — Building effective agentsAnthropic EngineeringarticleAI Product Manager interview questions (2026)ExponentarticleAnthropic PM interview (process & prep)IGotAnOfferarticleAI product sense is not the same as product senseAakash GuptaarticleHow to build AI product senseTal Raviv & Aman Khan

Checkpoint

You are asked to design AI drafts for support agents. First structural move?

ASketch the chat UI and pick GPT-4 so quality is maximizedBName the cost-of-error tail (wrong refund/security advice is costly) and define an eval/golden-set bar before committing to autonomy level or UXCPropose fully autonomous email sending on day one to maximize ROI story
Sign up free to answer and see why

Checkpoint

Which offline result should fail a launch gate even if average rubric scores look fine?

ASlightly slower p50 latency still under budgetBCritical policy failures or injection-suite misses on hard slices, even when mean groundedness is highCAgents say they "like" the feature in a hallway conversation
Sign up free to answer and see why

Checkpoint

Best guardrail metric for hallucinations in agent-assist?

ANumber of tokens generated per dayBGroundedness/policy audit fail rate and reopen rate on assisted tickets vs controlCModel parameter count
Sign up free to answer and see why

Checkpoint

Cost per draft doubles after a prompt change. Senior response?

AIgnore cost if quality is up — finance can deal with it laterBRevisit routing/context size, set a hard $/draft budget, and degrade to retrieval snippets if generative path cannot meet budget at required qualityCAlways switch to the largest model to improve caching
Sign up free to answer and see why

Checkpoint

When is auto-send appropriate in this product?

AImmediately for all intents to hit efficiency goalsBOnly after offline+online gates, limited to low-risk intents with strong groundedness, with a kill switch — refunds/security stay humanCNever, because AI must never send email
Sign up free to answer and see why

How ready are you to run an AI feature case with an Eval Spec and launch gates?

New to itGetting thereConfident

AI feature locked

  • Eval set first; averages are not enough — hard slices decide.
  • Deterministic shell + model + validators; UX for uncertainty.
  • HITL by action class; cost/latency budgets with degrade paths.
  • Launch gates offline → dogfood → limited online → GA; kill switch.

Next: Strategy / market entry — beachhead, asymmetric bet, kill criteria.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.