Lesson 6 of 6 · 47 min

Capstone: PRD for an agentic assistant

Assemble the whole track into one artifact: run the product-sense case for an agentic personal-shopping assistant, then write its PRD — goals/non-goals, the agent architecture choice, eval set, guardrails, failure modes, and the exec narrative — rehearsing the decisions an interviewer (or an incident) will push on.

Capstone

Run the case, then write the PRD

Time to assemble it: take a prompt that shows up verbatim in 2026 rounds — “outline a PRD for an agentic personal-shopping assistant” — and run it end to end. First the product-sense case (six lenses, wedge, segment, failure taxonomy, ship threshold), then the PRD (the 12-section structure, with the AI-specific cluster doing the heavy lifting), then the exec narrative around the bet. Every move below is a callback to a specific lesson. Treat this as the interview itself: a PRD round for an agentic assistant tests four skills — problem framing, architecture clarity, evaluation rigor, and risk/trust — and the difference between a mid and a senior signal is structure.
The single best opening move, exactly as in a RAG system-design round, is to ask the scoping questions before designing anything — it shows you know an “agent” hides a dozen different systems with different latency, cost, and liability profiles. Lock scope first.
  1. 01User & job? — who is this for, and what single job-to-be-done does the agent own (not “everyone who shops”)? (sets the wedge + segment).
  2. 02Autonomy? — does it suggest, draft a cart, or actually purchase? (sets the cost-of-error tail and the human-in-the-loop gate).
  3. 03Stakes? — is a wrong action embarrassing (bad recommendation) or costly (wrong purchase, charged card)? (sets eval rigor + guardrails).
  4. 04Architecture? — single-purpose tool, or orchestrated multi-agent? (Anthropic: start simple; add agency only when simpler solutions fail).
  5. 05Economics & latency? — cost per session, p95 budget? (sets model routing and whether you can afford a reasoning model in the loop).

The product-sense pass: wedge, segment, failure taxonomy

Run the six lenses (Lesson 1). User Reality: a busy gift-buyer who knows the occasion but not the product. Model Behavior: reliable at constrained recommendation + comparison, not reliable at unsupervised checkout. Economics: route a cheap model for browsing, a stronger one only for the final shortlist (Intercom’s pattern). System Design: the LLM sits inside a deterministic pipeline — catalog search and price/availability come from tools, not the model’s memory. Trust & Liability: the tail is a wrong purchase, so autonomy stops at “draft the cart; human confirms.” Go-to-Market: launch to a prosumer/beta segment that tolerates probabilistic quality. The wedge: not “an AI that shops,” but “an assistant that grounds every recommendation in live catalog data with a human-confirmed checkout” — trust and fit as the moat, not raw model quality.
Then the failure taxonomy (Lesson 2): an input error (vague request — “a gift for my dad”) → ask a clarifying question; a system error (the model invents a product that doesn’t exist) → ground in tool results, refuse/abstain when the catalog has no match; a context error (price or stock changed) → re-fetch at action time, never trust a cached price. Autonomy stops short of the irreversible step — Tome’s outline-gate pattern generalizes to “confirm the cart before purchase”: convert an open-ended agentic task into an iterative contract the user accepts. Interview angle. Naming the human-confirmation gate before drawing the agent loop is the move that signals you’ve shipped agents — it’s the single most common thing weak candidates skip.

The PRD skeleton — AI-specific cluster does the work

Now the artifact. Outline the 12 sections (Lesson 4), but spend your tokens on the AI-specific cluster. Anthropic’s rubric anchors the architecture: start with the simplest primitive (a single-purpose tool), add multi-step agency only when simpler solutions fail — so the PRD names the failure mode at which orchestration becomes necessary, rather than defaulting to multi-agent.
code
1AGENTIC PERSONAL-SHOPPER PRD -- the skeleton, with the lesson behind each section23  STRATEGIC FOUNDATION4    TL;DR ......................... 3-line bet: who, what job, what outcome      L55    Problem & Customer ........... gift-buyer; "knows occasion, not product"     L16    Success Metrics .............. input: cart-confirm rate, edit rate;          L37                                   output: completed-purchase rate, return rate8    Scope / Non-goals ............ NO unsupervised checkout in v1 (explicit)     L1/L39  AI-SPECIFIC LAYER10    Eval Framework ............... 2k labeled sessions; offline gold + online    L411    Quality Bar (numeric) ........ >=90% relevant shortlist; <2% hallucinated    L412                                   product; 98% precision on policy-violation13    Guardrails ................... input: PII, off-topic; output: no fake SKUs,  L4/L214                                   price grounded to tool; ACTION gate = human15    Model & Data Strategy ........ route cheap model browse / strong model       L316                                   shortlist; catalog + price via tools, not LLM17    Responsible AI ............... no dark-pattern upsell; disclose AI; bias on  L218                                   brand/price recommendations19  OPERATIONAL PLAN20    Monitoring & Drift ........... daily eval sweep; alert on hallucination-rate L421    Failure Modes & Fallback ..... input/system/context recovery per failure    L222    GTM & Adoption ............... prosumer beta; 15% weekly-active by 180d      L3/L523  LIVING EPILOGUE24    Decision log ................. Date | Decision | Reason | Reversible?         L52526  Architecture (Anthropic): single-purpose recommend+compare tool FIRST;27  multi-agent only if eval fails on multi-step tasks. Human confirms the buy.
The acceptance criteria block (Lesson 4) makes the quality bar a contract: “On a 2,000-session gold set: shortlist relevance >=90% (LLM-as-judge, calibrated quarterly against human review), hallucinated-product rate <2%, policy-violation precision 98% at the action gate; p95 session latency <3s; cost/session <$0.08; fallback path 100% pass; refreshed each catalog or model update.” Notice it carries the input/output metric split, a guardrail bar, an operational envelope, and a cadence — no adjectives. Interview angle. “Describe the success metrics, eval set, and guardrails for this agent” is a verbatim 2026 prompt — answer with this block, not “we’ll measure accuracy.”

The debug round: when the agent is wrong in production

A PRD round often pivots to an incident: “three weeks after launch, recommendation quality dropped 15% — debug it.” The senior move is to name the diagnostic ladder out loud and budget time on each rung the way an on-call PM would, rather than blaming “the model.” Walk it from cheapest to deepest, and state your stakeholder cadence (updates every ~2 hours) as you go — that incident-reflex is exactly what the round tests.
code
1THE DEBUG LADDER -- ascending from observation to root cause23  Order  Suspect              Diagnostic action4  -----  -------------------  ------------------------------------------------5  1      Measurement error    replay the eval pipeline; check labeler/judge drift6  2      Data drift           compare input distribution last 7d vs prior 30d7  3      Prompt/policy change  diff the system prompt + tool defs between versions8  4      Model version change  check deployment IDs; is a rollback available?9  5      Rate limit / quota    inspect token/sec caps; are tool calls throttling?10  6      User-behavior shift   segment new vs returning; did a campaign change mix?11  7      Infrastructure        tool/vector latency, embedding service health1213  Weak: blame "the model." Strong: name the ladder, log time per rung,14  update stakeholders every ~2h (1h surface, 4h root cause, 24h fix).
The single likeliest culprit, from Lesson 1’s case study, is a silent provider-side change (you shipped nothing, yet quality fell) — the safeguard is to pin model and prompt versions and gate every change on an eval. The fix in scope: roll back to the pinned version, re-run the gold set, and add the failure cluster as a permanent regression case. Interview angle. “How do you prevent this next month?” wants a named guardrail (a hallucination-rate kill-switch on the ramp, a pinned version, a drift alarm with a threshold), not a proposal to “monitor more.”

Generalize the pattern — the same skeleton, a new agent

The point of the capstone is the reusable shape, not the shopping example. Swap in a trip-planning agent and the skeleton holds: the wedge is grounding (real flight/hotel inventory via tools, not the model’s memory); the action gate is the booking step (draft the itinerary, human confirms before charging); the failure taxonomy is identical (input: vague “somewhere warm in March” → clarify; system: invented a hotel → ground + abstain; context: price/availability changed → re-fetch at booking). Swap in an enterprise coding agent migrating COBOL and it still holds: the gate is “open a PR, a human reviews before merge,” the eval set is labeled migration cases with a correctness bar, and autonomy stops short of pushing to production. Interview angle. When you’re handed any “PRD for an agent” prompt, run this exact pattern — scope → six lenses → 12-section PRD with the action gate → debug ladder — and you have a senior answer for an agent you’ve never seen before.
Start with simple prompts, optimize them with comprehensive evaluation, and add multi-step agentic systems only when simpler solutions fall short. — Anthropic, Building Effective AI Agents

Interview prep

The capstone round is the whole track under time pressure. Across every prompt type the most consistent rubric signal is the same: name a specific metric, a specific failure mode, and a specific user. If your answer carries all three, you outscore candidates who don’t — regardless of the round.
  1. 01“Outline a PRD for an agentic shopping assistant.” → scope first, then the 12 sections; spend tokens on eval + guardrails + the action gate, not the problem statement.
  2. 02“What’s the agent architecture?” → simplest primitive first (single-purpose recommend tool); multi-agent only when eval fails on multi-step tasks (Anthropic).
  3. 03“Success metrics, eval set, guardrails?” → input (cart-confirm rate) + output (completed-purchase rate) + a 2k-session gold set + an action-gate policy bar — not “accuracy.”
  4. 04“Where’s the human in the loop?” → at the irreversible step: draft the cart, human confirms the purchase (Tome’s outline-gate generalized).
  5. 05“The agent recommended a product that doesn’t exist — why, and the fix?” → system error; ground in tool results, abstain on no-match, add a fake-SKU output guardrail.
  6. 06“Quality dropped 15% post-launch — debug it.” → name the ladder (measurement → drift → prompt → model → rate-limit → behavior → infra); most likely a silent model update.
  7. 07“How do you prevent recurrence?” → pin model + prompt versions, eval gate on every change, a hallucination-rate kill-switch on the ramp.
  8. 08“Pitch this bet to execs.” → customer + problem in $/hours → hypothesis with thresholds → fund/kill/study; model is the enabler.
The 2026 meta-trend to close on: KORE1 and Exponent both report a shift toward eval-harness + online-measurement prompts at FAANG and AI labs — rehearse “offline eval set → shadow launch → guardrail metrics → ramp” as one fluent paragraph, because a candidate who nails product sense but can’t describe an eval harness reads as a 2024 PM. And remember the round weights (Aakash, 47 placements): Behavioral ~35%, Product Sense ~20%, Execution ~15%, Technical Depth ~15%, Presentation ~10% — your single biggest prep hour is behavioral storytelling with specific AI failure stories, not framework memorization.
videoInside a $400K AI Product Sense Interview (Amazon, Meta, Google, OpenAI)Aakash GuptavideoAI Product Sense Mock Interview: “Set goals for AI-only Social Network” (ex-Google Senior PM)IGotAnOfferarticleBuilding Effective AI Agents — start simple, add agency only when neededAnthropicarticleAI Product Manager Interview Questions 2026 — the eval-harness shiftKORE1

Checkpoint

You’re outlining the agentic personal-shopper PRD and have 30 minutes. Where should you spend the most tokens to signal senior judgment?

AA thorough problem statement and market sizingBThe AI-specific cluster — eval framework, numeric quality bar, guardrails, and the human-confirmed action gateCA polished GTM and pricing plan
Sign up free to answer and see why

Checkpoint

An interviewer asks you to choose the architecture for the shopping agent. Best-reasoned answer?

AA multi-agent system (planner, searcher, negotiator, buyer) for maximum capabilityBStart with a single-purpose recommend-and-compare tool grounded in catalog data; introduce multi-step agency only when the eval shows it fails on genuinely multi-step tasksCWhatever framework is most popular this quarter
Sign up free to answer and see why

Checkpoint

For v1 of the agent, where do you place the human-in-the-loop, and why?

AAt the irreversible step — the agent drafts the cart, but a human confirms before any purchase is chargedBNowhere — full autonomy is the whole point of an agentCAt every step, requiring approval for each recommendation
Sign up free to answer and see why

Checkpoint

“Three weeks after launch, recommendation quality dropped 15% and you shipped nothing.” What’s the most likely cause and the right safeguard?

ARandom variation — LLMs are probabilistic, so there’s nothing to doBThe user base changed; just wait for it to normalizeCThe provider silently updated the model — pin model + prompt versions and gate every change on an eval so you detect and prevent it
Sign up free to answer and see why

Checkpoint

Across the whole capstone round, what single rubric signal most reliably separates a strong answer from a weak one?

AThe most complete, perfectly-structured framework applied end-to-endBNaming a specific metric, a specific failure mode, and a specific user — the through-line of every round in the trackCCiting the newest models and the highest benchmark scores
Sign up free to answer and see why

Could you run the product-sense case and write the agentic-assistant PRD end-to-end, then field the debug and exec follow-ups?

New to itGetting thereConfident

Capstone takeaways

  • Scope before you design — an “agent” hides many systems with different latency, cost, and liability profiles.
  • Run the six lenses to a wedge + segment; classify the failure modes (input/system/context) before the UX.
  • Write the 12-section PRD; spend tokens on the AI-specific cluster and the human-confirmed action gate.
  • Make the quality bar a contract — input + output metrics, a gold set, a guardrail bar, an envelope, no adjectives.
  • Default to the simplest primitive (Anthropic); add agency only when the eval demands it; gate the irreversible step.
  • In the debug round, name the ladder and pin versions; one metric + one failure mode + one user wins every round.

You’ve completed AI Product Sense & PRDs. Keep a live eval set and 2-3 specific AI failure stories ready — they carry the round.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.