Put it together: take a concrete product and map model capabilities to a feature set — choosing prompting/RAG/fine-tuning/agents per feature, scoping each by inputs/outputs/tasks/subjects, pricing it, designing for failure, and specifying the eval. The end-to-end artifact a senior AI PM produces.
From capability to a shippable feature set
Everything so far converges into one artifact a senior AI PM is expected to produce: a capability-to-feature map. Given a real product, you map what models can reliably do onto a concrete set of features, and for each feature you make the four decisions this track taught — which architecture (prompt / RAG / fine-tune / agent), how to scope it (inputs, outputs, tasks, subjects), what it costs and how fast it is, how it fails gracefully, and how you will measure it. The gate before any of this, per the research, is the capability-to-problem check: start from the user’s job and the cost of error, never from “we have AI, where do we put it.”
The most-cited AI-PM failure is starting with the technology instead of the problem — building toward a capability narrative rather than a workflow. So the map begins with a frame for each candidate feature: whose workflow does it change and what is the current workaround; what is the cost of error (reversible / expensive / irreversible / regulated); which capability tier it needs (lesson 1); is it frequent and painful and AI-shaped enough to ship; and how will we know it worked without trusting our own demo. Features that cannot answer the last question do not go on the map yet — they go back for an eval plan.
code
1THE CAPABILITY-TO-FEATURE MAP (one row per feature)23 Feature | User job & | Capability | Architecture | Scope (in/out/ | Cost & | Failure | Eval4 | err-cost | tier | (prompt/RAG/ | task/subject) | latency | design | (metric +5 | | | FT/agent) | | plan | (HITL?) | golden set)6 --------+------------+------------+--------------+---------------+---------+-----------+-----------78 Fill one row per feature. A feature isn't "ready" until every column is answered.9 If the Scope column says "anything" anywhere, it's NOT scoped -- go narrow it.10 If the Eval column is empty, it's a demo, not a feature.
Scoping: inputs, outputs, tasks, subjects
NN/g defines a feature’s scope as the breadth of its data inputs and functionality, decomposed into four levers: the kinds/size of input accepted, the types of output produced, the tasks it can handle, and the subjects it covers. Broad-scope systems (a bare ChatGPT box) trade predictability for openness; narrow-scope features — Spotify’s playlist creator, Photoshop’s generative fill, a travel-itinerary planner — make generative output predictable for users and testable for you. Narrowing each lever has a payoff: guided inputs cut user error, structured outputs become gradable, a single named task ships faster with a lower error rate, and a bounded subject raises trust and bounds the RAG corpus.
The failure mode NN/g warns against is the one to call out by name in a review: a chat box pasted onto your app. It is broad in every dimension — the user does not know what to type, the model does not know what to do, and “success” is indistinguishable from hallucination. The scoping ritual that prevents it is a one-page answer to four questions per feature: what inputs are accepted (and rejected), what outputs are possible (and what fails closed), what tasks are in-scope (and what hands off elsewhere), and what subjects are covered (and what is “out of corpus”). If any answer is “anything,” you have not scoped — you have shipped a liability. Duolingo’s AI-content rollout is the positive mirror: narrow subject (one language), narrow output (lesson templates), expert-reviewed — scope discipline is what made AI-at-scale shippable.
code
1THE FRAME GATE (run this BEFORE a feature earns a row on the map)23 1 Whose workflow changes, and what's the current workaround? -> persona + baseline4 2 Cost of error? reversible / expensive / irreversible / regulated -> err-cost rating5 3 Capability tier needed? (tier 1 ship / tier 2 scaffold / tier 3 HITL)6 4 Frequent + painful + AI-shaped enough to ship? -> ship / defer / kill7 5 How will we know it worked WITHOUT trusting our demo? -> eval metric + baseline89 Fail step 5 and the feature is not ready -- it goes back for an eval plan, not forward.
Interview angle. The case prompt is almost always “how would you use AI to improve product X?” and the discriminator is the order you reason in. Weak candidates jump to “a chatbot that…”; strong candidates run the frame gate first — name the user’s job, the cost of error, the capability tier, and the eval — then propose the feature. Leading with the problem and the metric, not the capability, is exactly what the rubric rewards, and it is the same discipline that kept Notion from shipping a generic chatbot and let it ship summarize/translate/action-items that actually moved retention.
Worked example: an AI layer for a customer-support product
Take a B2B support product and build three rows of the map. (1) Ticket auto-reply over the help center. User job: deflect repetitive tickets; error-cost: expensive but reversible (a human reviews before send). Tier 1–2. Architecture: RAG over the help center (fresh, citable, access-controlled) with a structured output ({category, draft_reply, cited_article_id}); scope: input = a ticket; output = a draft + citation, fails closed to “escalate” when no article is found; task = draft only, not send; subject = product help only. Cost: small model on the easy majority, frontier on low confidence (lesson 5). Failure design: human approves before send; “I’m not sure → escalate.” Eval: groundedness (every claim cites a real article) plus a golden set of past tickets, graded on category accuracy and reply quality.
(2) Agent that resolves “where is my order” end-to-end. User job: instant status without a human; error-cost: low-to-medium and mostly reversible (read-only lookups), but it spans tools (order DB, shipping API). This is the rare agentic candidate — multiple tool calls the user wouldn’t make themselves. Apply the three filters: agency yes, reversibility high (read-only), value density moderate. Architecture: a tightly scoped agent (or even a workflow with routing) over MCP connectors to the order and shipping systems; keep any action (issuing a refund) behind a human gate per lesson 4’s compound-failure math. Eval: trajectory and outcome — did it call the right tools and return the right status — plus a cap on steps. (3) Brand-voice rewrite of canned macros at scale. Pure behaviour, high volume, fixed task → the fine-tune/distill case (lesson 3): few-shot first, fine-tune a small model only if examples stop fitting or cost stops penciling.
code
1WORKED MAP: support product (abridged)23 Feature Architecture Scope Cost/latency Eval4 ---------------- ----------------- ------------------------- --------------- --------------------5 ticket auto-reply RAG + structured in: ticket; out: draft+ small->frontier groundedness + golden6 output cite; task: DRAFT only; cascade; cache set (category acc +7 subject: product help the system prompt reply quality)8 order-status scoped agent over in: order id/query; out: cheap model; trajectory + outcome;9 resolver MCP (read-only) status; action=refund is cap steps step cap; refund gated10 HUMAN-GATED11 brand-voice fine-tune / distill in: macro; out: rewritten small fine-tuned style eval + human spot12 macro rewrite (few-shot first) macro; subject: our voice model, high QPS audit; regression set1314 Note how each row uses a DIFFERENT lever -- that's the point of the map.
Designing for failure and proving it before commit
Because output is probabilistic, every row of the map ships with failure design as a release gate, not an afterthought: confidence signaling, an inline correction/edit path, a graceful fallback (including human handoff), and feedback capture that closes back into the eval set. This is the Google PAIR / Stanford HAI consensus — variance is the design surface, and the human-in-the-loop is the default control plane until the system earns trust on that specific workflow. The cautionary tale is Google Duplex: a capability that could book a restaurant convincingly but failed the product test because it hid its agency, omitted consent, and routed accountability nowhere. Capability without consent, confidence UI, and an override path is a liability.
code
1THE THREE-PROTOTYPE RULE (each stage gets a written go/no-go)23 Stage What it proves Artifact4 ---------------- ------------------------- --------------------------------5 1 Wizard-of-Oz the PROBLEM + the UX, with mock/manual flow, humans in loop6 a human faking the model7 2 golden eval set QUALITY on hard, realistic 50-500 examples + LLM-as-judge8 inputs (+ red team) rubric + failure-mode probes9 3 shadow / canary it survives REAL traffic 1-5% of traffic, kill switch on1011 Refuse to enter engineering build until each stage clears product + design + ML.12 This is how you close the demo-to-production gap before it closes on you.
And you do not commit engineering until each row is validated. The disciplined sequence from the research is a three-prototype rule: (1) a Wizard-of-Oz or mock to validate the problem and the UX with humans in the loop; (2) a real-model run against a golden eval set (50–500 hard, realistic examples) with an LLM-as-judge rubric, plus a red-team pass on the failure modes that matter (prompt injection for the agent, hallucination for the RAG reply); (3) a shadow/canary on 1–5% of real traffic with a kill switch before general availability. Each stage gets a written go/no-go from product, design, and ML. The whole point is to close the demo-to-production gap that compound failure and the long tail otherwise spring on you in week one.
A passing demo is survivorship bias on cherry-picked inputs. The senior move is to design the experiment that would prove you wrong — the golden set, the red team, the shadow run — before you rely on the one that confirms you.
A note on the artifact’s real audience: the map is also how you talk to engineers credibly. Each column maps to a question they respect — “which architecture” invites “what’s the eval and the cost-per-task target,” “scope” invites “what inputs do we reject,” “failure design” invites “where’s the human gate.” The vocabulary from this track is the bridge: tokens, recall, groundedness, cost-per-task, trajectory, latency-to-first-token. When an engineer says “we need to fine-tune,” the credible PM answer is “what task, what eval, what cost-per-task target?” The map turns a vague AI ambition into a shared spec both sides can argue about in the same units.
Choosing the strategic shape
One level up from individual features is the shape of the whole bet, and the map should make it explicit. The three options, each with a different eval surface, data moat, and failure mode: an embedded copilot (assistance inside an existing workflow — fast to value, low error cost, fails via hidden behaviour drift); a vertical agent (a named job done end-to-end — medium time-to-value, higher error cost, fails on multi-step accountability gaps); and a horizontal platform (a general assistant — slow, broad, fails when generality becomes vagueness). The research’s practical rule: if you are not already the platform, pick the vertical slice that owns a definable user outcome with a bounded cost of error — the smallest slice that proves the loop, exactly what Notion did.
Interview prep
The capstone maps directly onto the hardest AI-PM interview format: a case (“how would you use AI to improve product X?”) or a scoping prompt. Strong answers define the user journey’s pain, name goals and non-goals, specify dual success metrics (product KPIs and model KPIs), and define the MVP as the smallest measurable-quality workload; weak answers list use cases. Bring the four decisions per feature and the eval plan, and you answer like the principal PM running the loop, not a candidate reciting features.
01“How would you use AI to improve product X?” → map features to capabilities; per feature pick architecture, scope, cost, failure design, and an eval — don’t list use cases.
02“How do you define an MVP for an AI feature?” → the smallest workload where you can measure model quality, not the smallest feature set; ship with an eval and room for the quality curve to climb.
03“Is AI even the right tool here?” → compare to a rule-based/classical-ML baseline; if a heuristic wins on cost and reliability, don’t force AI (you’re scored down for forcing it).
04“How do you scope an AI feature?” → narrow inputs, outputs, tasks, and subjects; if any is ‘anything,’ it isn’t scoped — a bare chat box fails this test.
05“What are your success metrics?” → two columns: product KPIs (task success, retention, NPS) and model KPIs (groundedness, accuracy, latency, hallucination rate).
06“How do you validate before building?” → Wizard-of-Oz for the flow, a golden eval set + red team, then a shadow/canary on 1–5% with a kill switch.
07“Probabilistic, sometimes-wrong output — how do you launch?” → confidence UI, correction path, graceful fallback/HITL, feedback into the eval set — variance is the design surface.
08“Copilot, vertical agent, or platform?” → if you’re not the platform, take the vertical slice with a definable outcome and bounded error — the smallest slice that proves the loop.
Going deeper, the follow-ups that separate a hire: “what’s non-goal #1?” (non-goals cap the eval surface and are where scope discipline shows); “how does this MVP leave room for the model to improve?” (AI MVPs metabolize data into quality over time — name the feedback loop); “how will you detect quality regression in production?” (pinned versions, online evals, cost-per-task and groundedness dashboards from lesson 5); and “the model is great for 90% of users but bad for a 10% segment — what do you do?” (subgroup audit, then augment/reweight/route before re-launch — not “more data” vaguely). Every answer ends in a metric and a guardrail.
A stakeholder pitches “a smart assistant in our app that can help with anything.” Using the scoping framework, what do you push back with first?
AGreat — ship a chat box and let users discover what it can doBNarrow it: define accepted inputs, possible outputs, in-scope tasks, and covered subjects for one concrete job — if any answer is “anything,” it isn’t scopedCAdd a bigger model so it can truly handle anything
For the support product, which feature is the legitimate agentic candidate, and how should it be bounded?
AThe brand-voice macro rewrite — make it an autonomous agentBThe order-status resolver — a tightly scoped agent over read-only MCP tools, with any refund action gated behind a human and a step capCThe ticket auto-reply — make it a fully autonomous agent that sends replies
You’re about to commit engineering to the RAG ticket auto-reply. What does the disciplined validation sequence look like?
AWizard-of-Oz/mock for the flow, then a golden eval set + red-team pass, then a shadow/canary on 1–5% of traffic with a kill switch — each with a go/no-goBShip to 100% of users and watch the support queueCDemo it to leadership; if the demo is clean, build and launch
An interviewer asks you to define the MVP for an AI feature whose model quality will improve over time. Strongest definition?
AThe smallest set of features we can ship quicklyBWhatever the demo already does wellCThe smallest workload where we can measure model quality — shipped with an eval and room for the quality curve to climb as data accrues
Your launched feature works well for 90% of users but produces poor results for a specific segment. What’s the senior response?
AJust collect more data and retrain — it’ll average outBRun a subgroup audit to localize the failure, then augment/reweight the data or route that segment to a stronger path before re-launch — and track the segment metricCShip as-is; 90% is good enough to keep
Could you take a real product into an interview and produce a capability-to-feature map — architecture, scope, cost, failure design, and eval per feature — and defend the strategic shape?
New to itGetting thereConfident
Takeaways
The deliverable is a capability-to-feature map: per feature, decide architecture, scope, cost/latency, failure design, and the eval.
Start from the user’s job and the cost of error, never from “we have AI” — features without an eval plan aren’t ready.
Scope on four levers (inputs, outputs, tasks, subjects); if any is “anything,” narrow it — a bare chat box fails the test.
Each feature tends to use a different lever: RAG for cited help, a gated agent for multi-tool tasks, fine-tune/distill for high-volume voice.
Design for failure as a release gate (confidence UI, correction, fallback/HITL, feedback→eval) and validate with Wizard-of-Oz → golden set + red team → shadow/canary.
Pick the strategic shape deliberately; if you’re not the platform, take the vertical slice with a bounded outcome — the smallest slice that proves the loop.
You can now scope, price, sequence, and defend an AI feature set — and talk credibly with engineers about every layer of the stack.