A traditional PRD is a contract for a deterministic system; AI breaks that contract twice. The shift from “output equals X” to “the eval passes >= threshold across the gold set,” the four-layer acceptance framework, numeric quality bars (not adjectives), the input/output guardrail taxonomy, and the living document that doesn’t rot by month two.
A PRD that breaks its own contract
A traditional PRD is a contract for a deterministic system: same input, same output, acceptance criteria written as expected values. AI breaks that contract twice. First, the model is non-deterministic by design — the “right” output for one query may differ next run. Second, quality is observable in distributions, not pass/fail — a single bad output isn’t a bug, but 5% bad outputs in a vertical is a critical regression. The operational definition to anchor on (Ainna): an AI PRD extends the standard spec with eval-driven acceptance criteria, guardrail definitions, model-dependency documentation, and living-document architecture. This lesson is how you write that document so an engineer — and an SRE — can build and monitor against it.
The shift is from “the output should equal X” to “the eval should pass >= Y threshold across the gold set.” The implication is concrete: budget 15-20% of the PRD’s length to the eval framework, not 0%. Aakash Gupta’s 2026 observation sharpens the bar — “40 lines of markdown replaced my 15-page PRD,” and the claim “I use ChatGPT for PRDs” is now the wrong answer to the most common AI interview question, because AI drafting is baseline, not a moat. Senior PMs lean toward terseness in the main spec and depth in a sibling eval-set doc: a ~2-3 page narrative PRD that links to a 3-5 page eval specification.
A 12-section structure synthesized from leading templates
Three independently-authored practitioner templates — Ainna’s 3-tier pyramid, Miqdad Jaffer’s 9 sections (OpenAI), Product School’s 5 sections — converge on the same backbone. There is no single industry-standard AI-PRD spec yet, but synthesizing these leading templates yields a 12-section structure in four clusters. Two design choices deserve defense. First, Quality Bar (numeric thresholds) is separate from Success Metrics: success metrics measure business outcomes; quality bars measure model behavior against a labeled eval set — conflating them is the single most common anti-pattern. Second, the Living Epilogue (a running decision log) is non-negotiable: requirements evolve as fast as the model under them, so a frozen PRD is a wrong PRD by month two.
code
1THE AI PRD -- 12 sections, 4 clusters (synthesized from leading templates)23 STRATEGIC FOUNDATION TL;DR | Problem & Customer | Success Metrics | Scope4 AI-SPECIFIC LAYER Eval Framework | Quality Bar (numeric) | Guardrails5 (input+output) | Model & Data Strategy | Responsible AI6 OPERATIONAL PLAN Monitoring & Drift | Failure Modes & Fallback | GTM7 LIVING EPILOGUE running log: Date | Decision | Reason | Reversible?89 Token budget (senior default):10 ~2-3 pages narrative (TL;DR + Problem + Outcome + Scope)11 ~3-5 pages AI-specific (Eval, Quality Bar, Guardrails, Model/Data, Resp. AI)12 ~1 page operational (Monitoring, Drift, Failure Modes, GTM)13 living log updated weekly -> a frozen PRD is wrong by month two
Why keep scientific rigor, executive narrative, and engineering tractability as discrete sections? Because most PRDs fail by collapsing two of the three into a bullet list. The synthesized playbook keeps them separate with their own owners: the eval framework (rigor) is owned with engineering/data science, the outcome narrative (the next lesson) is owned by the PM up the chain, and the guardrail/fallback table (tractability) is owned with the on-call/SRE function. Interview angle. A PRD-round prompt (“outline a PRD for an agentic assistant”) is graded on whether your structure surfaces all three — candidates who write a beautiful problem statement but no eval section, or numeric bars but no fallback, read as missing the AI-specific half.
Acceptance criteria for probabilistic output
For a deterministic system, “given input X, return Y” is the universal acceptance shape. For a probabilistic system the equivalent is: “given the eval set, the model produces a distribution of outputs whose [metric] is at least [threshold] across [slice], refreshed at [cadence].” Ainna’s framing is the cleanest — structured, repeatable tests that define “what good looks like,” effectively replacing traditional acceptance criteria. The evals are the acceptance criteria; they are not separate from them. The Xenoss four-layer framework gives the taxonomy that stops a single test from hiding a regression.
code
1FOUR-LAYER ACCEPTANCE FRAMEWORK (Xenoss) -- each layer catches a distinct regression23 Layer Question Why one test can't see it4 --------------- ------------------------------------ --------------------------5 Business result did it move a north-star metric? accurate but unadopted6 User behavior fewer steps / retries / escalations? safe but unhelpful7 Model behavior numeric bars on a labeled eval set? fluent but hallucinated8 Operational latency / error / cost in envelope? correct but too slow/costly910 A model can be accurate AND slow, safe AND unhelpful, fluent AND hallucinated.11 Layering surfaces each failure mode separately.
The worked example to memorize is Miqdad Jaffer’s (Product Lead @ OpenAI), from the Shopify Auto Write case: “The AI must achieve >=90% accuracy on a labeled test set of 10,000 queries, limit hallucination rates to <2% via RAG integration, and flag inappropriate outputs with 98% precision, validated monthly through human review.” That single sentence carries a model-behavior bar (90% / <2%), a guardrail bar (98% precision on the safety classifier), and an operational cadence (monthly human review) — roughly 270 characters of acceptance criteria replacing a dozen vague “the response should…” statements. Hamel Husain’s rule (canonical, drawn from teaching 700+ engineers and PMs) sets the integrity bar: an eval is only an eval if its source dataset is preserved, its rubric is calibrated against human review at least quarterly, and its metric is decomposed by user slice.
Refound/Lenny add a fifth dimension — executability: “the purest sense of what a PRD should be is this eval judge that’s telling you exactly what it should be, and it’s automatic and running constantly.” So write at least one acceptance criterion as a runnable LLM-as-judge prompt that returns pass/fail on every CI run. A PRD a CI bot can lint is a PRD that survives contact with reality.
code
1ACCEPTANCE CRITERIA -- PROBABILISTIC OUTPUT (drop this at the top of the AI-specific section)23 Eval set: 10,000 labeled queries, refreshed each model upgrade.4 Business result: 15% lift in AI-assisted task completion within 180 days.5 Model behavior: >=90% accuracy vs gold; <2% hallucination; <3% off-brand tone.6 Guardrails: 98% precision on toxicity/PII; 99.5% recall on jailbreak.7 Operational: p95 latency <2.5s; cost/query <$0.04; fallback path 100% pass.8 Cadence: automated sweep daily; human calibration monthly.910 Note what's absent: the word "fast." "Fast isn't a target." -- Miqdad Jaffer
The eval-set spec — the sibling doc that does the work
The terse main PRD links to a fatter eval-set specification (often 3-5 pages), and that doc is where the rigor lives. Hamel Husain’s integrity rules define what makes it real: the source dataset is preserved (you can re-run last quarter’s set), the rubric is calibrated against human review at least quarterly (the LLM-judge agrees with experts), and the metric is decomposed by user slice (so a regression in one vertical isn’t hidden by a healthy average). The composition matters too: a good eval set is built from production-derived traces, includes the hard and adversarial cases, and grows by turning every incident into a permanent regression case — the eval-driven-development loop.
code
1EVAL-SET SPEC -- the sibling doc the PRD links to23 Size & source ~2-10k cases, production-derived + adversarial seeds4 Slices by vertical / user segment / query type (decompose!)5 Graders code-based (format/grounding) + model-based (tone/coverage)6 + human (final bar) -- avoid rigid grading; use partial credit7 Calibration judge vs human review, >= quarterly (Hamel's rule)8 Refresh cadence each model/prompt upgrade; named owner for drift9 Growth rule every production incident -> a new permanent regression case1011 The main PRD stays ~2-3 pages and LINKS here. Depth lives in the eval doc.
Success metrics vs guardrails — the two-direction taxonomy
A PRD that separates business metrics from model metrics mistakes a feature for a product, so the metric ledger has four categories (user-experience, safety, operational, business) — and guardrails get their own taxonomy split by direction. Input guardrails: topic-relevance filters, prompt-injection detection, PII detection, blocklists. Output guardrails: toxicity filters, hallucination detection, format validation, brand-voice compliance, regulatory restrictions. Fiddler adds a second axis: gateway-level guardrails (PII, jailbreak, prompt injection — non-negotiable defaults) vs evaluation-level guardrails (groundedness, hallucination, bias — the staged-rollout flags that decide whether a model change can ship).
code
1GUARDRAIL TAXONOMY -- two axes23 By DIRECTION4 Input guardrails topic relevance, prompt-injection detect, PII detect,5 blocklists (act on the request)6 Output guardrails toxicity, hallucination detect, format/brand-voice,7 regulatory restrictions (act on the response)89 By LAYER (Fiddler)10 Gateway-level PII, jailbreak, prompt injection -> non-negotiable defaults11 Evaluation-level groundedness, hallucination, bias -> staged-rollout flags1213 A missing guardrail table is a missing definition of "deployment-ready."
The Shopify Auto Write case shows the discipline end-to-end: a text generator for merchant product descriptions on GPT-3 (Davinci-003), with streaming output for latency, regeneration limits to control abuse, content moderation to satisfy Apple’s App Store TOS, and a launch metric of “15% merchant adoption within 180 days.” A single, dated, named-persona metric prevents the most common AI org failure: shipping a model that “feels magical in demo” and “languishes in production.” The PRD also required “robust feedback loops … and human-in-the-loop QA measured by quarterly review cycles” — the living-document architecture that stops the doc (and the model) from rotting. Interview angle. Tie the launch metric to a named persona and a date (merchants / 180 days), not a generic “increase engagement” — that specificity is the senior signal.
Interview prep
PRD rounds reward eval-as-acceptance-criteria, numeric bars over adjectives, a two-direction guardrail table, and a living log. Lead with the structure, then a concrete threshold.
01“How is an AI PRD different from a normal PRD?” → it adds eval-driven acceptance criteria, guardrails, model-dependency docs, and a living log; quality is a distribution, not pass/fail.
02“Write acceptance criteria for a probabilistic feature.” → “given the eval set, [metric] >= [threshold] across [slice], refreshed at [cadence]” — evals are the acceptance criteria.
03“What are the success metrics?” → split business vs model bars; e.g. 15% adoption in 180 days (business) AND >=90% / <2% hallucination on a 10k gold set (model).
04“What guardrails would you spec?” → by direction (input: PII/injection; output: toxicity/hallucination/format) and layer (gateway defaults vs evaluation-level rollout flags).
05“How long should the PRD be?” → terse spec (~2-3 pp) + a sibling eval-set doc (~3-5 pp); 15-20% of length on the eval framework.
06“Make one criterion executable.” → write it as an LLM-as-judge prompt returning pass/fail in CI (Refound) — a PRD a bot can lint.
07“What keeps the PRD from going stale?” → a living decision log (Date / Decision / Reason / Reversible?) updated weekly + a fixed eval-refresh cadence.
08“The team says ‘make it fast.’ Your response?” → “fast isn’t a target” (Miqdad) — replace it with p95 latency <2.5s and a cost/query ceiling.
Follow-ups push on rigor and failure modes. Expect “how do you stop the eval set from drifting?” (fixed refresh cadence + quarterly human calibration, owned by a named person — Hamel’s integrity bar), “what’s your false-positive rate on the safety filter?” (track it as a first-class metric; a too-aggressive gateway classifier blocks legitimate flows — Fiddler), and “how do you avoid shipping to the metric?” (tie the quality bar to a business metric so the metric can’t deliver without the outcome — Cagan’s feature-shipping-disguised-as-outcome trap). The convergence across Cagan (outcome), Hamel (eval), Ainna (guardrail), and Miqdad (named-persona P&L) is the strongest evidence this synthesis is the emerging best practice — even without a single canonical standard.
A reviewer says your AI PRD’s success-metrics section is fine because it lists “>=92% model accuracy.” What’s the structural problem?
AQuality bar (model behavior) and success metrics (business outcome) are being conflated — they belong in separate sections; success metrics need a business outcome like adoption or deflectionBNothing — a clear accuracy bar is exactly what an AI PRD needsCThe accuracy should be expressed as a letter grade for executives
You must write one acceptance criterion for a feature that summarizes support tickets. Which is a proper probabilistic acceptance criterion?
A“The summary should accurately capture the ticket.”B“On a 2,000-ticket gold set, ROUGE-L >=0.45 AND an LLM-as-judge rates faithfulness >=4/5 on >=90% of samples, decomposed by ticket category, refreshed monthly.”C“The summary should be fast and high-quality.”
Your PRD has numeric model bars but no guardrail table. An interviewer asks what’s missing for “deployment-ready.” Best answer?
AA two-direction guardrail table — input (PII, prompt-injection, topic filters) and output (toxicity, hallucination, format/brand-voice) — plus gateway vs evaluation-level layeringBNothing material — guardrails are an appendix concernCA longer problem statement to justify the feature
Two months after launch, your AI PRD no longer matches the product — the model was upgraded twice and requirements shifted. What does the synthesized 12-section structure prescribe to prevent this?
AFreeze the PRD at launch so there’s a stable referenceBA Living Epilogue — a running decision log (Date / Decision / Reason / Reversible?) updated weekly, plus a fixed eval-refresh cadenceCRewrite the entire PRD from scratch each quarter
Shopify Auto Write’s PRD set the launch metric as “15% merchant adoption within 180 days.” Why is that framing stronger than “increase merchant engagement”?
AIt’s a bigger, more ambitious numberBIt avoids needing an eval setCA single dated, named-persona target prevents the “magical in demo, languishes in production” failure by making one team accountable to one measurable outcome by a date