Lesson 4 of 6 · 49 min

Writing AI PRDs

A traditional PRD is a contract for a deterministic system; AI breaks that contract twice. The shift from “output equals X” to “the eval passes >= threshold across the gold set,” the four-layer acceptance framework, numeric quality bars (not adjectives), the input/output guardrail taxonomy, and the living document that doesn’t rot by month two.

A PRD that breaks its own contract

A traditional PRD is a contract for a deterministic system: same input, same output, acceptance criteria written as expected values. AI breaks that contract twice. First, the model is non-deterministic by design — the “right” output for one query may differ next run. Second, quality is observable in distributions, not pass/fail — a single bad output isn’t a bug, but 5% bad outputs in a vertical is a critical regression. The operational definition to anchor on (Ainna): an AI PRD extends the standard spec with eval-driven acceptance criteria, guardrail definitions, model-dependency documentation, and living-document architecture. This lesson is how you write that document so an engineer — and an SRE — can build and monitor against it.
The shift is from “the output should equal X” to “the eval should pass >= Y threshold across the gold set.” The implication is concrete: budget 15-20% of the PRD’s length to the eval framework, not 0%. Aakash Gupta’s 2026 observation sharpens the bar — “40 lines of markdown replaced my 15-page PRD,” and the claim “I use ChatGPT for PRDs” is now the wrong answer to the most common AI interview question, because AI drafting is baseline, not a moat. Senior PMs lean toward terseness in the main spec and depth in a sibling eval-set doc: a ~2-3 page narrative PRD that links to a 3-5 page eval specification.
Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya ShankarLenny's Podcast

A 12-section structure synthesized from leading templates

Three independently-authored practitioner templates — Ainna’s 3-tier pyramid, Miqdad Jaffer’s 9 sections (OpenAI), Product School’s 5 sections — converge on the same backbone. There is no single industry-standard AI-PRD spec yet, but synthesizing these leading templates yields a 12-section structure in four clusters. Two design choices deserve defense. First, Quality Bar (numeric thresholds) is separate from Success Metrics: success metrics measure business outcomes; quality bars measure model behavior against a labeled eval set — conflating them is the single most common anti-pattern. Second, the Living Epilogue (a running decision log) is non-negotiable: requirements evolve as fast as the model under them, so a frozen PRD is a wrong PRD by month two.
code
1THE AI PRD -- 12 sections, 4 clusters (synthesized from leading templates)23  STRATEGIC FOUNDATION    TL;DR | Problem & Customer | Success Metrics | Scope4  AI-SPECIFIC LAYER       Eval Framework | Quality Bar (numeric) | Guardrails5                          (input+output) | Model & Data Strategy | Responsible AI6  OPERATIONAL PLAN        Monitoring & Drift | Failure Modes & Fallback | GTM7  LIVING EPILOGUE         running log: Date | Decision | Reason | Reversible?89  Token budget (senior default):10    ~2-3 pages narrative  (TL;DR + Problem + Outcome + Scope)11    ~3-5 pages AI-specific (Eval, Quality Bar, Guardrails, Model/Data, Resp. AI)12    ~1 page operational   (Monitoring, Drift, Failure Modes, GTM)13    living log updated weekly  ->  a frozen PRD is wrong by month two
Why keep scientific rigor, executive narrative, and engineering tractability as discrete sections? Because most PRDs fail by collapsing two of the three into a bullet list. The synthesized playbook keeps them separate with their own owners: the eval framework (rigor) is owned with engineering/data science, the outcome narrative (the next lesson) is owned by the PM up the chain, and the guardrail/fallback table (tractability) is owned with the on-call/SRE function. Interview angle. A PRD-round prompt (“outline a PRD for an agentic assistant”) is graded on whether your structure surfaces all three — candidates who write a beautiful problem statement but no eval section, or numeric bars but no fallback, read as missing the AI-specific half.

Acceptance criteria for probabilistic output

For a deterministic system, “given input X, return Y” is the universal acceptance shape. For a probabilistic system the equivalent is: “given the eval set, the model produces a distribution of outputs whose [metric] is at least [threshold] across [slice], refreshed at [cadence].” Ainna’s framing is the cleanest — structured, repeatable tests that define “what good looks like,” effectively replacing traditional acceptance criteria. The evals are the acceptance criteria; they are not separate from them. The Xenoss four-layer framework gives the taxonomy that stops a single test from hiding a regression.
code
1FOUR-LAYER ACCEPTANCE FRAMEWORK (Xenoss) -- each layer catches a distinct regression23  Layer            Question                              Why one test can't see it4  ---------------  ------------------------------------  --------------------------5  Business result  did it move a north-star metric?      accurate but unadopted6  User behavior    fewer steps / retries / escalations?  safe but unhelpful7  Model behavior   numeric bars on a labeled eval set?   fluent but hallucinated8  Operational      latency / error / cost in envelope?   correct but too slow/costly910  A model can be accurate AND slow, safe AND unhelpful, fluent AND hallucinated.11  Layering surfaces each failure mode separately.
The worked example to memorize is Miqdad Jaffer’s (Product Lead @ OpenAI), from the Shopify Auto Write case: “The AI must achieve >=90% accuracy on a labeled test set of 10,000 queries, limit hallucination rates to <2% via RAG integration, and flag inappropriate outputs with 98% precision, validated monthly through human review.” That single sentence carries a model-behavior bar (90% / <2%), a guardrail bar (98% precision on the safety classifier), and an operational cadence (monthly human review) — roughly 270 characters of acceptance criteria replacing a dozen vague “the response should…” statements. Hamel Husain’s rule (canonical, drawn from teaching 700+ engineers and PMs) sets the integrity bar: an eval is only an eval if its source dataset is preserved, its rubric is calibrated against human review at least quarterly, and its metric is decomposed by user slice.
Refound/Lenny add a fifth dimension — executability: “the purest sense of what a PRD should be is this eval judge that’s telling you exactly what it should be, and it’s automatic and running constantly.” So write at least one acceptance criterion as a runnable LLM-as-judge prompt that returns pass/fail on every CI run. A PRD a CI bot can lint is a PRD that survives contact with reality.
code
1ACCEPTANCE CRITERIA -- PROBABILISTIC OUTPUT  (drop this at the top of the AI-specific section)23  Eval set:       10,000 labeled queries, refreshed each model upgrade.4  Business result: 15% lift in AI-assisted task completion within 180 days.5  Model behavior:  >=90% accuracy vs gold; <2% hallucination; <3% off-brand tone.6  Guardrails:      98% precision on toxicity/PII; 99.5% recall on jailbreak.7  Operational:     p95 latency <2.5s; cost/query <$0.04; fallback path 100% pass.8  Cadence:         automated sweep daily; human calibration monthly.910  Note what's absent: the word "fast." "Fast isn't a target." -- Miqdad Jaffer

The eval-set spec — the sibling doc that does the work

The terse main PRD links to a fatter eval-set specification (often 3-5 pages), and that doc is where the rigor lives. Hamel Husain’s integrity rules define what makes it real: the source dataset is preserved (you can re-run last quarter’s set), the rubric is calibrated against human review at least quarterly (the LLM-judge agrees with experts), and the metric is decomposed by user slice (so a regression in one vertical isn’t hidden by a healthy average). The composition matters too: a good eval set is built from production-derived traces, includes the hard and adversarial cases, and grows by turning every incident into a permanent regression case — the eval-driven-development loop.
code
1EVAL-SET SPEC -- the sibling doc the PRD links to23  Size & source      ~2-10k cases, production-derived + adversarial seeds4  Slices             by vertical / user segment / query type (decompose!)5  Graders            code-based (format/grounding) + model-based (tone/coverage)6                     + human (final bar)  -- avoid rigid grading; use partial credit7  Calibration        judge vs human review, >= quarterly (Hamel's rule)8  Refresh cadence    each model/prompt upgrade; named owner for drift9  Growth rule        every production incident -> a new permanent regression case1011  The main PRD stays ~2-3 pages and LINKS here. Depth lives in the eval doc.

Success metrics vs guardrails — the two-direction taxonomy

A PRD that separates business metrics from model metrics mistakes a feature for a product, so the metric ledger has four categories (user-experience, safety, operational, business) — and guardrails get their own taxonomy split by direction. Input guardrails: topic-relevance filters, prompt-injection detection, PII detection, blocklists. Output guardrails: toxicity filters, hallucination detection, format validation, brand-voice compliance, regulatory restrictions. Fiddler adds a second axis: gateway-level guardrails (PII, jailbreak, prompt injection — non-negotiable defaults) vs evaluation-level guardrails (groundedness, hallucination, bias — the staged-rollout flags that decide whether a model change can ship).
code
1GUARDRAIL TAXONOMY -- two axes23  By DIRECTION4    Input guardrails   topic relevance, prompt-injection detect, PII detect,5                       blocklists  (act on the request)6    Output guardrails  toxicity, hallucination detect, format/brand-voice,7                       regulatory restrictions  (act on the response)89  By LAYER (Fiddler)10    Gateway-level      PII, jailbreak, prompt injection -> non-negotiable defaults11    Evaluation-level   groundedness, hallucination, bias -> staged-rollout flags1213  A missing guardrail table is a missing definition of "deployment-ready."
The Shopify Auto Write case shows the discipline end-to-end: a text generator for merchant product descriptions on GPT-3 (Davinci-003), with streaming output for latency, regeneration limits to control abuse, content moderation to satisfy Apple’s App Store TOS, and a launch metric of “15% merchant adoption within 180 days.” A single, dated, named-persona metric prevents the most common AI org failure: shipping a model that “feels magical in demo” and “languishes in production.” The PRD also required “robust feedback loops … and human-in-the-loop QA measured by quarterly review cycles” — the living-document architecture that stops the doc (and the model) from rotting. Interview angle. Tie the launch metric to a named persona and a date (merchants / 180 days), not a generic “increase engagement” — that specificity is the senior signal.

Interview prep

PRD rounds reward eval-as-acceptance-criteria, numeric bars over adjectives, a two-direction guardrail table, and a living log. Lead with the structure, then a concrete threshold.
  1. 01“How is an AI PRD different from a normal PRD?” → it adds eval-driven acceptance criteria, guardrails, model-dependency docs, and a living log; quality is a distribution, not pass/fail.
  2. 02“Write acceptance criteria for a probabilistic feature.” → “given the eval set, [metric] >= [threshold] across [slice], refreshed at [cadence]” — evals are the acceptance criteria.
  3. 03“What are the success metrics?” → split business vs model bars; e.g. 15% adoption in 180 days (business) AND >=90% / <2% hallucination on a 10k gold set (model).
  4. 04“What guardrails would you spec?” → by direction (input: PII/injection; output: toxicity/hallucination/format) and layer (gateway defaults vs evaluation-level rollout flags).
  5. 05“How long should the PRD be?” → terse spec (~2-3 pp) + a sibling eval-set doc (~3-5 pp); 15-20% of length on the eval framework.
  6. 06“Make one criterion executable.” → write it as an LLM-as-judge prompt returning pass/fail in CI (Refound) — a PRD a bot can lint.
  7. 07“What keeps the PRD from going stale?” → a living decision log (Date / Decision / Reason / Reversible?) updated weekly + a fixed eval-refresh cadence.
  8. 08“The team says ‘make it fast.’ Your response?” → “fast isn’t a target” (Miqdad) — replace it with p95 latency <2.5s and a cost/query ceiling.
Follow-ups push on rigor and failure modes. Expect “how do you stop the eval set from drifting?” (fixed refresh cadence + quarterly human calibration, owned by a named person — Hamel’s integrity bar), “what’s your false-positive rate on the safety filter?” (track it as a first-class metric; a too-aggressive gateway classifier blocks legitimate flows — Fiddler), and “how do you avoid shipping to the metric?” (tie the quality bar to a business metric so the metric can’t deliver without the outcome — Cagan’s feature-shipping-disguised-as-outcome trap). The convergence across Cagan (outcome), Hamel (eval), Ainna (guardrail), and Miqdad (named-persona P&L) is the strongest evidence this synthesis is the emerging best practice — even without a single canonical standard.
articleA Proven AI PRD Template by Miqdad Jaffer (Product Lead @ OpenAI) — Shopify Auto WriteProduct CompassarticleLLM Evals: Everything You Need to Know (the eval-as-acceptance-criterion authority)Hamel HusainarticleHow Do You Write a PRD for AI Products? — eval-driven acceptance, guardrail taxonomyAinnaarticleAI Guardrails Metrics — gateway vs evaluation-level guardrailsFiddler AI

Checkpoint

A reviewer says your AI PRD’s success-metrics section is fine because it lists “>=92% model accuracy.” What’s the structural problem?

AQuality bar (model behavior) and success metrics (business outcome) are being conflated — they belong in separate sections; success metrics need a business outcome like adoption or deflectionBNothing — a clear accuracy bar is exactly what an AI PRD needsCThe accuracy should be expressed as a letter grade for executives
Sign up free to answer and see why

Checkpoint

You must write one acceptance criterion for a feature that summarizes support tickets. Which is a proper probabilistic acceptance criterion?

A“The summary should accurately capture the ticket.”B“On a 2,000-ticket gold set, ROUGE-L >=0.45 AND an LLM-as-judge rates faithfulness >=4/5 on >=90% of samples, decomposed by ticket category, refreshed monthly.”C“The summary should be fast and high-quality.”
Sign up free to answer and see why

Checkpoint

Your PRD has numeric model bars but no guardrail table. An interviewer asks what’s missing for “deployment-ready.” Best answer?

AA two-direction guardrail table — input (PII, prompt-injection, topic filters) and output (toxicity, hallucination, format/brand-voice) — plus gateway vs evaluation-level layeringBNothing material — guardrails are an appendix concernCA longer problem statement to justify the feature
Sign up free to answer and see why

Checkpoint

Two months after launch, your AI PRD no longer matches the product — the model was upgraded twice and requirements shifted. What does the synthesized 12-section structure prescribe to prevent this?

AFreeze the PRD at launch so there’s a stable referenceBA Living Epilogue — a running decision log (Date / Decision / Reason / Reversible?) updated weekly, plus a fixed eval-refresh cadenceCRewrite the entire PRD from scratch each quarter
Sign up free to answer and see why

Checkpoint

Shopify Auto Write’s PRD set the launch metric as “15% merchant adoption within 180 days.” Why is that framing stronger than “increase merchant engagement”?

AIt’s a bigger, more ambitious numberBIt avoids needing an eval setCA single dated, named-persona target prevents the “magical in demo, languishes in production” failure by making one team accountable to one measurable outcome by a date
Sign up free to answer and see why

Could you draft the 12-section AI PRD, write eval-driven acceptance criteria, spec a guardrail table, and field the PRD questions above?

New to itGetting thereConfident

Takeaways

  • An AI PRD trades “output equals X” for “the eval passes >= threshold across the gold set”; evals ARE the acceptance criteria.
  • Use the 12-section structure; keep the numeric Quality Bar separate from business Success Metrics.
  • Four acceptance layers — business, user behavior, model behavior, operational — each catches a distinct regression.
  • Replace adjectives with thresholds; if an SRE can’t alert on it and a CI bot can’t lint it, it isn’t acceptance criteria.
  • Spec guardrails by direction (input/output) and layer (gateway defaults vs evaluation-level rollout flags).
  • A living decision log + eval-refresh cadence keep the PRD from being wrong by month two.

Next: stakeholder narrative & exec comms — the story and communication around an AI bet.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.