An AI feature is a distribution, not a binary — so “done” stops being a checkbox. The deterministic-to-distributional shift, eval-driven definitions of done written before code, the rule-based baseline test, and the “the model is 85% accurate — do you ship?” question that opens every AI PM case round.
Why “done” isn’t a checkbox anymore
A deterministic feature is a vending machine: same button, same can, every time — and “done” means the tests pass. An AI feature is a gumball machine: the same prompt produces different outputs across users, sessions, model updates, and contexts (Jeff Gothelf’s framing). The first senior failure mode is copying the deterministic PM playbook onto a probabilistic product — defining done as “spec met / no P0 bugs” when the thing you shipped is a distribution of behaviours. This lesson rebuilds the definition of success from that one fact, and it is the exact ground the AI PM case round opens on.
Anthropic states the gap bluntly in Demystifying evals for AI agents: “the capabilities that make agents useful also make them difficult to evaluate,” and “a task that passed on one eval run might fail on the next” because each task has its own success rate. So a single green run tells you almost nothing. The consequence for a PM is categorical: your definition of done can no longer be a binary test-pass — it has to be a statement about the distribution of outcomes, plus a tripwire for when that distribution drifts. Get this wrong and you will ship on a demo that worked once and be surprised in production.
The discipline that fixes it is what Anthropic calls eval-driven development: teams “build evals to define planned capabilities before agents can fulfill them,” which “force[s] product teams to specify what success means for the agent” before the first prompt ships. This inverts the usual order — the eval spec lives in the PRD, not after the prototype works. The reason it’s powerful is subtle: writing the eval is a forcing function on the requirement itself. As Anthropic puts it, two engineers reading the same spec interpret edge cases differently; the eval is what resolves the ambiguity. If you can’t write the eval, the requirement isn’t concrete enough to build.
Here is the shift made concrete. Gothelf’s now-standard definition of done for an AI feature reads: “For 80% of inputs in category X, the system returns a response that meets quality bar Y; for the remaining 20%, the failure mode is degraded but not embarrassing.” Read that carefully — it has three moving parts a deterministic spec never had: a target rate (80%), a scoped input class (category X, not “all inputs”), and an explicit promise about the tail (the 20% fails gracefully, not catastrophically). The tail clause is the part juniors drop, and it’s the part that keeps you off the front page of the news.
code
1DETERMINISTIC "done" vs AI-FEATURE "done"2 --------------------------- ----------------------------------------3 all tests pass / no P0 bugs for 80% of inputs in category X, meets4 quality bar Y; the other 20% degrades5 gracefully, not embarrassingly6 acceptance = binary acceptance = distributional + tripwire7 spec authored by PM, fixed at ship eval spec authored BEFORE code; revisited8 ~every 6 months as the model changes9 "it works" = it returned "it works" = the DISTRIBUTION of outputs10 the right value once clears the bar across a labelled set1112 Source: Gothelf, "What done means when shipping AI features"; Anthropic evals.
Two consequences fall out of the distributional view, and both are senior signals. First, a single bad output is not a bug — a shift in the failure rate is. If your spec says 80% and one user hits the 20%, the system is behaving as designed; you escalate only when the rate moves. Second, the definition has a shelf life: Notion’s AI lead Sarah Sachs says the team must “start over every six months” because frontier-model capability re-defines the distribution you can achieve. Interview angle. When asked “how do you define done for an AI feature,” lead with Gothelf’s distributional sentence — a rate, a scoped input class, and a graceful tail — not “when it passes QA.” Saying “done is a stance, not a checkbox” lands the point.
A scalar that operationalises the distribution is Anthropic’s pass@k: instead of pass@1 (one stochastic attempt, which understates a system that’s allowed to retry), pass@k measures “the likelihood of at least one correct solution in k attempts.” It’s the honest metric when your product re-generates or lets the user retry, and it forces you to think in attempts, not single shots. Pair it with Anthropic’s base operational metrics — latency, token usage, cost per task, and error rate — and you have a launch gate that describes a distribution rather than a lucky demo.
Problem fit before model fit: the rule-based baseline
Before you define success for an AI feature you must earn the right to build one. Google PAIR’s User Needs + Defining Success chapter insists the first question is “Can AI solve this problem in a unique way?” and warns that “even the best AI will fail if it doesn’t provide unique value to users.” The operational test a senior PM bakes into every AI PRD: what does a dumb rule-based baseline score on this exact task? If a heuristic or a regex hits 95% precision and recall, ship the heuristic — a non-deterministic model that’s harder to evaluate, costs per token, and can hallucinate is a strict downgrade for a problem a rule already solves.
PAIR also makes you choose a mode before a metric: automation (the system acts for the user — optimise efficiency and safety) versus augmentation (the system assists — optimise the human-plus-AI outcome). The mode dictates the metric stack downstream, so picking it is step zero. And PAIR publishes the kind of tripwire numbers teams should write down in advance: “average rate of rejection of smart playlists and routes goes above 20%” or “users open the app frequently but only complete runs 25% of the time” — concrete failure thresholds, not vibes. Interview angle. A favourite probe is “prove the AI version beats a non-AI baseline” — strong candidates propose a head-to-head A/B against the rule-based version with a falsifiable hypothesis; weak ones assume the AI is better because it’s AI.
If a rule-based system can hit the bar, ship the rule. AI is the right tool only when the problem is genuinely fuzzy, open-ended, or personalised — and you can show it beats the dumb baseline on your own labelled set.
The eval set is the prerequisite — and it’s smaller than you think
Defining success means writing the artifact that measures success: a labelled eval set. The most common mistake is treating it like a training set — “get as much data as possible.” Every primary source pushes back. Notion deliberately starts tiny: “evaluation sets typically start with 10–20 examples for a given scenario,” then grow weekly by curating “hand-written examples for edge cases and real-world usage automatically logged during production.” Hamel Husain and Shreya Shankar recommend purpose-built CI datasets of “100+ examples” and a fresh look at “100+ traces” every 2–4 weeks. The headline: coverage by dimension beats volume by row.
Why shape beats size: coverage is what fails in LLM products, not statistical power. If a whole query category — say a user who mixes Japanese, Korean, and English in one workspace (a real Notion test case) — is absent from the eval set, your confidence in the metric is structurally wrong, not just noisy. Anthropic operationalises this as “build balanced problem sets” that “test where behavior should and should not occur” — i.e. include the negative cases the model must refuse, not only the positive cases it must pass. A senior eval set is a portfolio: a tight regression set, a sampled-production set, and a small frontier set designed to differentiate competing models.
code
1EVAL SET: shape over size (what each primary source actually does)23 Notion start 10-20 examples/scenario; grow weekly from prod traces4 Hamel/Shankar 100+ purpose-built CI examples; 100+ fresh traces / 2-4 wks5 Anthropic BALANCED sets: positive + negative ("should NOT happen") cases6 + multiple replications per task -> report pass@k, not pass@178 Cover by DIMENSION (user type, intent, language, edge/adversarial),9 not by row count. A missing category is a SILENT confidence error.
Anthropic adds a statistical subtlety PMs miss: tests need independence of trials so the variance you see is the model’s, not the harness’s — otherwise the numbers are uninterpretable. Practically, that means running each task several times and reporting pass@k across the replications rather than trusting one deterministic pass@1. The PM takeaway: the eval set isn’t a one-time deliverable, it’s a recurring investment that belongs on quarterly planning — the distribution you’re measuring moves under you as the model and your users change.
An interviewer says: “Our summarizer is 85% accurate on our test set. Do we ship?” What’s the strongest opening?
AYes — 85% is well above chance and clearly good enough for most usersBRefuse the flat number: ask what task, what the cost of a wrong output is, how the 15% is distributed, and what the human-fallback path is — then define a ship gateCNo — never ship below 95% accuracy on any AI feature
A PM writes the definition of done as “all eval cases pass.” Why does a senior reviewer push back for an LLM feature?
ABecause eval cases are too expensive to run on every buildBBecause “all pass” is binary, but the feature is a distribution — done should be a target rate on a scoped input class plus a graceful tail and a tripwireCBecause eval cases can’t test non-English inputs
For a new feature that detects whether an email is a meeting request, a regex + keyword baseline already scores 0.96 precision / 0.95 recall on your labelled set. What’s the senior call?
AShip the rule-based baseline; only reach for an LLM if you can show it beats those numbers on the same setBUse an LLM anyway — it will generalize better to future email formatsCUse an LLM because rules don’t scale
You’re standing up the first eval set for an AI assistant. A teammate wants to wait until you’ve collected 50,000 labelled examples. Better move?
AAgree — more data always means a more trustworthy evalBWait, but only for the regression set; the rest can come laterCStart with 10–20 examples per scenario covering distinct dimensions (incl. negative/“should-refuse” cases), then grow weekly from production traces
A model that succeeds 70% of the time per attempt passed your demo three times in a row. A stakeholder calls it “proven.” How do you frame the launch metric instead?
AReport pass@1 from the demo — three successes is a representative sampleBCharacterize the distribution: run each task many times, report pass@k across replications plus latency / cost-per-task / error rate, and gate on thoseCLower the temperature to 0 so the result is deterministic and re-run the demo once
The “defining success” round tests one thing: do you treat an AI feature as a distribution with a written, cost-bound definition of done — or as a magic box you eyeball? The canonical opener is “the model is 85% accurate — do you ship?” and graders reward candidates who refuse the flat number and build a layered, conditional ship decision. Lead with the mechanism (it’s a distribution), then the artifact (the eval set and the gate), then the business binding (cost of error).
01“The model is 85% accurate — do you ship?” → refuse the binary: ask the task, the cost of a wrong output, how the 15% is distributed, the human-fallback path — then write a ship gate.
02“How do you define done for an AI feature?” → Gothelf’s distributional sentence: 80% of inputs in category X meet bar Y, the other 20% degrades gracefully, plus a tripwire metric.
03“How do AI success criteria differ from traditional ones?” → traditional is a binary spec-match; AI is a distributional contract that can drift silently and expires as the model changes.
04“When should you NOT use AI here?” → when a rule-based baseline already hits the bar; AI must earn its place with a head-to-head win (PAIR’s unique-value test).
05“How big should the eval set be?” → start 10–20 per scenario (Notion), grow from traces; cover by dimension incl. negative cases — coverage beats volume.
06“pass@1 vs pass@k?” → pass@1 understates a retry-allowed system; pass@k = at least one of k attempts succeeds, with independent trials so variance is the model’s.
07“Your demo passed three times — proven?” → no; characterize the distribution with replicated runs and pass@k, not a 3-sample anecdote.
08“Why revisit the definition of done?” → frontier-model capability shifts the achievable distribution; Notion re-builds its evals roughly every six months.
To go deeper, expect the follow-ups that separate “read a blog” from “shipped one”: “how do you set the threshold?” (bind it to downstream cost — a refund vs a lawsuit — and stratify the test set by difficulty so 85%-overall isn’t hiding 0% on the hard 20%); “how would you know within 24 hours that a slice of users is being harmed?” (a divergence alarm and a stratified eval, the subject of L2 and L5); and “prove the AI beats the non-AI version” (a head-to-head A/B with a falsifiable hypothesis). In every case, name the metric and the gate before proposing a fix.
Could you, on a whiteboard, turn “the model is 85% accurate” into a distributional definition of done with a tripwire and a fallback?
New to itGetting thereConfident
Takeaways
An AI feature is a distribution, not a binary — done = a target rate on a scoped input class, a graceful tail, and a tripwire.
Eval-driven development: write the eval spec in the PRD before code; if two engineers can’t write the same eval, the requirement is too vague.
Earn the AI: if a rule-based baseline hits the bar, ship the rule (PAIR’s unique-value test).
Eval sets are shape-over-size: 10–20 per scenario, grow from traces, cover by dimension incl. negative cases.
Use pass@k with independent trials for stochastic, retry-allowed systems — not a 3-run demo.
The definition of done expires: revisit it as the frontier model shifts the achievable distribution.
Next: the metric stack — connecting model quality to product and business outcomes, and the metrics that mislead.