RICE still applies, but AI forces three new modifiers: confidence (eval-backed, not gut), coverage (fraction of inputs handled gracefully), and cost-of-error. The input/output metric split, eval saturation as the ship threshold, the ship-posture matrix, dynamic model routing as a prioritization move, and the discipline of pruning infinite surface area.
Confidence is now a number, not a gut call
Classic ICE/RICE still applies under uncertainty — but AI forces three extra modifiers that change the math: confidence (eval-backed pass-rate on real-user traces, not a gut call), coverage (the fraction of user inputs the feature handles gracefully), and cost-of-error (the worst-case harm and its remediation cost). The senior reframing: in a world where the model can produce infinite surface area, prioritization is mostly about pruning — saying no — and the discipline that lets the eval discipline work. This lesson is how you bet on probabilistic features and sequence them under risk.
code
1RICE -> AI-ADJUSTED RICE23 Factor Classic AI-adjusted4 -------------- --------------------- --------------------------------------5 Reach users affected users x distributions reliably covered6 Impact 1-3 discretionary 1-3 x confidence-weight7 Confidence gut call eval pass-rate on real-user traces8 Effort eng units eng x model-cost/call x tail-risk fix9 Cost-of-error (implicit) EXPLICIT: liability + recovery cost (new)1011 The new row is the point: a feature with huge reach and a catastrophic tail12 scores LOWER than a smaller feature whose cost-of-error is bounded.
The Redpoint constraint makes the pruning concrete: resist the flood of feature requests and prioritize shipping critical products over secondary features like dashboards. OpenAI’s internal rule is the same insight from a different angle — “North Star” critical needs first. Tome narrows further: prioritize prosumers as enterprise ambassadors and “do not build dashboards.” When the model can generate any surface, the scarce resource is focus, and saying no is the highest-leverage prioritization act. Interview angle. “What would you cut?” is a near-certain follow-up in any AI case — the senior answer names the secondary surface you’re deliberately not building (and why), not a grudging trim.
The coverage modifier deserves its own beat because it’s the one teams forget. Coverage is the fraction of real user inputs the feature handles gracefully — not the fraction it answers, but the fraction it either answers well or fails well (clarifies, abstains, hands off). A feature with 99% accuracy on the 40% of queries it attempts and a confident wrong answer on the other 60% has terrible coverage and will feel broken, while a feature that handles 95% of inputs — some by answering, some by gracefully deferring — feels reliable. This is why narrow-but-graceful beats broad-but-brittle in early AI products: you’d rather scope the feature to a slice you cover well and expand, than ship wide and let the uncovered tail define the user’s impression.
Aakash Gupta’s framing is now the default opening frame for any “how would you measure success” question, and candidates who can’t name both buckets lose the room: input metrics measure what users DO; output metrics measure what the BUSINESS gets. Split the answer cleanly into these two before you name a single number. Input metrics (thumbs, retry/regenerate rate, edit-after-suggestion rate, task completion logged) are the Day-1, low-latency signals that reflect model quality. Output metrics (task success rate, deflection/containment, retention of AI-feature users, ARPU lift, support cost saved) are the slower, less-noisy business impact.
code
1INPUT vs OUTPUT METRICS (split BEFORE naming a number)23 Input metrics (what users DO) Output metrics (what BUSINESS gets)4 ----------------------------------- -----------------------------------------5 thumbs up/down task success rate6 retry / "regenerate" clicks deflection / containment rate7 edit-after-suggestion rate retention of AI-feature users8 task completion logged ARPU lift, support cost saved9 -> Day-1 signal, reflects quality -> business impact, slower + less noisy1011 Weak answer conflates "model accuracy" with "user value." Strong answer names12 one input metric AND one output metric, then a guardrail (kill-switch).
The trap to avoid is conflating model accuracy with user value — a model can be 95% accurate and still move no business metric if the 5% lands on the queries that matter, or if users don’t adopt it. So a complete answer adds a third element: a guardrail metric (a kill-switch threshold — e.g. hallucination rate, false-positive rate on a safety filter, p95 latency) that can halt a ramp regardless of how good the headline metric looks. One input, one output, one guardrail: that triad is the senior shape.
Input metrics measure what users DO; output metrics measure what the BUSINESS gets — split any “how would you measure success” answer into those two buckets before you name a single number. — Aakash Gupta, AI Product Success Metrics
The ship threshold: eval saturation, not 100%
The single hardest judgment in an AI prioritization case is when is a probabilistic feature good enough to ship? Anthropic’s published answer is the cleanest: ship on eval saturation — the moment the agent passes all solvable tasks — not when accuracy hits 100%. The complement is evaluator design: code-based graders (verifiable, atomic checks), model-based graders (soft attributes like tone/coverage), and human graders (the final bar), with an explicit warning to avoid rigid grading that penalizes minor formatting and to use partial credit for multi-component tasks so success is a continuum, not a brittle pass/fail.
Reconcile the two voices you’ll hear in interviews: Redpoint says “just do it” (ship fast for feedback); Anthropic says ship on eval saturation. The contradiction dissolves once you see Redpoint’s “just do it” assumes you have an eval at all — what you skip is product polish, not the eval. Perplexity’s deliberate start as a “wrapper” to quickly gain user feedback is the same move: treat the LLM as a black box, concentrate product sense on the shell, and let real usage drive the eval. Interview angle. “When would you ship this?” → “when it saturates the eval on the distribution I care about, with a guardrail in place” beats any single accuracy number.
Error budgets for AI — borrowing the SRE lens
Because a probabilistic product can never be perfect, the SRE concept of an error budget ports cleanly and gives prioritization a quantitative spine. Instead of chasing zero errors, you allocate an acceptable failure rate and spend it deliberately across three AI-specific dimensions: hallucination (quality errors), latency (p95/p99 budget), and spend (cost per query). The budget reframes the roadmap: as long as you’re inside budget, ship features and take risk; when you blow the budget, work stops on new features and goes to reliability. This is the mechanism behind a guardrail kill-switch — the kill-switch fires when a ramp would exhaust the error budget on the dimension that matters.
code
1AI ERROR BUDGET -- allocate failure, don't chase zero23 Dimension Budget example When it's spent...4 ------------- ---------------------- --------------------------------5 Hallucination <2% on the gold set halt ramp; the kill-switch fires6 Latency p95 < 2.5s shed load / route to faster model7 Spend < $0.04 / query downshift model tier; cap regen89 Inside budget -> ship features, take risk.10 Budget blown -> freeze features, fix reliability (the SRE rule).1112 This is why "guardrail metric" and "kill-switch" are the same idea:13 the switch fires the moment a ramp would exhaust the budget.
The ship-posture matrix — pick one and defend it
Once the eval saturates, you still choose a launch posture that decides which trade-off you accept. The five documented postures map to five companies, and the senior move in a case is to pick one explicitly and name its cost — not to gesture at “we’ll be careful.”
code
1SHIP-POSTURE MATRIX -- pick one, name the trade-off23 Posture When to use it Trade-off accepted4 ---------------------- ------------------------ -------------------------5 Beta + transparency high learning rate, some users over-trust the6 (Notion) forgiving user "beta" label anyway7 Citation-required domain where authority coverage drops; facts must8 (Perplexity) matters chase each other9 Indemnified commercial/B2B, price higher cost basis; liability10 (Adobe Firefly) absorbs risk paid in cash11 Trust-layer wrapped platform, many use cases layer's defaults become a12 (Salesforce) on one trust layer product constraint13 Refusal-as-feature high-stakes, asymmetric coverage falls; UX loses on14 (Anthropic agents) downside non-stakes questions
Intercom is the prioritization mechanism for probabilistic services: it dynamically routes between GPT-4 and GPT-3.5 based on the scenario — the cheaper model where the cost-of-error is low, the expensive model where it isn’t. This is ICE/RICE applied to model selection per request, and it’s the textbook AI cost/quality move: you don’t pick one model globally, you route. Pair this with the cost-of-error row of AI-adjusted RICE and you have a defensible sequencing story — ship the high-confidence, bounded-cost surfaces first; route expensive models only where the tail justifies them.
Interview prep
Prioritization-and-metrics rounds reward the input/output split, an eval-backed confidence story, and an explicit ship posture with its trade-off. Lead with the framework, then a number.
01“How would you measure success of this AI feature?” → split first: one input metric (what users do), one output metric (what the business gets), plus a guardrail kill-switch.
02“When is it good enough to ship?” → at eval saturation (passes all solvable tasks) on the slice I care about, with a guardrail wired — not at 100% accuracy.
03“How would you prioritize a RAG feature vs a fine-tune vs a UI rewrite?” → AI-adjusted RICE; weight by eval-backed confidence and the cost-of-error row, not gut impact.
04“What would you cut?” → the secondary surface (dashboards, settings) — Redpoint/Tome’s “do not build dashboards”; pruning is the prioritization act.
05“Ship fast or get it right?” → both: Redpoint’s “just do it” assumes you have an eval; you skip polish, not the eval gate.
06“How do you decide which model to use?” → route per scenario (Intercom GPT-4/3.5) by cost-of-error; cheap model where a miss is cheap, expensive where it isn’t.
07“What’s your guardrail metric?” → a kill-switch threshold (hallucination rate, safety false-positive rate, p95 latency) that halts the ramp regardless of the headline metric.
08“How do you sequence a risky AI roadmap?” → high-confidence, bounded-cost surfaces first; reserve expensive models / autonomy for where the tail justifies them.
Follow-ups probe the numbers and the failure modes. Expect “what confidence interval are you shipping on?” (name eval pass-rate + sample size, e.g. >=90% on a 10k-trace set), “what happens on day 30 if the input metric is up but the output metric is flat?” (adoption-without-value — investigate whether the 5% miss lands on high-value queries), and “how do you avoid eval-set drift?” (a fixed refresh cadence — monthly for high-velocity, quarterly otherwise — owned by a named person). The meta-point per Exponent (2026): naming a metric, a failure mode, and a user beats any framework recited cleanly.
You’re ranking three AI bets with RICE. Feature A: huge reach, but a wrong answer could expose customer financial data. Feature B: modest reach, bounded cost-of-error, eval pass-rate 92%. How should AI-adjusted RICE treat A vs B?
AA wins — reach dominates the RICE scoreBThey tie — cost-of-error isn’t part of RICECB can rank higher — the explicit cost-of-error row down-weights A’s catastrophic, reach-multiplied tail, while B has bounded harm and eval-backed confidence
An exec asks “what accuracy do we need before we launch the assistant?” Best senior reframing?
A“We’ll launch at 99% accuracy to be safe.”B“We launch at eval saturation — when it passes every solvable task on the slice we care about — with a guardrail kill-switch wired and a forgiving segment first.”C“We’ll know it’s ready when the demo feels good.”
A support-bot PRD lists the success metric as “90% model accuracy on our test set.” What’s the biggest gap?
AThe accuracy target is too low; it should be 95%BIt conflates model accuracy with user value and names no business outcome — it needs an input metric, an output metric (e.g. deflection), and a guardrailCIt should specify which model and provider are used
Intercom routes between GPT-4 and GPT-3.5 per scenario. What prioritization principle does this encode?
AAlways use the most capable model for consistencyBPick one model globally and optimize prompts around itCApply ICE/RICE to model selection per request — route the cheap model where the cost-of-error is low, the expensive model where it isn’t
Your team wants to build an analytics dashboard for the new AI feature before the core flow is solid. Per Redpoint/Tome, the senior call is…
ABuild the dashboard first — you can’t manage what you can’t measureBBuild both in parallel to avoid reworkCPrune it — ship the critical core product first; “do not build dashboards” is the canonical example of resisting the feature flood