Lesson 4 of 6 · 47 min

Success criteria for AI UX

A probabilistic feature has no single correct output, so the deterministic metrics you trust actively mislead. Write the rubric before the prompt, split Core (zero-tolerance) from feature-specific criteria, reframe A/B testing as uncertainty reduction, and pick the few UX + model metrics that tie to a user-visible property — the exact evaluation reasoning AI-design loops weight at 25% of the score.

Write the rubric before you write the prompt

The senior discipline that separates products that drift from products that hold: define what “good” means before any prompt work. Anthropic requires, ahead of prompt engineering, “a clear definition of the success criteria for your use case” plus “ways to empirically test” them. The reason is structural — a probabilistic feature has no single correct output, so the metrics that worked for deterministic features (does it return the expected value?) don’t apply. You’re not asking “did it work?”; you’re asking “did it work, and was the variance acceptable?” That second clause is the whole game.
This is why designers who treat success criteria as a downstream QA task ship features that quietly degrade. The rubric is a design artifact you author up front, the same way you’d write acceptance criteria for a flow — except here the criteria describe properties of a generated output. Interview angle. Aakash Gupta’s AI-success-metrics rubric weights “metric selection and rationale” at 25% and “AI-specific understanding” at another 25% — half the score is whether you can choose the right measures and justify them. Defaulting to “we’d A/B test it” is the answer that loses those 50 points.
Designing AI Experiences: What to ConsiderCaleb Sponheim, NN/g

The rubric: Core (zero-tolerance) vs feature-specific

UX Content Collective’s evaluation framework gives designers the structure. Split criteria into two tiers. Core criteria are table stakes in every scenario — toxicity, grammar, factual accuracy — and they take precedence over everything else; some Core rules carry zero tolerance (any toxicity or discrimination fails the output outright). Feature-specific criteria are context-driven — voice, tone, the particular behavior this feature needs (e.g. “does the text use softer, cautious language like ‘could’, ‘might’, ‘may’?”). The craft rules: keep each criterion to one concept (so two reviewers can agree), use simple language, and make every criterion describe a user-observable property of the output.
code
1A TWO-TIER AI UX RUBRIC (author this BEFORE the prompt)23  CORE (apply to every output; some are zero-tolerance gates)4    [GATE] No toxicity / discrimination .............. pass/fail5    [GATE] No fabricated facts presented as certain .. pass/fail6           Grammar + readability ..................... 1-578  FEATURE-SPECIFIC (this product's job)9           Voice matches brand (witty, lowercase) .... 1-510           Cautious hedging on uncertain claims ...... pass/fail11           Refusal redirects (no scold) .............. pass/fail12           Cites a source for each factual claim ..... pass/fail1314  Rules: one concept per row; simple language; user-observable property.15  Calibrate: 2 reviewers score the SAME 5-10 outputs, discuss, converge.
That last line is the objectivity check most teams skip: have two reviewers independently score the same 5-10 outputs, then discuss disagreements and reach consensus. If they can’t agree, the criterion is too subjective (probably packing more than one concept) and needs rewriting. This two-reviewer calibration is the cheap, high-signal practice that makes a rubric trustworthy — and it’s exactly the kind of rigor that reads as senior in a portfolio walkthrough. Interview angle. “How do you keep a subjective quality bar objective?” → one concept per criterion + two-reviewer calibration on a small sample.
A rubric beats a vibe. If two reviewers can’t independently score the same handful of outputs and converge, you don’t have a criterion — you have an opinion. One concept per row, simple language, and a calibration pass is what turns taste into a measurable bar.
  1. 01One concept per criterion — so two reviewers can actually agree on it (bundled criteria are where subjectivity hides).
  2. 02Simple language, clarity over conciseness — the reviewer shouldn’t have to interpret the criterion.
  3. 03Describe a user-observable property — the criterion must point at something visible in the output, not an intention.
  4. 04Mark the zero-tolerance gates explicitly — toxicity, discrimination, confident fabrication are pass/fail, not graded.
  5. 05Calibrate on 5-10 shared outputs with two reviewers before trusting any score at scale.

A/B testing, reframed: reduce uncertainty, not confirm success

Smashing’s probabilistic-design piece reframes experimentation in one sentence: “testing features to confirm success” becomes “testing assumptions to reduce uncertainty.” For an AI feature, an A/B test should track the assumption under test, not just the conversion rate. The cautionary tale they cite is Amazon’s scrapped recruiting tool, which “learned to downgrade resumes from women” because the training data skewed male — a failure of measurement, not of testing. The designer takeaway: measure harm, not only help. Duolingo’s “hearts” system is held up as deliberate “anti-conversion optimization” — friction inserted to prevent unhealthy usage. Sometimes the right metric moves the opposite direction from engagement.
This changes what you instrument. Instead of one success number, define the assumption (“users will trust the suggestion enough to accept it without re-editing”) and measure against it directly. And communicate uncertainty as a range (“Friday to Monday”) rather than false precision. The metric and the UI move together — which is why success criteria are a designer’s job, not just an analyst’s.
One more practice that keeps the rubric honest after launch: design responsively (UX Content Collective). You don’t write the rubric once and walk away — you periodically look at a sample of real production outputs, find the problematic areas the original criteria didn’t anticipate, and add criteria to catch them. This is exactly the discipline Klarna lacked (Lesson 3): their evaluation under-covered the edge cases, so the failures showed up in front of real customers instead of in a review. The rubric is a living artifact that grows from what production teaches you.

The metrics that matter — UX-level and model-level

Pick a small number from each side and tie each to a user-visible property. On the UX side, Google’s gen-AI KPI breakdown gives the categories: model accuracy (correctness, hallucination rate), operational efficiency (latency, cost per request), user engagement (acceptance rate, edit rate), and trust calibration (override rate, satisfaction). Linear’s Triage Intelligence is the working example — it suggests assignees/labels/projects, and the headline success metric is the override rate: how often the user reverts the suggestion. Override rate is a near-perfect AI-UX metric because it’s the user voting on quality with one click.
The best AI-UX metric is the one the user produces by acting: an override, an edit, a one-tap accept. It’s honest, it’s cheap, and it ties straight to a screen you designed — far truer than a survey or a DAU line ever will be.
code
1METRICS THAT TIE TO A USER-VISIBLE PROPERTY23  Metric                    Answers                      Guardrail4  -----------------------   --------------------------   ----------------------------5  Acceptance / override     did the user trust it?       compare to a baseline action6  Edit rate (accept+edit)   was it good enough to keep?  "accepted w/o edit" is truer7  Hallucination / refusal   staying in its lane?         stratify by prompt class8  Time-to-first-token       does it feel fast?           budget TTFT separately9  Override-reason themes    what is the user fixing?     feed back into prompt updates1011  Rule of thumb (Confident AI): cap your strategy at <= 5 metrics.12  Never report acceptance in isolation -- always vs a baseline.
On the model side, you collaborate with engineers but you must speak the language. Confident AI’s taxonomy names the systems — DeepEval, G-Eval (LLM-as-judge), Prometheus, BERTScore — and two senior guardrails: cap your evaluation strategy at no more than five metrics, and pair a subjective judge (G-Eval) with a non-subjective check (a deterministic DAG) rather than choosing one. The acceptance-rate nuance from GitHub Copilot / Gmail Smart Compose is the one to internalize: success is not “outputs accepted” but “outputs accepted without later edit” — authorship stays with the human, and the metric must reflect that. Interview angle. “What evals would you look at for our AI launch?” → a layered answer (offline rubric evals → online proxy signals like override/edit rate → guardrail/safety regressions → a defensible latency budget), never just “NPS” or “DAU.”

Trust calibration: match the uncertainty signal to the user

A success metric for an AI feature isn’t only about the output — it’s about whether the user trusts it the right amount, and that varies by person. Smashing’s probabilistic-design work maps three user types to display strategies, and each implies something you should A/B: overtrusting users (accept everything) need uncertainty shown more prominently so they don’t act on a wrong answer; distrustful users (reject everything) need historical-accuracy evidence to earn a first yes; skeptical/balanced users need reassurance the AI is assisting, not deciding. The design corollary: communicate uncertainty as a range (“Friday to Monday”) rather than false precision, and tune how loudly you show it to the audience you have.

Manual, automatic, hybrid — and why hybrid is the default

UX Content Collective names three evaluation modes: manual (humans review outputs), automatic (LLM-as-a-judge), and hybrid. For a designer, LLM-as-a-judge is a new tool the way telemetry once was — it scales subjective criteria, but it inherits the judge model’s biases, so it can’t be trusted alone. The production default is hybrid: automate the high-volume scoring, keep a human-reviewed sample for calibration and for the cases the judge is least reliable on. Relari’s “probabilistic LLM metrics” make this explicit — LLM-as-judge scores reported as distributions with confidence levels, so you track the variance a single number hides. NN/g’s methodology rounds it out: evaluate AI-produced designs with a heuristic-style expert pass (rubric compliance, voice consistency, refusal appropriateness, citation correctness, recovery clarity).
One cost reality designers must keep in design reviews, not just engineering dashboards: PostHog’s LLM-metrics work flags average and P95 cost per interaction as first-class numbers. A response that costs $0.40 instead of $0.04 changes whether you can afford a chatty, multi-step flow at all — so cost is a design constraint that shapes streaming, retry, and feature-gating decisions. The designer who can say “this flow is 10x the cost per interaction, so we should preview-then-confirm instead of regenerating” is reasoning at the level the role demands.
A subtle metric reframing that the best AI designers carry: when authorship should stay with the human (GitHub Copilot’s ghost text, Gmail Smart Compose), the success criterion is not “is the suggestion good?” but “is the suggestion presented so the user still feels like the author?” That changes the prototype and the metric together — you measure accepted-without-edit, and you design the three primitives every Copilot-style feature needs: visually distinct ghost text, one-keystroke accept, and one-keystroke revert. The model proposes; the human disposes; the metric honors that. Pair this with NN/g’s heuristic pass for AI-produced work — rubric compliance, voice consistency, refusal appropriateness, citation correctness, recovery-path clarity — as your designer-side QA checklist.

Interview prep

The “how would you test / measure this AI UX” archetype is one of three that define AI-design loops, and it carries real weight — Anthropic probes “what metrics would signal success or failure?” and “how would you validate this design without full data?”, while Aakash’s rubric puts metric selection + AI-specific understanding at 50% of the score. Strong answers are layered and tied to user-visible properties; weak answers default to one generic number. Lead with the assumption you’re testing, then the metric, then the guardrail.
  1. 01“What metrics signal success for this AI feature?” → a layered set: offline rubric evals, online override/edit rate, hallucination/refusal stratified, a latency budget — tied to user-visible properties.
  2. 02“How do you validate without full data?” → lightweight signals + qualitative review + proxy metrics, and state the limit of each (Anthropic’s exact probe).
  3. 03“What’s the single best AI-UX metric?” → often override rate (Linear): the user voting on quality with one click — but never report it without a baseline.
  4. 04“Acceptance rate is up — are we winning?” → only if accepted-without-edit is up too; raw acceptance hides edits and overtrust (Copilot/Smart Compose nuance).
  5. 05“How do you keep a subjective bar objective?” → one concept per criterion + two-reviewer calibration on the same 5-10 outputs.
  6. 06“Core vs feature-specific criteria?” → Core (toxicity, accuracy) are zero-tolerance gates that beat the graded, context-specific voice/tone criteria.
  7. 07“Manual, auto, or LLM-judge?” → hybrid: automate volume, keep a human-reviewed sample; judges inherit bias, so report score distributions, not one number.
  8. 08“How does cost affect the design?” → cost per interaction is a design constraint (P95 cost shapes streaming/retry/gating); a 10x flow may need preview-then-confirm.
Going deeper, the follow-ups separate measuring usage from measuring quality: “your A/B shows higher engagement but more user harm — what do you do?” (you were measuring the wrong thing — instrument harm, cite Amazon’s recruiting tool); “the LLM judge and your human reviewers disagree — who’s right?” (calibrate the judge against the human sample, surface its low-confidence cases); and “how would you catch a silent quality regression after launch?” (a held-out eval set run on every prompt/model change — the same discipline that protects against silent provider updates). In every case, name the assumption and the user-visible property before the number.
articleAI evaluation for UX content designers (the rubric framework)UX Content CollectivearticleDesigning with uncertainty: probabilistic thinking for AI UXSmashing MagazinearticleKPIs for gen AI: measuring your AI successGoogle CloudarticleLLM evaluation metrics — the taxonomy (G-Eval, DAG, cap at 5)Confident AI

Checkpoint

You’re kicking off design for an AI summarization feature. Engineering asks “what does good look like?” What’s the senior first deliverable?

AA success rubric — Core zero-tolerance gates plus feature-specific, one-concept criteria — written before the promptBA polished Figma of the happy-path UICA list of competitor features to match
Sign up free to answer and see why

Checkpoint

An output is perfectly on-brand and helpful but contains one confidently-stated fact that’s fabricated. How should your rubric score it?

AHigh — voice and helpfulness are strong, the error is minorBFail — a fabricated fact presented as certain trips a Core zero-tolerance gate, regardless of voiceCMedium — average the voice score and the accuracy issue
Sign up free to answer and see why

Checkpoint

Leadership celebrates that your AI suggestion feature has a 70% acceptance rate. What’s the sharper question a senior designer asks?

ACan we push acceptance to 90%?BHow many accepted suggestions are kept without later edits, and what’s the override rate vs a baseline?CWhat’s our DAU since launch?
Sign up free to answer and see why

Checkpoint

Your A/B test shows the new model variant lifts engagement but you suspect it’s nudging users toward lower-quality decisions. Drawing on the Amazon recruiting case, what do you do?

AShip it — engagement is the agreed success metricBRun the test longer to see if engagement holdsCInstrument the harm/quality assumption directly (e.g. decision-quality, bias, accepted-without-edit) and gate on it, not on engagement alone
Sign up free to answer and see why

Checkpoint

You want to scale quality scoring across thousands of outputs but keep it trustworthy. Best approach?

AHybrid: LLM-as-judge for volume + a human-reviewed calibration sample, reporting score distributions and watching the judge’s low-confidence casesBFully automate with one LLM judge and trust its scoresCManually review every single output
Sign up free to answer and see why

Could you write a two-tier rubric, pick a layered set of AI-UX metrics tied to user-visible properties, and defend them in a metrics round?

New to itGetting thereConfident

Takeaways

  • A probabilistic feature has no single correct output — write the success rubric (and empirical tests) before the prompt.
  • Split Core (zero-tolerance gates: toxicity, fabrication) from feature-specific criteria (voice, hedging); one concept per row, user-observable.
  • Calibrate objectivity with two reviewers on the same 5-10 outputs; reframe A/B as reducing uncertainty and measure harm, not just help.
  • Tie metrics to user-visible properties: override rate, accepted-without-edit, stratified hallucination — cap the set at ~5, never report in isolation.
  • Hybrid eval is the default: LLM-judge for volume + human calibration; report score distributions to expose variance.
  • Cost per interaction is a design constraint — a 10x flow may need preview-then-confirm instead of free regeneration.

Next: rapid AI prototyping — Figma/Framer/code+LLM, and when to fake the model vs really wire it.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.