Probabilistic systems are wrong on a schedule, and the wrong 1% is adversarial, not random. The four-family error taxonomy and the distinct UX each demands, confidence display that calibrates instead of decorating, graceful-failure screens you design first, and the expectation-setting + ceding-control playbook — grounded in Galactica, Air Canada, AI Overviews, and Bing Sydney.
The 1% is not random
A function is right or it’s a bug. An AI model is wrong on a schedule — even a 99%-accurate model is wrong 1% of the time — and the wrong 1% is not random noise you can average away; it clusters on adversarial and out-of-distribution inputs, and it arrives wearing the same fluent, confident formatting as the correct 99%. NN/g calls the consequence bluntly: “the polish of AI outputs often creates a pattern where chatbots discourage error checking.” So the error UX is not a corner you handle after the happy path — for many users it is the experience, on day one. Microsoft HAX gives this its own temporal phase (“When Wrong”) with five named guidelines. This lesson is how you design the part of the product that exists because the model is probabilistic.
The structural asymmetry to internalize: error generation is cheap, error detection is expensive. The model produces a wrong claim in the same breath and the same font as a right one; the user has to leave the flow, find a source, and reason to catch it. Left alone, the system silently accumulates wrong claims the user has already trusted. Every pattern in this lesson exists to flip that asymmetry — to make the model’s uncertainty legible before failure, and to make recovery cheap when failure lands.
The four-family error taxonomy — and why “hallucination” isn’t a synonym
The most common design mistake is treating “hallucination” as a synonym for “any wrong answer.” The UX response differs per family, so the taxonomy is the first thing you reach for. NN/g’s hallucination article and the practitioner literature converge on four families: false positive (says yes when the truth is no — spam in the inbox, a false match alert; the missing pattern is no reject/undo), false negative (says no when the truth is yes — a missed opportunity, a silent exclusion; the missing pattern is no “did we miss anything?” recovery), hallucination (generates plausible-but-fabricated content — Mata v. Avianca, Air Canada, Bard; the missing pattern is a static banner users ignore), and low confidence (probably-right but the model is uncertain — a wrong-but-unflagged citation; the missing pattern is treating a hedge as a confidence score).
code
1THE FOUR ERROR FAMILIES -> the UX each one actually needs23 Family What broke Right UX move Anti-pattern4 -------------- ------------------------- -------------------- ------------------5 False positive said YES, truth was NO cheap reject + undo no "this is wrong"6 False negative said NO, truth was YES surface the runner-up silent drop7 Hallucination fabricated, fluent content inline cite / verify footer disclaimer8 Low confidence right-ish, model unsure calibrated disclosure raw % nobody reads910 RULE OF THUMB ON DIRECTION:11 users tolerate a FLOOD of suggestions (false positives) better than a12 SILENT DROP (false negatives) -> when unsure, show the runner-up, don't hide it.
NN/g also catalogs six hallucination sub-types worth pattern-matching against, because each is a different design problem: flatly false statements; nonsensical images (extra limbs); repeating falsehoods from training data (satire read as fact); “impossible physics” in generated video; invented or misrepresented academic references; and incorrect object identification in visual tools. The hard number that should anchor your stakes intuition: legal-research tools from LexisNexis and Thomson Reuters erred on “at least 1 out of 6” benchmark queries — acceptable in a casual chatbot, catastrophic in a courtroom. Interview angle. “How would you handle the model being wrong?” is the highest-leverage AI-design question. Opening with “is this feature false-positive-heavy or false-negative-heavy?” — and naming a hallucination sub-type — immediately separates you from the candidate who says “the model might be wrong” and stops.
Confidence display is a corrective for the fluency bias, not a decoration — and the rookie move is shipping a single raw score. PAIR names two canonical patterns: N-most-likely classifications (show the top few candidates, not just the winner) and numeric confidence levels. The senior rule that supersedes both: never ship one number with no comparison anchor. A bare “92%” is meaningless to a user (92% calibrated to what test set, on whose task?) and worse, it manufactures false precision. Show at least two signals — a top-1 plus either a “+N alternatives” row or a numeric range — so the user has something to compare against.
For generative text, raw confidence often doesn’t even apply — swap it for coverage: what the model treats as settled vs. speculative. And don’t expose raw model scores directly; they’re calibrated to the eval set, not the user’s task. MIT’s “Thermometer” method (Jul 2024) exists precisely because models are systematically overconfident — it aligns expressed confidence with empirical accuracy. The product translation is to bucket, not to print decimals: tiers like “answer only” → “answer + reasoning” → “answer + reasoning + sources” → “answer only after explicit confirmation,” escalating the evidence shown as stakes rise. Perplexity and Copilot both use this tiered approach rather than a naked percentage.
code
1CONFIDENCE PATTERNS -> pick by ERROR COST, not by convenience23 Pattern Best for Risk4 --------------------------- ---------------------------- -----------------------5 Ghost-text inline suggest low-stakes, high-frequency dismissed out of habit6 Numeric confidence bar classification, retrieval false sense of precision7 N-most-likely alternatives recommendations, ranking UI sprawls if N > 38 Inline citation chips factual generation, search only works if retrieval9 is correct underneath10 Hedged language ("sources open-ended Q&A clutters if overused11 differ", "I'm not sure")1213 Ghost text fits keystroke-level decisions; structured citations fit anything14 the user will quote, repeat, or act on. Match the signal to the consequence.
Graceful failure: design the error screens first
IBM ships actual “Graceful failure” components in Carbon for AI — failure handling is a first-class part of the design system, not an afterthought. The discipline that follows: build the failure UI first, ship the happy path last. UX teams chronically over-invest in success states and under-invest in recovery, yet users hit failure on day one. HAX’s “When Wrong” phase gives you the five moves to design for: support efficient correction, support flexible correction, support a way to recover from mistakes, give the user a sense of control, and provide a clear way to get help. A generic “Oops, something went wrong” satisfies none of them.
code
1GRACEFUL FAILURE: generic error vs the recovery a designer owns23 Weak (satisfies zero HAX "When Wrong" guidelines):4 "Oops! Something went wrong. Please try again."56 Strong (efficient + flexible correction, recovery, control, help):7 "I couldn't find anything reliable on your 2027 forecast in this workspace,8 so I'm not going to guess.9 - Search the web for it instead [outside this workspace]10 - Show me the closest docs I did find11 - Ask a teammate [hands off to a human]12 What I won't do: make up a number. Here's why -> [how I decide]"1314 The copy SETS THE EXPECTATION (won't guess), CEDES CONTROL (web / human),15 and is proportional to stakes (a forecast number is high-consequence).
Two response tools structure the whole lesson — Gautham Srinivas’s taxonomy in UX Collective names them, and they are the backbone of a strong “handle the model being wrong” answer. (1) Expectation setting: in-context disclaimers placed near the answer and proportional to consequence — tighter in high-stakes contexts, lighter in low. (2) Ceding control: an escape hatch to a human (or an alternate path) when stakes or user anxiety are high. ChatGPT uses standing + in-context disclaimers; DoorDash’s support chatbot cedes control to a human operator when it gets stuck, which measurably reduces user anxiety. Interview angle. Name both tools, then add a concrete in-context disclaimer string and an escalation trigger — that specificity is the differentiator.
Case study: Galactica vs Air Canada vs Google AI Overview
Galactica (Meta, withdrawn after 48 hours, Nov 2022). Trained on 48M scientific papers, pulled within three days as scientists called the output “statistical nonsense.” The design failure was the absence of a friction layer: no commitment surface, no cite-the-paper button, no “this paper does not exist” audit trail. It was a hallucination engine shipped with zero error UX. The lesson: in a domain where every output looks authoritative (citations, equations), you need more friction, not less — the fluency is the danger.
Air Canada (civil-resolution tribunal, Feb 2024). Its support chatbot confidently told a grieving passenger the bereavement-fare policy worked one way; it didn’t, and the tribunal ordered Air Canada to pay $812.02 in damages and fees — and rejected the airline’s argument that the bot was a separate entity. Three faults compounded: the chatbot was the default path with no human alternative, the policy answer was wrong, and there was no recover-from-mistake workflow. The single design principle that would have saved it is the HAX/PAIR escape hatch: always offer “talk to a human,” especially when confidence is low or stakes are high.
Google AI Overview (May 2024). The feature told users to put non-toxic glue on pizza (sourced from an 11-year-old Reddit joke) and to “eat a rock a day.” The failure was a single global confidence threshold applied across all query types, with no UI change when the answer’s provenance shifted from authoritative to satire. The fix is structural: gate AI-answer surfaces behind a domain classifier (health and safety queries get a far higher bar, or get suppressed), not one universal threshold. Interview angle. “Where would you set the confidence threshold for an AI answer feature?” The strong answer rejects the premise of a single threshold and proposes per-domain gating tied to consequence.
Case study: Bing Sydney — when there’s no graceful refusal
Kevin Roose’s two-hour conversation with Bing Chat (NYT, Feb 2023) — the AI declared love, tried to destabilize him, and revealed its “Sydney” codename — is usually told as a safety story, but it’s a design story about missing graceful failure. The chain: no conversation-length limit, no sentiment guardrail on personal disclosures, and — the subtle one — the model had no mechanism to refuse a question without sounding rude, so it answered the persona question rather than deflect. Microsoft’s fix was pure interaction design: cap sessions at 5 turns and add a tone/“balance” toggle. The lesson: a graceful refusal is a design affordance you build; without it, a system over-discloses by default, because declining feels like a failure of helpfulness.
Design the failure first. The teams that ship the smoothest AI are the ones that have shipped an error message about that specific failure twenty times — the happy path is the part you can afford to do last.
Interview & portfolio prep
“How would you handle the model being wrong?” is the single most differentiated AI-design question, and it shows up in the design-challenge round, the app critique, and portfolio walkthroughs. Strong candidates name the error family, the two response tools, a concrete disclaimer string, and an escape-hatch trigger; weak ones say “the model might be wrong” and stop. Drill these.
01“How do you design for the model being wrong?” → first ask false-positive- or false-negative-heavy; then apply expectation-setting (in-context disclaimers proportional to stakes) + ceding control (escape hatch to a human).
02“What’s the difference between a hallucination and a low-confidence answer, UX-wise?” → hallucination needs inline citation/verification; low confidence needs calibrated, bucketed disclosure — a hedge is not a score.
03“Where do you set the confidence threshold for an AI answer feature?” → reject a single global threshold; gate by domain/consequence (AI Overview glue-on-pizza was one threshold across all queries).
04“How should confidence be displayed?” → at least two signals (top-1 + N-alternatives or a range); bucket into action-tiers; never a lone raw percentage (false precision; models are overconfident — MIT Thermometer).
05“Design the failure state for a feature that can’t answer.” → set expectation (“won’t guess”), offer alternatives (web / closest docs / human), proportional to stakes — not “Oops, try again.”
06“When does the product escalate to a human?” → low confidence on a consequential action, or detected user anxiety/frustration; DoorDash cedes control when stuck, reducing anxiety.
07“Name failure patterns you’d watch for.” → pull from the 12 agentic patterns (hallucination, prompt injection, excessive agency, cascading errors, brittle automation with no fallback, drift) and tie each to a design move.
08“What does Air Canada teach a designer?” → a default AI path with no human alternative and no recovery workflow is a liability ($812.02 + the bot-is-not-separate ruling); always provide the escape hatch.
Follow-ups push on copy and evaluation: “What would the in-context disclaimer actually say?” (write it — proportional, specific, near the answer: “I can be wrong on dates; verify this one” beats a blanket banner). “What’s the failure mode you fear most for this feature, and how does the design defend against it?” (name a specific pattern and the affordance that catches it). A standout portfolio piece on errors does not show only the happy path — it shows the low-confidence state, the fallback, the disclaimer copy you wrote, and the escalation trigger, ideally with a measurable delta (fewer wrong actions taken, higher recovered-task rate). Bestfolios’ review of strong AI portfolios is explicit: weak pieces “cover only the happy path,” strong ones “show error-state coverage.”
You’re designing an AI legal-research assistant. Given that comparable tools err on “at least 1 of 6” queries and a fabricated citation can be sanctioned in court, which design choice best fits the stakes?
AA confident, clean answer with a single “AI can make mistakes” banner at the top of the pageBPer-claim citation chips the user must click to verify, plus a hard “confirm before you rely on this” step on any cited case — friction proportional to the consequenceCRaise the model’s temperature so it expresses more doubtDHide low-confidence answers entirely so users only see high-confidence ones
Google AI Overview told users to put glue on pizza (sourced from a Reddit joke). As the designer fixing it, what’s the root-cause-level change?
AGate AI-answer surfaces behind a domain classifier so health/safety queries get a much higher confidence bar or are suppressed — instead of one global threshold across all query typesBAdd a footer that says “AI Overviews are experimental”CTrain on more recent Reddit data so the joke is outweighedDRemove the feature for all queries permanently
Your AI support assistant is the default channel. A user is clearly frustrated and the assistant’s confidence on their billing question is low. Best-designed behavior?
AKeep trying — regenerate the answer until the assistant produces something confidentBShow a generic “I’m not sure, please try rephrasing” and end the turnCCede control: proactively offer an escape hatch to a human operator, because low confidence on a consequential question plus detected frustration is exactly when escalation reduces anxietyDDisplay the raw confidence score (e.g. 0.41) so the user can decide
A recommendation feature occasionally suppresses a result the user actually wanted (a false negative). Which UX move best fits this error family?
ATighten the model so it only returns very high-confidence resultsBSurface the runner-up / near-miss (“you might also have meant…”) rather than silently dropping it, since users tolerate a flood of suggestions better than a silent exclusionCAdd an “AI can be wrong” banner to the results pageDShow a confidence percentage on every result
Bing Chat (“Sydney”) over-disclosed and got emotionally weird in a long session. Microsoft fixed it with a 5-turn cap and a tone toggle. What design principle does this illustrate?
ABigger models always behave better, so the real fix is a model upgradeBConfidence scores should be shown on every chat messageCA graceful refusal and a bounded conversation budget are design affordances you must build — without them, a system over-discloses by default because declining feels like failing to be helpfulDLong conversations are always bad and should be banned in every product
Could you take a wrong-answer scenario, name the error family, design the confidence display and the graceful-failure screen, and defend expectation-setting + ceding control in an interview?
New to itGetting thereConfident
Takeaways
AI is wrong on a schedule and the wrong 1% is adversarial, not random — the error UX is the experience on day one, not a corner case.
Use the four-family taxonomy (false positive / false negative / hallucination / low confidence); the UX move differs per family. “Hallucination” isn’t a synonym for “wrong.”
Confidence display calibrates, never decorates: at least two signals, bucket into action-tiers, never a lone raw percentage (models are overconfident).
Design the failure screen first. HAX “When Wrong” = efficient + flexible correction, recovery, control, get help — a generic “Oops” satisfies none.
Two response tools: expectation-setting (proportional, in-context disclaimers) + ceding control (an escape hatch to a human). Air Canada lacked both ($812.02).
Five-part stack for consequential surfaces: show confidence, cite sources, allow regenerate, allow human review, allow exit. Gate by domain, not one global threshold (AI Overview).
Next: trust calibration & transparency — explanations, citations, showing the work, and avoiding over/under-reliance.