The design goal is calibrated trust, not maximal trust — over-trust is as much a defect as under-trust. Google PAIR’s input/system/context failure taxonomy, refusal-as-a-feature, citations as the trust boundary, and the recovery patterns that turn a wrong answer into a designed moment rather than an incident.
The goal is calibrated trust, not maximal trust
The canonical reference here is Google’s People + AI Guidebook (PAIR), and its central claim reframes the whole problem: the design goal is “properly calibrated user trust,” not maximal trust. Over-trust is as much a defect as under-trust — a user who believes the system is correct while it is hallucinating absorbs the full error cost. Marketing wants “AI you can trust”; the senior PM wants “AI you trust exactly as much as it deserves, per answer.” This lesson is how you design that calibration, and how the case round probes it.
Calibration is a judgment skill, and Cagan ties it directly to product sense: he is most excited by strong product judgment + generative AI, and worried about handing the same tools to people without the product foundation — because without calibration, AI features degrade into hallucinations users accept. PAIR’s corollary for design: explain for understanding, not completeness. Share the information a user needs to make a decision and move forward; do not try to explain everything the system is doing. A calibrated explanation trades completeness for usefulness — the citation a user can check beats a model-internals essay they can’t.
PAIR’s Errors + Graceful Failure chapter forces a discipline that is the single most useful move in an AI design round: before you ship, classify every failure mode into one of three buckets and design the response. The buckets are not academic — each implies a different product fix, and naming the right one is what separates a senior answer from “we’ll add error handling.”
code
1PAIR FAILURE TAXONOMY -- classify, then design the recovery23 Failure type Who/what is wrong Product response4 ------------- ------------------------- -----------------------------------5 User error bad input from the user prompt repair, inline guidance,6 clarifying question7 System error model failed on a fallback, temperature 0, route to a8 reasonable input stronger model, REFUSE ("I don't know")9 Context error the world changed retrieval grounding, "source of truth"10 (drift, stale docs) prompt, stale-data warning1112 Senior move: name the bucket BEFORE the UX. "We'll add error handling" is the13 junior answer; "this is a context error, so we ground + warn on staleness" is not.
The two named failures from Lesson 1 are textbook mis-classifications. Klarna treated a system error (the bot failing on reasonable support questions) as a routine retrievable failure with no designed recovery — so the tail became a quality crisis and a public reversal. Microsoft Recall treated capability acceptance as user consent — a context/trust error — shipping continuous screen capture default-on, then retreating to opt-in after a privacy and security backlash. Interview angle. “The model gets this wrong sometimes — what do you do?” is really “can you classify the failure?” Answer with the bucket first, then the matching recovery; an interviewer who hears “is this an input, system, or context error?” knows you’ve designed real AI products.
Refusal is a feature, not a defect
The cleanest graceful-failure mode available to a probabilistic system is to decline. Anthropic’s guidance for agents is explicit: ask the model to return “Unknown” when there is insufficient information. A RAG system that cannot say “I don’t know” will confidently hallucinate on every out-of-corpus question — so abstention is the difference between a trustworthy assistant and a plausible liar. The tradeoff is real and you should name it: refusal lowers coverage (the system answers fewer questions) in exchange for trust, and over-used it becomes “refusal-washing” — declining to dodge hard cases. The calibration is choosing where on the coverage/trust curve a given feature should sit.
Perplexity operationalizes the same idea on the answer side: every claim must come from a source with domain authority or a trust score, which is what turns an answer from something that “comes across like an opinion” into a source of truth. Citations are graceful failure in disguise — they bound the cost of a wrong answer to a verifiable boundary (the user can check the source) rather than leaving the user to guess. Aravind Srinivas’ bar — “almost never wrong, so you can trust what it says” — is a coverage-narrowing bet: cite-or-abstain answers fewer things, more trustworthily.
Surfacing uncertainty — the calibration UI
Calibration also has to be legible in the interface, and PAIR’s “explain for understanding, not completeness” rule decides what to show. The pattern stack, from cheapest to richest: (1) provenance — show the source/citation so the user calibrates by checking, not by trusting; (2) hedging language — “based on the docs I found…” vs a flat assertion sets the prior honestly; (3) graded states — a high-confidence answer, a “here’s my best guess” answer, and an “I don’t know” are visibly different UI states, not the same chat bubble; (4) human handoff — at low confidence or high stakes, route to a person rather than guess. The anti-pattern is a single confident-looking bubble for every answer, which guarantees over-trust on exactly the answers that don’t deserve it.
A subtle trap: raw model confidence is not calibrated probability. A model’s token-level likelihood, or its self-reported “I’m 90% sure,” does not mean it’s right 90% of the time — models are routinely over-confident, and a confidence number shown to users is worse than none if it isn’t validated against actual accuracy on a labeled set. So if you surface confidence, you owe a calibration check (does “90% sure” actually correspond to ~90% correct on the eval set?) before the number ships. Interview angle. “Would you show a confidence score?” → only if it’s calibrated against real accuracy; otherwise prefer provenance and graded states, which let the user calibrate without trusting a possibly-miscalibrated number.
Recovery patterns and the feedback loop
PAIR’s recovery pattern list is operational, not abstract. Design explicit feedback affordances — thumbs up/down, hide this recommendation, flag or report — and set the tone with two small rules: acknowledge that you received the feedback, and tell the user how the system will respond to it. These apply equally to a single-turn feature and a multi-turn agent. The non-obvious payoff is in the next lesson and the PRD lesson: that thumbs-down is not just UX politeness — it is a labeled training/eval signal. Redpoint’s framing is that your customers are your RLHF annotators: the user base is the eval pipeline, not just the consumer base.
Adobe Firefly shows the architecture: it pairs explicit signals (likes, dislikes) and implicit signals (saves, shares) to refine the model, inside a co-pilot pattern where the user stays in control — a graceful-failure architecture with a human in the loop by construction. Tome adds a structural recovery: it forces user agreement on an outline before building full content, which contains the downstream failure (a bad outline) by converting an open-ended probabilistic task into an iterative contract the user has already accepted. Interview angle. “How do you handle the model being wrong over time?” → the strong answer closes the loop: capture the correction as a labeled signal, feed it into the eval set, and gate the next model change on it — not “we’ll retrain.”
Beta framing and indemnification as trust tactics
Trust calibration is not only UI — it can live in positioning and contracts. Notion set the prior with language: launching Q&A as an “isn’t perfect” beta set calibrated expectations before the first error, paired with a measurable customer win to keep the beta credible. Adobe Firefly bet the enterprise launch on being “commercially safe” — training-data discipline plus IP indemnification — so the ship threshold became “is the customer indemnified if it isn’t perfect?” Salesforce collapses this into a platform primitive: the Einstein Trust Layer applies system policies that limit hallucinations and reduce harmful outputs, so trust becomes a reusable layer many use cases inherit. Each is a different answer to the same calibration question, and naming which one fits a given product is a senior signal.
Properly calibrated user trust — not maximal trust — is the goal; over-trust and under-trust are both failures, and graceful failure is a designed response to a classified error. — Google People + AI Guidebook (PAIR)
Interview prep
Trust-and-failure questions reward classifying the failure, naming the coverage trade, and closing the feedback loop. Lead with the PAIR taxonomy, then the matching recovery.
01“The model is sometimes wrong — what do you do?” → classify it: input vs system vs context error; each implies a different recovery (prompt repair / fallback+refuse / grounding+staleness warning).
02“How do you stop hallucination?” → ground-and-cite + enforced abstention + a faithfulness check; name the coverage cost — not “use a better model.”
03“Should the AI ever refuse to answer?” → yes; refusal (“I don’t know” / “Unknown”) is the cleanest graceful failure — a coverage-for-trust trade you tune per stakes.
04“Is more user trust always better?” → no; the goal is calibrated trust (PAIR). Over-trust means users absorb the cost of hallucinations they didn’t catch.
05“How do you design for when the model is wrong?” → make recovery a first-class feature: acknowledge, explain the system’s response, and capture the correction as a labeled signal.
06“How much should the UI explain?” → explain for understanding, not completeness — the decision-relevant info (a checkable citation), not model internals.
07“How would you set expectations at launch?” → calibrate via framing (Notion’s beta), contracts (Adobe indemnification), or a trust layer (Salesforce) — match the lever to the segment.
08“What do you do with thumbs-down data?” → treat it as RLHF-grade labeled feedback; route it into the eval set and gate the next model change on it.
Follow-ups push on the trade-offs. Expect “what’s the downside of citations / refusal?” (coverage drops; cite-or-abstain answers fewer questions), “how do you avoid refusal-washing?” (set an explicit abstention rate and review borderline refusals), and “who is your human in the loop, and at what threshold?” (name the reviewer and the confidence/stakes cutoff). Marily Nika’s framing is the bar: every AI design answer should tie to a messy world — hallucination, trust, and a user-override path — not a clean happy path.
Your AI assistant answers from a knowledge base. Users report it gives confidently wrong answers when the source doc was updated last week but the index is stale. Which PAIR bucket, and the fix?
AUser error — users are asking badly; add input guidanceBSystem error — the model is broken; route to a stronger modelCContext error — the world changed; fix with retrieval grounding, a freshness/“last verified” signal, and a stale-data warning
A PM says “let’s remove the ‘I don’t know’ response — it makes the product look weak, and coverage looks better without it.” Best senior pushback?
ARefusal is the cleanest graceful failure (Anthropic’s “Unknown”); it’s a deliberate coverage-for-trust trade. Removing it means confident hallucination on every question outside the corpusBAgree — always answering improves the demo and the coverage metricCKeep it but hide it behind a setting most users won’t find
Microsoft Recall shipped continuous screen capture default-on and was pulled to opt-in after backlash. Through the trust lens, the core mistake was…
AThe feature was too slow, so users disabled itBIt mis-classified capability acceptance as user consent — a trust/context miscalibration that introduced a high-stakes capability with no graduated, opt-in switchCThe model hallucinated screen contents
Your team adds thumbs up/down to an AI feature. Beyond UX politeness, what’s the highest-leverage way to use that signal?
AShow an aggregate satisfaction score on a dashboardBEmail the user back to thank them for the feedbackCTreat it as RLHF-grade labeled feedback — route corrections into the eval set and gate the next model change on whether it improves them
Adobe launched Firefly to enterprises foregrounding “commercially safe” (training-data discipline + IP indemnification). What trust problem does that solve that a UI change cannot?
AIt makes the images generate fasterBIt reframes the ship threshold to “is the customer indemnified if it isn’t perfect?” — moving the residual risk of a wrong/infringing output off the customer and onto a contractCIt guarantees the model never produces an infringing image