Lesson 3 of 6 · 49 min

Trust calibration & transparency

Trust is calibrated, not granted — it rises when behavior matches the user’s model and falls when it diverges. The transparency primitives that move it (explanations, citations, showing the work), the over- vs under-reliance curve and its measured costs, and the cognitive-forcing moves that beat automation bias — grounded in Perplexity, the o1-vs-Claude reasoning split, and Anthropic’s Citations API.

Trust is a curve, not a goal

The brief is never “build trust” — it’s calibrate trust. Trust calibration is the correspondence between how much a user trusts the AI and how trustworthy it actually is, and the science is older than LLMs (Lee & Moray studied it on nuclear-plant operators in 1992). Picture a curve: at low trust, users ignore correct advice (under-reliance / algorithm aversion); at high trust, they accept wrong advice (over-reliance / automation bias). Both ends cost money, and both are design failures, not user failures. Your transparency tools — explanations, citations, showing the work — are the levers that move a user along that curve toward the calibrated middle. This lesson is how you aim them.
The mechanism: trust moves up when the model’s behavior matches the user’s mental model, and down when it diverges. Explanations are the corrective signal — they let the user update their model in real time. That’s why PAIR’s defining rule for explanations is “explain for understanding, not completeness” — a pattern focused on “sharing information users need to make decisions rather than explaining everything in the system.” A ten-page model card does less for trust than a two-line “I matched you to this because…”, because completeness over-explains and under-decides. NN/g adds the behavioral fact: users engage in information foraging — they hunt for evidence, not summaries — and they do it while the model writes, not after.
AI Is Your New Design MaterialJosh Clark · AIGA

The two failure modes and their measured costs

Both ends of the curve have a paper and a price tag, and naming them is a strong-candidate signal. Over-reliance (automation bias): Buçinça et al. (Harvard, 2021) showed users over-accept AI suggestions in simulated decision tasks because the output sounds fluent; the cost is clinical misdiagnosis and fabricated legal citations. Under-reliance (algorithm aversion): Dietvorst et al. (Wharton, 2015) showed users abandon an algorithm after seeing it err once — even when it still outperforms humans; the cost is discovered productivity gains thrown away. The asymmetry that should shape your design: a single visible mistake can collapse trust below where it should sit, so recovery UX matters as much as the trust-building copy.
There’s a third, sneakier failure between them: fluent-but-wrong. Jakesch (PNAS, 2023) showed people can’t detect AI-generated language because the fluency heuristic equates “sounds right” with “is right.” The mitigation is to make the generation visible, not just the output — show “I’m not sure, but…”, expose the citation graph, surface the model version. Interview angle. “How do you avoid users over-trusting your AI?” A weak answer says “add a disclaimer.” A strong answer names automation bias and the fluency heuristic, then proposes a cognitive-forcing move (below) plus visible evidence — and notes the opposite risk, algorithm aversion, so you don’t over-correct into a product nobody trusts.
code
1THE TRUST CURVE -> two failure modes, two design responses23   under-reliance            CALIBRATED               over-reliance4   (algorithm aversion)        (target)              (automation bias)5        |------------------------|------------------------|6   ignores correct advice                          accepts wrong advice7   Dietvorst 2015: abandons                        Bucinca 2021: over-accepts8   after ONE visible error                         fluent suggestions910   move UP the curve              hold the middle           move DOWN the curve11   - progressive trust build      - visible evidence        - cognitive forcing12   - "you can edit it" controls     at the foraging moment     (commit-before-reveal)13   - show what changed since       - citations + confidence  - friction on high stakes14     last time                       proportional to stakes  - show reasoning, not polish
The most actionable anti-over-reliance technique is the cognitive forcing function: make the user commit before the AI reveals its answer. Buçinça’s intervention — “mark your answer first, then we’ll show what the AI suggests” — reduced over-reliance while preserving the AI’s benefit. GitHub Copilot’s “take the task, then suggest” flow is a shipped version. A large Nature Medicine study of AI assistance across 140 radiologists found the effect is heterogeneous — for some readers AI assistance actually degraded accuracy — so how and when the AI’s answer is surfaced changes diagnostic-correctness patterns, not just whether it’s present. The design move is to not show the AI suggestion first on high-stakes decisions — show the user’s own answer box first, or surface what the AI is uncertain about before its conclusion.

Three transparency primitives — and when to use each

Transparency is plural, and the error is shipping one primitive (usually the model card) and assuming it covers the others. Three recur across all four frameworks: explanations (PAIR partial/progressive disclosure, IBM explainability), citations (the Perplexity pattern), and showing the work (visible reasoning, audit trails). Different surfaces need different ones — match the primitive to the decision, not to convenience:
  1. 01“Why this result?” link (PAIR partial/progressive) — one-tap explanation on demand. Best for recommendations and classifications; default to a one-liner, expand on click, reserve the model card for power users.
  2. 02Inline citation footnote/chip (Perplexity, ChatGPT) — numbered reference inline, hover/click for the source. Best for factual generation and search; date-stamp sources for an extra calibration signal.
  3. 03Visible reasoning / “show the work” (extended thinking) — streamed reasoning block. Best for code, math, multi-hop analysis; lets the user audit a wrong answer instead of trusting it wholesale.
  4. 04Numeric confidence + N alternatives (PAIR, HAX) — visible ranking with scores. Best for recommendations/retrieval; always paired with an anchor, never a lone number.
  5. 05Model card / system info (IBM explainability) — full capability disclosure. Best in settings / first-run, NOT inline; completeness here over-explains and under-decides.
Two design rules cut across all three primitives. First, stream explanations during generation, not after — users forage while the model writes, so that’s when evidence has to appear. Second, default “show the work” for high-stakes domains (medical, legal, financial, code review) and hide it for low-stakes fluency tools (autocomplete, casual chat) where it’s just clutter. Interview angle. “Should you always show the model’s reasoning?” The strong answer is no — it’s stakes-dependent; show it where the user will audit and act, hide it where it adds noise. Reaching reflexively for “always be transparent” misses that transparency has a cost in attention.

Case study: Perplexity’s five-level citation stack

Perplexity layers trust signals at five levels, and it’s the canonical citation reference: (1) Domain +N chips inline on each claim, (2) a Sources summary row, (3) a Links tab for a full audit, (4) Check sources on selection — highlight any sentence to see its support, and (5) a Wrong sources feedback affordance. Two things make this senior-grade. The “+N” notation is exactly PAIR’s “N-most-likely” pattern, repurposed from classifications to evidence. And crucially, the audit affordance is on-demand and contextual — it sits next to the claim, not buried in a settings menu — which directly solves NN/g’s “polish discourages error checking” failure by making checking cheap.
The failure case keeps you honest: Reddit threads document Perplexity “failing to recognize valid links when submitted in bulk.” The citation UI does its job only when the retrieval underneath does its job — a beautiful audit affordance over broken retrieval just lets users watch the system be wrong with footnotes. Interview angle. “Citations make AI trustworthy — agree?” The nuanced answer: citations are necessary and high-leverage, but they shift trust onto the retrieval layer; a cited-but-wrong source can be more persuasive than an uncited guess. The design has to make the source legible enough that a user can catch a bad citation, not just see that one exists.

Case study: o1 hidden reasoning vs Claude visible thinking

The frameworks do not converge on whether to show the work — and as a designer you have to choose. OpenAI’s o1 (Sept 2024) introduced reasoning tokens that are billed and counted but “not visible in the API response” — reasoning as an implementation detail, hidden by default. Anthropic took the opposite stance: Claude 3.7 lets users “toggle extended thinking mode on or off,” and the API exposes “thinking content blocks where it outputs its internal reasoning” — reasoning as a trust surface. The product consequence is sharp: when reasoning is shown, a user can audit a wrong answer and salvage the good parts (NN/g’s “Apple Picking” — keep the right pieces, revise the rest); when it’s hidden, they must trust or distrust the whole output.
code
1SHOW-THE-WORK: the two stances, and when each is right23  o1 (hidden reasoning)            Claude 3.7 (visible, toggleable)4  ----------------------          --------------------------------5  reasoning = implementation      reasoning = trust surface6  user trusts/distrusts WHOLE     user can audit + "apple pick"7  cleaner, lower cognitive load   higher load, but recoverable errors89  DESIGN HEURISTIC:10    high stakes + user will act on it   -> show the work (auditable)11    low stakes + speed/flow matters     -> hide it (less clutter)12    pair VISIBLE reasoning with PER-CLAIM citations = strongest current pattern13    (the user sees HOW it thought AND WHAT evidence backs each claim)
The strongest current pattern pairs the two: visible reasoning + per-segment citations. The user sees how the model thought and what evidence supports any individual claim — which together let them calibrate at the granularity of a sentence rather than a whole answer. This is also where regulation is heading: EU AI Act Article 13 requires high-risk systems to be “designed to be transparent, so that those using them can understand and use them correctly,” which makes explanations and provenance a legal expectation for consequential surfaces, not a nice-to-have.
A citation turns “is the model right?” into “is this source right?” — a smaller, answerable question. Visible reasoning turns “trust the whole answer” into “keep these parts, fix those.” Both shrink what the user has to take on faith.

Interview & portfolio prep

Trust and transparency show up in the design-challenge round (“how do users know to trust this?”), the app critique, and portfolio walkthroughs (“what was the strongest evidence your design works?”). The signal: you treat trust as a calibrated curve with two failure modes, you can name the right transparency primitive per surface, and you know transparency has a cost. Drill these.
  1. 01“Build trust vs calibrate trust — what’s the difference?” → calibration targets the middle of a curve: under-reliance (algorithm aversion, Dietvorst) at one end, over-reliance (automation bias, Buçinça) at the other; both are design failures.
  2. 02“How do you prevent over-reliance?” → cognitive forcing (commit before reveal; don’t show the AI’s answer first on high stakes) + visible evidence; name automation bias and the fluency heuristic.
  3. 03“Why not just always show the model’s reasoning?” → stakes-dependent; show the work for code/medical/legal where users audit and act, hide it for low-stakes fluency tools where it’s clutter (PAIR: understanding, not completeness).
  4. 04“Walk me through Perplexity’s citation UX.” → five levels (domain +N chips, sources row, links tab, check-sources-on-selection, wrong-sources feedback); on-demand + contextual makes checking cheap.
  5. 05“Do citations make AI trustworthy?” → necessary and highest-leverage, but they shift trust to the retrieval layer; a cited-but-wrong source can be more persuasive — make the source legible enough to catch.
  6. 06“o1 hides reasoning, Claude shows it — which is right?” → depends on stakes; visible reasoning enables auditing and apple-picking, hidden lowers load. Pair visible reasoning + per-claim citations for the strongest pattern.
  7. 07“When does explanation belong inline vs in settings?” → one-line ‘why?’ inline at the foraging moment, expand on click; model card lives in settings/first-run, never inline.
  8. 08“What’s the strongest evidence a trust design works?” → a calibration metric: do users accept correct answers AND catch wrong ones at higher rates? Over-trust and under-trust both move, ideally measured.
Follow-ups dig into measurement and tradeoffs: “How would you measure calibrated trust?” (not satisfaction — measure whether acceptance tracks correctness: appropriate reliance, not blanket reliance; you want users to take good suggestions and reject bad ones, and you can A/B a forcing function against a control). “Doesn’t a commit-before-reveal step slow people down?” (yes, deliberately — that friction is the point on high-stakes decisions; you scope it to where over-reliance is costly, not everywhere). A standout portfolio piece names which transparency primitive it shipped and why, streams evidence at the foraging moment, and shows a before/after on a calibration or appropriate-reliance proxy — not just “users said they trusted it more.”
docsExplainability + Trust (explain for understanding, not completeness)Google PAIRpaperCognitive Forcing Functions Can Reduce Overreliance on AIBuçinça et al., 2021docsAI UX Patterns — Citations (inline cues, source panels, verification; Perplexity-popularized)Shape of AIarticleClaude’s extended thinking (reasoning as a visible trust surface)Anthropic

Checkpoint

You’re designing an AI diagnostic-support tool for clinicians. Pilot data shows clinicians accept the AI’s suggestion even when it contradicts their own initial read. What’s the best-evidenced design intervention?

AA cognitive forcing function: require the clinician to record their own read before the AI’s suggestion is revealed, so they engage their own judgment firstBMake the AI’s recommendation more prominent and confident so clinicians act on it fasterCAdd a footer disclaimer that the tool is decision-support onlyDHide the AI’s confidence so clinicians don’t anchor on it
Sign up free to answer and see why

Checkpoint

Your AI research tool already cites every claim with a numbered source. A reviewer says trust is still low because users can’t tell good citations from bad. Best next move?

AAdd more citations per claim so each answer looks better-supportedBMake each source legible at the foraging moment — surface the domain, date, and a snippet on hover/selection so users can judge the source, not just see that one existsCRemove citations and show a single overall confidence score insteadDMove the sources into a separate page so the answer looks cleaner
Sign up free to answer and see why

Checkpoint

A casual-chat AI feature (low stakes, speed matters) ships with the full model reasoning trace expanded under every reply. Users complain it feels cluttered and slow to read. What does this reveal?

AUsers dislike transparency and the reasoning should be removed from all productsBThe model is reasoning too much; lower its reasoning effort globallyCAdd a confidence percentage to compensate for the clutterDTransparency is stakes-dependent: “show the work” fits high-stakes, auditable tasks; for a low-stakes fluency tool it’s clutter, so collapse it behind an optional toggle
Sign up free to answer and see why

Checkpoint

A PM proposes boosting trust in your AI assistant by removing all hedging and making every answer sound assertive and authoritative. As the designer, your strongest objection is:

AIt optimizes for over-trust: confident, fluent output is exactly what drives automation bias and the fluency heuristic, so users will accept more wrong answers — calibrated trust, not maximal trust, is the goalBAssertive copy is harder to write and will slow the team downCUsers prefer cute, friendly copy over assertive copyDAssertive answers will increase server costs
Sign up free to answer and see why

Checkpoint

You want to measure whether your transparency redesign actually calibrated trust (rather than just raising it). Best metric?

AOverall trust rating on a post-task surveyBTime-on-task onlyCAppropriate reliance: whether users accept correct AI outputs AND reject incorrect ones at higher rates than before — measuring both directions, ideally against a controlDNumber of citations clicked
Sign up free to answer and see why

Could you explain the trust curve and its two failure modes, choose the right transparency primitive per surface, and defend a calibration metric in an interview?

New to itGetting thereConfident

Takeaways

  • Calibrate trust, don’t maximize it: under-reliance (algorithm aversion, Dietvorst) and over-reliance (automation bias, Buçinça) are both design failures with price tags.
  • A single visible error can collapse trust below where it belongs — so recovery UX matters as much as trust-building copy.
  • Cognitive forcing (commit before the AI reveals) is the measured anti-over-reliance move; don’t show the AI’s answer first on high-stakes decisions.
  • Transparency is plural — explanation, citation, show-the-work — and stakes-dependent; stream it at the foraging moment, hide it where it’s clutter (PAIR: understanding, not completeness).
  • Perplexity’s five-level, on-demand citation stack makes checking cheap; citations shrink the risk envelope but shift trust onto retrieval — make sources legible.
  • o1 hides reasoning, Claude shows it: choose by stakes; visible reasoning + per-claim citations is the strongest current pattern, and EU AI Act Art. 13 makes it an expectation for high-risk systems.

Next: user control over automation — undo, edit, override, and choosing the right level of automation.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.