Users run a private simulation of what your AI can do — and every acceptance, every surprise, every collapse of trust traces back to whether that model is honest. How mental models form, why intent-based interfaces raise the stakes, the sycophancy trap, and the onboarding/empty-state moves that seed an accurate model instead of a hype one.
The design surface you can’t see
Before a user clicks anything, they have already built a private theory of your feature: what it can do, what it won’t do, who is accountable when it’s wrong. Google’s PAIR Guidebook opens its very first chapter on exactly this — a mental model is “a person’s belief system that helps set expectations for what a product can and can’t do.” That invisible model is the real interface. When it’s accurate, users accept good output, catch bad output, and recover gracefully. When it’s inflated — and AI marketing inflates it by default — they over-trust, get burned, and abandon. Most AI UX failures are mental-model failures wearing a different costume. This lesson is the ground floor for the whole track.
A mental model is the user’s running simulation of your system. It sets the prior on every interaction: how much to trust an answer, how surprised to be on failure, whether to even check. The reason this matters more for AI than for a button is that AI moved us into what NN/g calls a third UI paradigm — intent-based interaction, where users state goals (“clean up my inbox”) rather than issue commands (“delete this row”). Command UIs are constrained: the button does one thing, so the model is easy to keep accurate. Intent UIs have enormous degrees of freedom, so the gap between “what the user thinks it does” and “what it actually does” can be vast — and the user finds out only after they’ve delegated.
IBM makes mental models a first-class generative-AI principle: “designers must carefully consider how to impart useful mental models to help users understand how a system works.” Microsoft’s HAX Toolkit puts two of its 18 guidelines in the “Initially” phase for the same reason — G1: Make clear what the system can do and G2: Make clear how well the system can do what it can do. Both fire before the first interaction precisely because, once a wrong model forms, every downstream pattern you build — confidence displays, citations, undo — has to fight it rather than ride it. Interview angle. If you can name where in the timeline (initially / during / on-error / over-time) a given design move belongs, you signal you think in the HAX framework, not in vibes.
Four canonical frameworks — treat them as one craft
There are four pattern libraries every senior AI designer is expected to know, and the rookie mistake is adopting one as orthodoxy. They converge on the same rules — calibrate trust, expose uncertainty, return control, design for access — but each leads with a different lever because each was built by a team solving a different problem at a different scale:
01Google PAIR Guidebook — leads with mental models + trust. Deepest on the “why,” because Google ships consumer AI at planetary scale. Weakest on shipped components.
02Microsoft HAX Toolkit — 18 evidence-based guidelines in 4 temporal phases (Initially / During / When Wrong / Over Time). The most enforcement-friendly: a checklist eng + design can grep in review. Prescribes no UI.
03NN/g — usage patterns and anti-patterns (sycophancy, “discourages error checking,” the ambiguous sparkles icon). The most user-research-grounded; tells you what NOT to ship more than what to ship.
04IBM Design for AI — 6 ethics pillars (Accountability, Value alignment, Explainability, Fairness, User data rights, Everyday ethics) PLUS Carbon for AI, the only library that ships the actual components (AI label, AI explainability, accessible chat).
How the wrong model forms — and what it costs
Three forces inflate the mental model before you get a say. (1) Marketing copy: positioning AI as a coworker imports social expectations (competence, deference, implied consent). (2) Fluency: well-formatted, confident prose reads as authoritative regardless of correctness — Jakesch et al. (PNAS, 2023) showed people cannot detect AI-generated self-presentations because judgment is “misguided by intuitive but flawed heuristics.” The output looks expert, so the user models the system as expert. (3) Early wins: the first three correct answers set an anchor the fourth, wrong one then sails past unchecked. NN/g’s field number is brutal — across 200 quotes, ChatGPT misattributed roughly 76%, and only 7 of 153 error cases carried any signal of uncertainty. The model says nothing, so the user keeps the inflated prior.
The cost shows up in the financial record, not just in surveys. Google’s Bard demo gave one wrong fact about the James Webb Space Telescope in a promo video; Alphabet shed about $100B of market value and the stock fell ~9% (Guardian, Feb 2023). The lesson for designers is not “the model hallucinated” — it’s that a single confidently-wrong output, with no friction layer, is a front-page event. You design for the worst-case mistake on day one, not after the postmortem.
Two honest-model patterns: scoped AI and the labeled guide
The strongest mental-model design doesn’t just describe the system honestly — it structurally constrains the system so the honest description is true by construction. Google’s NotebookLM (designed by Jason Spielman’s team) is the cleanest example: you upload your sources first, and the AI can only synthesize from those sources. That’s hallucination-by-design — the output space is bounded by the uploaded material, so the user’s natural model (“it answers from my docs”) actually matches reality. When the use case has known inputs, scoped AI is the most powerful mental-model guardrail there is, because it removes the gap between expected and actual behavior rather than papering over it with copy.
Spotify’s AI DJ is the counterpoint that shows a coworker-style model done well: it’s framed as “an AI guide that knows your music taste,” it’s clearly named and identified as the DJ (not disguised as a human), it tells you why it’s playing something (a trust-builder), and a thumbs-down shapes future picks (a feedback loop). Across 100B+ Discover Weekly tracks streamed in a decade, Spotify’s repeated lesson is that the personalization feedback signal matters more than the model’s underlying “talent.” The mental model is honest because the identity, the reasoning, and the controls are all visible — the opposite of a sparkle-avatar bot pretending to be more capable than it is.
Case study: Notion’s “24/7 AI team” and the sycophancy trap
Notion markets its AI as “Meet your 24/7 AI team” — agents that “build, edit, and take action” while you sleep, with OpenAI, Figma, Ramp, and Nvidia named as adopters. That is a deliberate coworker mental model, and it’s a powerful one: it makes the product feel capable and delegable. But it imports a risk NN/g names directly — sycophancy, where LLMs “prioritize pleasing the user over factual accuracy.” A coworker model silently trains users to extend social trust: to skip the double-check, to read agreement as competence. A sycophantic model will agree with a flawed premise to keep the relationship warm. So the very framing that drives adoption also disarms the user’s skepticism.
The design response is not to abandon the coworker framing — it sells — but to inject friction proportional to the framing. Notion’s actual UI keeps AI output inline in the user’s own document with accept / edit / decline, so a correction lives in the artifact rather than vanishing in chat; and any side-effecting agent action wants a confirm rail precisely because the coworker model has taught the user to skip it. Interview angle. “Notion calls its agents a 24/7 team — what’s the risk and how would you mitigate it?” A strong answer names sycophancy and the implied-consent problem, then proposes friction (confirmation, are-you-sure rails) on irreversible actions — not a vague “add guardrails.”
Seeding an honest model: onboarding, empty states, examples
You shape the mental model in the first sixty seconds, with three surfaces. Empty states and onboarding are where you teach capability and — more importantly — the refusal frontier. The single most under-used move: show the model failing on a representative task, on purpose. An adversarial example (“here’s a question it will get wrong, and how to tell”) calibrates faster than ten success demos, because it installs the checking habit you actually want. Capability chips beat catch-all labels: “Drafts replies · Won’t send on your behalf · Can be wrong on dates” sets sharper expectations than “AI Assistant.” And example prompts should demonstrate breadth and edge, not just the happy path.
code
1EMPTY-STATE COPY: catch-all label vs honest capability framing23 Weak (inflates the model):4 [ ✨ AI Assistant ]5 "Ask me anything."67 Strong (seeds an accurate model):8 [ ✨ Draft assistant ] · knows your docs from the last 30 days9 "I draft replies and summaries from THIS workspace. I can be wrong on10 dates and numbers -- check those. I never send or delete on my own."11 Try: "summarize the Q3 launch thread" (breadth)12 "what did Priya commit to on Tuesday?" (shows it cites the source)13 "what's our 2027 revenue?" (shows it says 'I don't know')1415 The third example is the important one: it demonstrates the REFUSAL16 frontier, which is what stops over-trust before it starts.
The deeper principle PAIR names is “explain for understanding, not completeness.” You are not trying to teach the architecture; you are trying to install a good-enough simulation that predicts behavior. A two-line “here’s what I read and what I won’t touch” does more for calibration than a ten-page model card the user never opens. Interview angle. “How do you set expectations for a new AI feature?” Strong answers cite the surfaces (onboarding, empty state, first-run examples), name the adversarial-example move, and tie it to a framework guideline (HAX G1/G2, PAIR mental models) rather than saying “good onboarding.”
The interface users actually operate is the model in their head. Your screens are just the evidence they use to build it — so design the evidence to build an accurate one.
Interview & portfolio prep
AI-design loops probe mental models in the app-critique and “design an AI feature” rounds, and in portfolio walkthroughs. The signal they want: you can name how users form expectations, where in the timeline you intervene, and which framework backs each move. Have crisp answers ready.
01“What is a mental model and why does it matter for AI?” → the user’s running simulation of what the system can/can’t do; it sets the prior on trust, surprise, and whether they check. Inflated models cause over-reliance and abandonment.
02“Why are mental models harder for AI than traditional UI?” → intent-based interaction (NN/g’s third paradigm) gives the system far more degrees of freedom, so the gap between expected and actual behavior is larger and discovered post-delegation.
03“Name the four canonical frameworks and what each leads with.” → PAIR (mental models/trust), HAX (18 guidelines, 4 phases), NN/g (anti-patterns), IBM/Carbon (ethics pillars + shipped components).
04“Notion markets a ‘24/7 AI team’ — risk?” → sycophancy + implied consent; the coworker framing disarms skepticism. Mitigate with friction on irreversible actions, inline accept/edit, confirm rails.
05“How do you set expectations in onboarding?” → capability chips over catch-all labels; show an adversarial failure on purpose; example prompts that demonstrate the refusal frontier, not just wins.
06“Explain for understanding vs completeness — what does that mean?” → install a good-enough simulation (two-line ‘what I read / won’t touch’), not the architecture or a model card nobody reads.
07“What does the Bard $100B story teach a designer?” → a single confidently-wrong output with no friction layer is a front-page event; design for the worst-case mistake on day one.
08“Where do mental models live in a portfolio piece?” → lead with the user’s state and the expectation gap; show the onboarding/empty-state choices that closed it, with before/after comprehension evidence.
Follow-ups go deeper on evidence and tradeoffs: “How would you test whether users have the right mental model?” (comprehension probes in usability testing — ask users to predict what the feature will do before they use it; mismatch = your model-setting failed). “Doesn’t showing failures hurt adoption?” (counter-intuitively it builds durable trust — Dietvorst’s algorithm-aversion work shows users tolerate an imperfect system far better when they expected imperfection and can edit it). A strong portfolio piece on this names the AI’s role explicitly (off-the-shelf vs custom vs designed), shows the expectation-setting surfaces, and ties a measurable delta (comprehension, task success, trust) to the design move.
Your team is shipping an AI inbox assistant. Marketing wants the empty state to say “Your AI inbox genius — ask it anything.” You’re worried about downstream trust. Strongest design counter-move?
AKeep the copy — an aspirational label drives adoption, and you can add an “AI can make mistakes” banner above the chatBReplace the catch-all with capability chips and three example prompts — one of which shows the assistant saying “I don’t know” — so the empty state teaches the refusal frontierCAdd a confidence percentage next to every reply so users can judge for themselvesDGate the feature behind a tutorial that explains the transformer architecture
In a usability test, five users each accept an AI summary that contains one fabricated statistic without noticing. The output is fluent and well-formatted. What does this most likely indicate?
AThe model needs a larger context window so it stops fabricatingBUsers are careless and should be told to read more carefullyCThe fluency heuristic is suppressing verification — the polished format reads as authoritative, so you need a design that makes checking cheap and prompts it at the moment of relevanceDThe summary feature is working as intended; five users is too small a sample to act on
An interviewer asks you to place a design move on the HAX timeline: “show a worked example of the AI getting something wrong, so users learn to verify.” Where does it belong, and why does the phase matter?
AThe “Over Time” phase — failure examples are about long-term learningBThe “Initially” phase — it sets capability and the refusal frontier before the user forms an inflated model, which every later pattern then has to fightCThe “When Wrong” phase — it’s about errorsDIt doesn’t belong on the timeline — failure demos hurt adoption and should be cut
A PM says: “Let’s frame our agent as a tireless 24/7 teammate — it tested great in messaging.” As the designer, what’s the most important thing to add alongside that framing?
ANothing — if the messaging tests well, ship it; framing is marketing’s callBA faster model so the “tireless” claim feels trueCA brighter “AI” sparkle icon so users always know it’s the agentDFriction proportional to consequence — confirm rails and inline accept/edit on irreversible actions — because the coworker frame trains users to skip the very checks those actions need
You want to verify, before launch, that users hold an accurate mental model of what your AI feature does. Best research move?
AShip it and watch the thumbs-up rate — high ratings mean the model is accurateBIn usability sessions, ask users to predict what the feature will and won’t do before they touch it, then compare predictions to reality — mismatches reveal where your expectation-setting failedCA/B test two accent colors for the AI labelDCount how many users open the help docs
Could you walk an interviewer through how users form mental models of an AI feature, name the four frameworks, and defend the onboarding moves that seed an accurate one?
New to itGetting thereConfident
Takeaways
The real interface is the model in the user’s head; your screens are the evidence they build it from. Design the evidence for accuracy.
Intent-based AI (NN/g’s third paradigm) has huge degrees of freedom, so the expectation gap is large and discovered after the user delegates.
Know all four frameworks — PAIR (why), HAX (checklist), NN/g (anti-patterns), IBM/Carbon (components) — and ship their intersection.
Fluency inflates the model: polished output suppresses checking (Jakesch PNAS 2023; NN/g’s 7-of-153 uncertainty signals). A single confident miss can be a $100B event (Bard).
Seed an honest model in onboarding/empty states: capability chips, a deliberate failure demo, examples that show the refusal frontier. Explain for understanding, not completeness.
Counter the coworker frame (Notion’s “24/7 team”) with friction proportional to consequence — sycophancy + implied consent are the trap.
Next: designing for uncertainty & errors — confidence display, graceful failure, and recovery when the model is wrong.