Lesson 1 of 5 · 48 min

A pattern language for AI

AI UX is not one component — it is a small, reusable vocabulary: suggestion, citation/provenance, permission/control, and uncertainty. The canonical libraries (HAX, PAIR, NN/g), the named patterns inside each domain, the shipped products that got them right and wrong, and how to choose a pattern by blast radius rather than by what looks nice.

Why a sparkle icon is not a strategy

Most AI features get designed one screen at a time: a chat box here, a “✨ generate” button there, an “AI can make mistakes” footer at the bottom. The result is a product where every AI surface looks invented from scratch and none of them tell the user the same story. The senior move is to treat AI UX as a pattern language — a small, named vocabulary of reusable patterns that you reach for deliberately. The Gradient put the failure precisely: pattern libraries are “organised like component libraries… but product teams building AI features are still choosing the wrong patterns,” because no library answers the question that matters — given this product, these users, and this level of autonomy, which pattern fits? This lesson builds that vocabulary, and the selection logic on top of it.
Three bodies of work anchor the discipline, and a senior designer should be able to name what each is for. Microsoft HAX Toolkit — 18 evidence-based Guidelines (validated in Amershi et al., CHI 2019) plus a reusable Design Patterns library — is the normative source: what to do, organised across the interaction lifecycle. Google PAIR (People + AI Guidebook) frames everything around calibrated user trust — mental models, explainability, autonomy, safety. Nielsen Norman Group is the empirical spine: the only corpus grounded in ongoing usability testing, so it is what you cite when HAX or PAIR collide with how users actually behave. The pattern encyclopedias — AIUX Playground (172+ patterns across 11 categories) and Shape of AI (six lenses: Wayfinders, Inputs, Tuners, Governors, Trust builders, Identifiers) — sit on top as named, shipped-example catalogs.
Guidelines for Human-AI InteractionDr. Saleema Amershi

The four pattern domains every AI surface touches

Cut across all the libraries and the same four problem domains recur. Memorise these as the spine of the language — every AI feature you design will reach into at least one, usually three.
  1. 01Suggestion — the AI proposes something before you ask (prompt chips, smart autocomplete, suggested actions, drafts). The design question: how visible, how dismissible, how it sets expectations.
  2. 02Citation / provenance — who or what produced this, and where did it come from (source rows, inline refs, “AI-generated” badges, content credentials). The question: can the user verify and attribute correctly?
  3. 03Permission / control — who is allowed to start, stop, undo, or approve the AI (invocation, dismissal, correction, approval gates, autonomy budgets). The question: does control match the blast radius?
  4. 04Uncertainty / consistency — the model output varies and is sometimes wrong, so the surface must stay coherent (confidence signals, hedged language, regenerate, graceful failure). The question: do we promise a stable experience even when the output is not?
Notice these are not components — they are problem domains, and each resolves into specific named patterns at three grain sizes: vocabulary tokens (e.g. an --ai-confidence-low token), components (a <CitationRow>), and flows (a plan-then-approve handoff). A mature AI design system produces patterns at all three sizes; a junior one ships a single chat component and calls it “the AI.”

Suggestion: the most over-flattened domain

Designers routinely collapse three different suggestion patterns into one “AI button.” NN/g separates them by capability and complexity. Use-case prompt suggestions are content, not chrome — they ladder from pills (lightweight, ChatGPT’s “Analyze data”) → cards (richer, Claude’s pre-auth carousel) → carousels → example libraries (Midjourney pairs prompts with their outputs). NN/g’s rule: pills should be unobtrusive, dismissible, and randomized to avoid positional bias, and pulled from an analytics-driven list of real user goals, refreshed regularly — not a marketing wishlist.
Smart Autocomplete is a separate, lower-visibility pattern: intent- and context-aware inline completion (GitHub Copilot’s ghost text, Gmail Smart Compose, Tabnine, IntelliSense). The pedagogical difference AIUX Playground draws is that Copilot consumes semantic context (the whole file/project) and its trigger is probabilistic, whereas IntelliSense is lexical and deterministic — so they feel similar but carry different trust contracts. Then Microsoft’s Copilot design system adds a third tier: Suggested User Actions (SUAs) — persistent, mid-flow capability-discovery surfaces, not one-shot chips you click and lose. Interview angle. If a prompt asks you to “design an AI suggestion,” the first thing that separates a strong answer is naming which suggestion pattern (notice → nudge → compose) and why, instead of drawing one generic chip.
code
1THE SUGGESTION LADDER -- four named patterns, by trigger and visibility23  tier      pattern                         trigger              shipped example4  -------   -----------------------------   ------------------   -------------------5  Notice    Use-case prompt suggestions     empty state / first  ChatGPT pills,6            (pills/cards/carousels/library)  encounter            Claude cards, Midjourney7  Nudge     Suggested User Actions (SUAs)   mid-flow, context    Microsoft 365 Copilot8                                             shift9  Compose   Smart Autocomplete              inside an editor     Copilot ghost text,10            (inline, context-aware)                              Gmail Smart Compose11  Hand off  Throw & Catch (entry/canvas/    crossing a surface   Microsoft 365 Copilot12            chat handoff)                    boundary1314  Pick the row by WHERE and WHEN the suggestion appears, not by how it looks.
One organizing decision worth making early: adopt a single taxonomy and route every new feature through it before pixels. AIUX Playground organizes 172+ patterns across 11 categories (Chatbot, Design Tools, Agents, Audio, Commerce, Inputs, Outputs, Trust, Collab, Onboarding, Performance); Shape of AI collapses the same surface into six lenses (Wayfinders, Inputs, Tuners, Governors, Trust builders, Identifiers). You do not need both — you need one, used consistently, so that “what category is this?” is the first question asked of any AI feature. The encyclopedias are for routing and naming; the four domains above are for reasoning. Teams that skip this step end up with a folder of one-off screens and no shared language to review them against.

Permission: ladder the control surface to the blast radius

The governing principle of the permission domain: the more side-effects an action has, the heavier its control surface should be. HAX sets the floor with four guidelines — G7 efficient invocation (cheap to start), G8 efficient dismissal (cheap to stop), G9 efficient correction (cheap to undo), G17 global controls (suite-wide toggle). For agents that take real actions, AIUX Playground catalogs the next two layers. Approval Workflows require “human review and approval before AI-generated content or actions are executed” (Power Automate, Slack Workflow Builder, Asana, Jira). Autonomy Budgets are “hard bounds on time or action count” — Cursor Agent’s max-iterations, Claude Code’s session caps, Devin’s work units, GitHub Actions job timeouts — that default to pause when exhausted.
The design-system implication is a ladder, not a menu: a low-stakes inline suggestion needs only a cheap dismiss (G8); a draft the user will send needs a clear accept/discard; an agent that files tickets or deletes rows needs an approval gate plus a visible, decrementing budget. Interview angle. “How would you design permissions for an AI agent?” — the weak answer is “a confirmation dialog.” The strong answer ladders controls to blast radius and names the budget-defaults-to-pause rule, because runaway sessions destroy trust faster than AI features build it.

Case study: Microsoft Copilot — one capability, three surfaces, one orchestrating pattern

Microsoft’s public Copilot design system is the clearest production example of pattern-language thinking. A single capability (“Copilot”) is delivered through three coordinated primitives — the Dynamic Action Button (context-aware entry into Chat), On-Canvas (lightweight text/object selection), and Chat (long-running reasoning) — bound by one orchestrating pattern they call Throw & Catch: the user moves a task between surfaces without losing context. The three primitives are described not as parallel features but as seams. This operationalizes several HAX guidelines at the UI layer at once — efficient dismissal/correction (G8/G9), working memory (G12), cautious adaptation (G14).
The lesson to steal: pick one orchestrating pattern and bind it to a small, fixed set of entry points. The anti-pattern Microsoft is explicitly avoiding is the one most products fall into — a different, unrelated AI affordance bolted onto every surface, so the user re-learns “the AI” in each place. Interview angle. A portfolio piece that shows the same pattern applied across three surfaces (and the reasoning for the seams) reads as systems thinking; three pretty-but-unrelated AI screens read as a junior who designs screen-by-screen.

Choosing a pattern: the selection logic libraries do not give you

A pattern library tells you what exists; it cannot tell you what to use. The selection logic that separates senior work runs on three axes. (1) Blast radius — does the action only suggest text, or does it commit a side-effect? Heavier radius pulls you up the permission ladder. (2) Capability confidence — how well does the model actually do this task (HAX G2)? Low/variable capability pulls in uncertainty patterns (confidence, abstention, regenerate). (3) Frequency and surface — empty state vs mid-flow vs inside an editor picks your suggestion tier. Run every new AI feature through these three before you open Figma, and you will choose patterns instead of decorating screens.

Three tensions you will be asked to resolve

The libraries agree on the big principle — AI features must declare their capability and limits — but they diverge in ways a senior designer should be able to name, because interviewers probe exactly these seams. Tension 1: “make clear” vs “make light of.” HAX G1/G2 push toward visible disclosure, yet NN/g finds users still cannot tell what a chatbot is for; the resolution is that disclosure must be outcome-shaped (what it’s for), not a capability inventory. Tension 2: “global controls” vs “ambient AI.” HAX G17 demands suite-wide controls while Microsoft’s Throw & Catch promises ambient, seamless interaction; the resolution is layered — ambient within a session, with a settings surface that still owns opt-out and reset. Tension 3: “junior assistant” vs “production agent.” Brad Frost calls AI a “smart-but-unsophisticated junior developer,” while shipped agents run multi-hour autonomous sessions; the resolution is that one system must hold both mental models and make the boundary legible — which is the Draftsman/Co-author/Operator stratification.
docsHAX Toolkit — 18 Guidelines for Human-AI Interaction (+ Design Patterns library)Microsoft Research + UWarticleA simplified system — the Microsoft Copilot design system (DAB, On-Canvas, Chat, Throw & Catch)Microsoft DesignarticleAI Design Patterns: How to Choose What Actually FitsThe GradientdocsThe Shape of AI — six lenses + named, shipped pattern examplesShape of AI

Checkpoint

A PM asks you to “add AI suggestions” to a document editor. You have an empty-state, a mid-writing flow, and an inline-completion need. What is the strongest first move?

ADesign one reusable “✨ Suggest” button and drop it in all three places for consistencyBName which suggestion pattern each surface needs — prompt cards (empty state), SUAs (mid-flow), smart autocomplete (inline) — and design each to its triggerCStart in Figma with a polished inline ghost-text treatment since that is the most impressive
Sign up free to answer and see why

Checkpoint

Your team is shipping an agent that can send emails, file tickets, and (rarely) delete records. A designer proposes a single confirmation dialog for all agent actions. What is the issue?

AConfirmation dialogs are outdated; use a toast notification insteadBNothing — a uniform confirmation is the safest, most consistent choiceCControl should ladder to blast radius — cheap dismiss for suggestions, accept/discard for drafts, an approval gate plus a visible autonomy budget for side-effecting actions
Sign up free to answer and see why

Checkpoint

In a review, a stakeholder cites HAX guideline “make capabilities clear” to justify a long feature list on the AI panel. Your usability sessions show users still cannot tell what the tool is for. How do you resolve the conflict?

AReshape the disclosure to be outcome-shaped (what the tool is FOR), letting NN/g’s empirical finding refine the HAX principleBFollow HAX literally and expand the capability list furtherCDrop the disclosure entirely since users ignore it
Sign up free to answer and see why

Checkpoint

You are asked to map an AI feature to a pattern before any visual work. Which set of questions best drives the selection?

AWhat color is the brand accent, which icon library do we use, and is there a dark mode?BWhich competitor has the prettiest version, and can we match it?CWhat is the blast radius, how well does the model do this task, and on which surface/frequency does it appear?
Sign up free to answer and see why

Checkpoint

A portfolio reviewer sees three AI screens in your case study, each with a different, unrelated AI affordance. What is the most likely read?

AStrong range — you can design many different AI interactionsBScreen-by-screen design without a pattern language — a junior signal the reviewer will probeCIt depends entirely on the visual polish of each screen
Sign up free to answer and see why

Interview & portfolio prep

AI-designer loops probe whether you have a vocabulary and a selection logic, not whether you can name one library. Lead with the four domains and the blast-radius ladder; cite the canonical libraries by what each is for. Answer each of these in 60–90 seconds, mechanism first, then the product implication.
  1. 01“What are the canonical AI-UX pattern libraries?” → HAX (normative: 18 guidelines + patterns), PAIR (trust calibration), NN/g (empirical), with AIUX Playground / Shape of AI as shipped-example catalogs.
  2. 02“Walk me through the main reusable AI patterns.” → four domains: suggestion, citation/provenance, permission/control, uncertainty/consistency — each with named patterns at token/component/flow grain.
  3. 03“How do you choose between AI patterns?” → blast radius, capability confidence (HAX G2), surface/frequency — not aesthetics; ladder control to side-effects.
  4. 04“Differentiate the suggestion patterns.” → prompt suggestions (content, ladder pills→library) vs smart autocomplete (semantic, inline, probabilistic) vs SUAs (persistent capability discovery).
  5. 05“Floor for agent permissions?” → HAX G7/G8/G9/G17 (invoke/dismiss/correct/global), then Approval Workflows + Autonomy Budgets that default to pause.
  6. 06“When do HAX and NN/g disagree, who wins?” → empirical reshapes normative; e.g. disclose outcome (what it’s for), not a capability inventory, when users can’t tell.
  7. 07“Why is a sparkle + warning footer not enough?” → it’s undifferentiated chrome; disclosure must be outcome-shaped and activated at the low-confidence moment, scaled to stakes.
  8. 08“What does a strong AI-pattern portfolio piece show?” → one pattern applied consistently across multiple surfaces, with the selection reasoning made explicit.
Follow-ups dig into the seams. Expect “show me where the pattern breaks across surfaces,” “what happens when the model’s capability is low here,” and “who owns this pattern in the org and how is it kept consistent?” Strong candidates tie each answer back to a named guideline or shipped product (HAX G-number, Copilot’s Throw & Catch, NN/g’s chatbot finding) rather than asserting taste. The Gradient’s framing — that the candidate must own the selection logic between patterns, not just import them — is the bar.

Could you name the four AI pattern domains, place a feature on the blast-radius ladder, and defend a pattern choice with a named guideline — on a whiteboard?

New to itGetting thereConfident

Takeaways

  • AI UX is a small, reusable vocabulary across four domains: suggestion, citation/provenance, permission/control, uncertainty/consistency.
  • HAX and PAIR are normative; NN/g is empirical — when they conflict, let evidence reshape the principle.
  • Suggestion is three patterns, not one: prompt suggestions (content), smart autocomplete (semantic/inline), SUAs (persistent discovery).
  • Ladder the control surface to blast radius — HAX G7/G8/G9/G17, then approval gates + autonomy budgets that default to pause.
  • Choose patterns by blast radius, capability confidence, and surface — then decorate; never the reverse.
  • One orchestrating pattern across surfaces (Copilot’s Throw & Catch) reads as a system; unrelated AI screens read as junior.

Next: naming & terminology — the highest-leverage, lowest-cost AI pattern work, and how the big three standardised it.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.