Lesson 4 of 5 · 47 min

Consistency across unpredictable output

You cannot promise the same paragraph twice — but you can promise the same uncertainty surface. The five hallucination-mitigation patterns NN/g names, the bounded-container approach to variable-length output, confidence at the point of prediction, why warnings must activate at the low-confidence moment instead of sitting at page-bottom, and how a token-and-component system keeps a design coherent when the model does not.

Lesson 4 · Consistency

Coherent UI over incoherent output

Same prompt, different paragraph — every time

A designer cannot promise the same output twice: temperature, sampling, and silent model updates mean one prompt returns 12 words or 12,000, a confident answer or a hedged one, a chart or a code block. The naive response is to fight the variability; the senior response is to stop promising a stable output and start promising a stable uncertainty surface. NN/g frames this bluntly — solving hallucinations is mostly a UX problem — and Smashing reframes it as designing probabilistically rather than deterministically. This lesson is the named patterns that keep a design system coherent when its content is not.
Start with NN/g’s five concrete patterns for the single hardest case — when the model is wrong (hallucinations). (1) Intentionally uncertain language — the system speaks the way a careful expert would (“I’m not completely sure, but…”). (2) Verification tools — citations, inline links, end-of-output reference lists. (3) Source transparency — count of supporting resources, site trustworthiness, source chrome. (4) Explainable AI — surface the influential factors and warn explicitly on low-confidence predictions. (5) Decision support — debate mechanisms and multiple perspectives when stakes are high (this is HAX G10, scope services when in doubt). These five are the uncertainty vocabulary your system should be able to render anywhere.
AI design — what to consider (with Caleb Sponheim, PhD, NN/g)NN/g

The bounded container: designing for variable length

The structural problem is length: a component sized for a one-line answer breaks on a 2,000-word one, and vice versa. The convergent solution across NN/g, IBM Carbon, and the progressive-disclosure literature is a single rule: treat AI output as a bounded container — a fixed entry surface with a flexible interior (collapse / expand / revert). Indulge.digital puts it directly: progressive disclosure “turns fully auto-generated outputs into guided, usable experiences,” because generative interfaces “could quickly become chaotic” without structure — show a minimal state, then resolve or reveal more on interaction. Interview angle. “How do you design a component when the output length is unpredictable?” — the strong answer is the bounded container (fixed chrome, flexible interior, progressive disclosure), not “make the box big enough.”
code
1THE BOUNDED CONTAINER -- a fixed surface around a flexible interior23  FIXED chrome (always the same)4    +----------------------------------------------+5    | [sparkle] AI Label   [explain] [confidence]  |  <- entry surface6    +----------------------------------------------+7    |  minimal state by default                    |  <- FLEXIBLE interior8    |  v expand for full output                    |     (collapse/expand)9    |  ...                                         |10    |  [revert to AI]   [copy]   [sources: 4]      |  <- fixed actions11    +----------------------------------------------+1213  The container is the unit of design. Content varies; the surface does not.14  Progressive disclosure: minimal first, reveal/resolve on interaction.

Confidence at the point of prediction

PAIR’s most prescriptive guidance for variable output: show model confidence displays — “graphic-based indications of certainty, such as error bars or shaded areas indicating the range of alternative outcomes based on the system’s confidence level” — and put them at the prediction, not in a buried tooltip. A thin shaded band around a forecast, a low/medium/high tone on an extracted field, a “2 of 5 sources agree” marker — these are vastly stronger than a generic page-level disclaimer. The design-system move is to make confidence a visible, redundantly-encoded signal (color + label + placement) driven by a token like --ai-confidence-low, so the same confidence state renders identically everywhere it appears.
code
1CONFIDENCE-DRIVEN SURFACE -- one token, three coordinated treatments23  --ai-confidence    language            citations          marker4  ----------------   -----------------   ----------------   -------------------5  high               direct, declarative cite on request    calm; no flag6  medium             slight hedge        sources shown       subtle amber dot7  low                "I'm not sure, but" sources REQUIRED    visible warning +8                                         + source count       scope-down (HAX G10)910  The warning is a FUNCTION of the confidence token, not a page-wide footer.11  Same token drives every surface, so "uncertain" looks the same everywhere.
A subtlety the research flags: confidence has to be honest, because over-confident UI is worse than no confidence UI. NN/g’s ELIZA-effect finding is that users already over-attribute certainty to anything an AI says; a green “high confidence” chip on a wrong answer compounds the harm. So the design contract is that the confidence encoding tracks a real signal (model logits, retrieval source-count, agreement across samples), and when no honest signal exists, the surface defaults to hedged, not confident. Interview angle. “Where does the confidence number come from?” is a sharp follow-up — a strong answer names a concrete source (token probabilities, source agreement) and the default-to-uncertain rule when there isn’t one.

Persisting the surface across regeneration

Variable output also breaks over time: a user regenerates and the layout jumps, the confidence marker blinks, the citations reshuffle. The consistency contract is that the uncertainty surface persists across regeneration — if the next result is also low-confidence, the UI should not flash to confident and back. Two patterns help. A regenerate / branch carousel (AIUX Playground’s Regeneration Carousel) shows alternative outputs as siblings, so the user sees a distribution rather than one canonical “truth” that silently changed — the shipped reference is ChatGPT’s message-version arrows (the < 2/3 > navigator that appears after a regenerate, letting you step between sibling answers instead of overwriting the last), and Midjourney’s 4-image variation grid, which renders the distribution as the default output rather than pretending to one answer. The consistency failure to design against is the opposite: a surface that silently swaps the previous response on regenerate, with no sibling history and a layout that reflows each time, trains users to distrust it — a result they liked can vanish with no way back. And a “show your reasoning” affordance (chain-of-thought trace) is itself a reusable pattern — it gives the user a stable way to interrogate a varying answer. Interview angle. “What happens when the user hits regenerate?” tests whether you’ve thought past the happy path; persistent uncertainty + a branch carousel (ChatGPT’s sibling arrows, not a silent overwrite) is the senior answer.

Case study: pattern libraries are the consistency mechanism

How do teams actually hold this line at scale? By making consistency a first-class axis of the system, not a tag. Shape of AI organises an entire lens around Trust builders; AIUX Playground catalogs 170+ patterns with Trust as a category (Citations, Confidence Score, Source Browser, Knowledge Graph). IBM Carbon makes it structural: every AI-aware component is required to embed an AI Label and an explainability popover — so an engineer cannot ship an AI surface without the uncertainty affordances, because the component won’t let them. That is the lesson: consistency is enforced by the system’s defaults and requirements, not by designer vigilance. The Gradient’s warning still applies — a library of pretty patterns is not a strategy; the team must own which uncertainty pattern fits which stakes.
Tie it back to tokens, because that is where consistency physically lives. The minimum viable uncertainty system is a small set of primitives the research points to: --ai-confidence-{low,medium,high} tokens that drive color/label/placement; a <ConfidenceBadge> that composes into any surface; a <CitationList> with hover-preview; and an uncertainty-aware container. Defined once, these render the same confidence and the same hedging everywhere — so when the model is uncertain on screen A and screen B, the user sees the same visual language, even though the words differ. Interview angle. “How do you keep AI output consistent across a design system?” → a layered token taxonomy plus required-by-default uncertainty components, not manual review.

The unhappy path is the main path

For probabilistic systems, the empty / error / refused / failed states are not edge cases — they are frequent, so the design system must treat them as first-class, consistent states rather than afterthoughts. The skill taxonomy is blunt about it: design for model variability and wrong answers as the default case. A coherent system names four recovery states and gives each a consistent treatment: no result (the model has nothing — offer a reframed query, not a blank box), low-confidence result (render it, hedged, with the warning surface on), refusal (the model declines — explain why and offer an alternative path, per HAX G11 explanations), and hard failure (timeout/error — a graceful fallback, never a stack trace). The same four states, treated the same way across every AI surface, are what keep the experience coherent when the model is not.
This is also where consistency meets trust calibration as a measurable outcome. The “Evaluating AI experience quality” discipline treats hallucination and failure as UX defects to be measured, not bugs to hand-wave: you instrument whether users over-trust a confident-looking wrong answer or under-trust a hedged-but-correct one, and you tune the uncertainty surface against that. Interview angle. “How would you QA an AI feature as a designer?” — a strong answer is designer-side eval of output quality plus a trust-calibration metric, not “file a bug when it hallucinates.” A portfolio piece that shows the failure states designed and the trust-calibration measured reads far more senior than one that only shows the happy path.
Solving LLM hallucinations is (mostly) a UX problem. — the through-line across NN/g, Smashing, and the probabilistic-design literature: you cannot make the output certain, so you design a surface that tells the truth about how certain it is, every time, on every screen.
articleAI Hallucinations: What Designers Need to Know (the five mitigation patterns)Nielsen Norman GrouparticleDesigning With Uncertainty: how AI supercharges probabilistic thinkingSmashing MagazinearticleProgressive Disclosure is the Design Pattern for AI-generated interfacesIndulge DigitalarticleWhat Is Your Site’s AI Chatbot for? Users Can’t TellNielsen Norman Group

Checkpoint

A summary component looks great in design with a 3-sentence output, but in production some summaries run 1,500 words and the layout breaks. What is the right fix?

AA bounded container: fixed chrome (label, confidence, actions) with a flexible, progressively-disclosed interior (collapse/expand)BCap every summary at 3 sentences so it always fits the componentCMake the component tall enough to fit the longest possible output
Sign up free to answer and see why

Checkpoint

Legal asks you to “add a disclaimer that the AI can be wrong.” You add a footer; users still over-trust confident-looking wrong answers. What does the research recommend instead?

AMake the footer larger and put it at the top of the pageBDrive the warning off the confidence state — hedge language, stronger citations, and a visible uncertainty marker activate when confidence is lowCAdd the disclaimer to the onboarding so users see it once and remember
Sign up free to answer and see why

Checkpoint

PAIR recommends model confidence displays. Where should a confidence signal live for an extracted invoice field that the model is unsure about?

AIn a global settings page where users can read about the model’s overall accuracyBIn a single tooltip that also holds capability, disclosure, and change infoCAt the point of prediction — on the field itself, redundantly encoded (color + label + placement) and driven by a confidence token
Sign up free to answer and see why

Checkpoint

A user hits “regenerate” and the new answer is also low-confidence, but the UI flips to a clean, confident-looking layout each time before settling. Why is that a defect?

AIt is not a defect — a fresh confident layout feels more responsiveBThe uncertainty surface should persist across regeneration; flashing confident-then-uncertain misrepresents certainty and erodes trustCRegenerate should be removed so users cannot see the variation
Sign up free to answer and see why

Checkpoint

You want uncertainty affordances to appear on every AI surface without relying on designers to remember. What is the most reliable mechanism?

AA written guideline in the design-system docs reminding designers to add confidence and citationsBA periodic manual audit that flags AI surfaces missing uncertainty UICMake uncertainty structural — every AI-aware component embeds the AI label + explainability + confidence by default (Carbon for AI)
Sign up free to answer and see why

Interview & portfolio prep

Consistency-under-uncertainty is the prompt that most directly tests whether you can design for probabilistic systems. Lead with “you can’t promise the same output, only the same uncertainty surface,” then name the mechanisms: bounded container, confidence at the prediction, confidence-driven warnings, persistence across regeneration, and structural enforcement.
  1. 01“How do you keep AI output consistent across a design system?” → stable uncertainty surface: layered tokens + required-by-default uncertainty components, not a stable output.
  2. 02“Design for unpredictable output length?” → a bounded container — fixed chrome, flexible interior, progressive disclosure (minimal first, reveal on interaction).
  3. 03“Where do confidence signals go?” → at the point of prediction (PAIR: error bars / shaded ranges), redundantly encoded, driven by a confidence token.
  4. 04“A generic ‘AI can be wrong’ footer isn’t working — why?” → blanket warnings become ignorable; activate mitigation at the low-confidence moment, scaled to stakes.
  5. 05“What are NN/g’s hallucination patterns?” → uncertain language, verification tools, source transparency, explainable AI, multi-perspective decision support (HAX G10).
  6. 06“What happens on regenerate?” → the uncertainty surface persists; a regeneration/branch carousel shows alternatives as siblings, not one silently-changing truth.
  7. 07“How is consistency enforced at scale?” → structurally — components embed AI label + explainability + confidence by default (Carbon), so it can’t be skipped.
  8. 08“What does a strong consistency portfolio piece show?” → the same component handling short/long/low-confidence/failed output with one coherent uncertainty surface.
Follow-ups push on the unhappy path and the system mechanics: “show me the empty / error / refused states,” “how does the token drive the warning,” “what stops a team from shipping an AI surface without these affordances?” Strong candidates tie each answer to a token or a required component default and to a named source (NN/g’s five patterns, PAIR’s confidence displays, Carbon’s embedded-label requirement), and they treat hallucination and failure as design concerns to be measured, not edge cases to hand-wave.

Could you design one AI component that stays coherent across short, long, low-confidence, and failed output — and explain the token/enforcement mechanism?

New to itGetting thereConfident

Takeaways

  • You can’t promise the same output — only the same uncertainty surface; design the frame to be stable, not the content.
  • NN/g’s five hallucination patterns: uncertain language, verification tools, source transparency, explainable AI, multi-perspective decision support.
  • Solve variable length with a bounded container — fixed chrome, flexible interior, progressive disclosure.
  • Put confidence at the point of prediction, redundantly encoded and token-driven; activate warnings at the low-confidence moment, not at page-bottom.
  • Persist the uncertainty surface across regeneration; show alternatives with a branch carousel, not one silently-changing answer.
  • Enforce consistency structurally — required-by-default uncertainty affordances (Carbon) beat designer vigilance.

Next: the capstone — design one AI-suggestion pattern and apply it consistently across three surfaces.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.