Lesson 6 of 6 · 50 min

Capstone: critique & redesign an AI experience

Put the whole track to work: a repeatable audit method (the four-pillar review + the timeline walk), a worked critique-and-redesign of a real support chatbot against every principle, and the portfolio/interview craft to present it — the five-stage AI design whiteboard format, the rubric interviewers grade, and what a standout AI portfolio piece actually shows.

Lesson 6 · Capstone

Critique & redesign an AI experience

The skill the whole track was building

Everything so far converges on one repeatable act: take a real, shipped AI feature, audit it against the principles, and redesign it. This is also the highest-frequency AI-design interview task (the app critique and the “design an AI feature for X” round) and the spine of a standout portfolio piece — Bestfolios’ survey of strong AI portfolios found the winners “go beyond just talking about AI” and document how they design for it. The difference between a weak critique (“here’s what the app is”) and a strong one (“here’s why they made this choice, whether I agree, and what I’d change”) is exactly the difference this lesson trains. This is the capstone: a method, a worked example, and the craft to present it.
The method has two passes. Pass 1 — the four-pillar review: run the feature through all four frameworks, because each catches what the others miss — (1) PAIR mental-model audit (does onboarding seed an honest model?), (2) HAX 18-guideline grep (which of the 18, across the four phases, are violated?), (3) NN/g anti-pattern lookup (ignored banner, ambiguous sparkle, polish-discourages-checking?), (4) IBM/Carbon primitive check (is there an AI label, an accessible regenerate, a graceful-failure component?). Pass 2 — the timeline walk: trace the feature across HAX’s four phases — Initially / During interaction / When Wrong / Over Time — and find the gap in each. The “When Wrong” phase is the one most teams skip and the one your audit should hit hardest.
Microsoft’s Guidelines for Human-AI InteractionDelta CX · Ep 267

The audit rubric — what you score

Score the feature against the through-lines of the five lessons. Each row maps to a lesson and to a concrete design move, so the audit produces a redesign backlog, not just complaints:
code
1AI-EXPERIENCE AUDIT RUBRIC  (score each: present / weak / missing)23  Dimension              Question to ask                       Lesson4  ---------------------  ------------------------------------  ----------------5  Mental model           Does onboarding/empty-state seed an   L1 mental models6                         honest model + the refusal frontier?7  Error taxonomy         Is the dominant error family named    L2 uncertainty8                         (FP/FN/halluc/low-conf) + handled?9  Confidence display     >=2 signals, bucketed, not a lone     L2 uncertainty10                         raw %?11  Graceful failure       Is the failure screen designed (HAX   L2 uncertainty12                         When-Wrong) or a generic "Oops"?13  Trust calibration      Anti over-reliance (forcing fn) AND   L3 trust14                         anti under-reliance handled?15  Transparency           Right primitive per stakes (explain   L3 trust16                         / cite / show-work) at foraging?17  Control                undo/edit/override matched to         L4 control18                         commitment; findable at point of act?19  Feedback               accept/correct/explain, not binary;   L4 control20                         not penalized?21  Accessibility          confidence-as-text, kbd controls,     L5 responsible22                         aria-live, reduced-motion?23  Responsibility         AI label, provenance, data rights,    L5 responsible24                         escape hatch, bias auditability?

Name the failure: the agentic-AI pattern catalog

The single most differentiated audit signal is the ability to name the failure pattern, not just say “the AI might be wrong.” Agentic AI ships with a known catalog (Concentrix’s “12 failure patterns”): hallucination, prompt injection, excessive agency, cascading errors, brittle automation (no human fallback), drift / data-shift exposure, lack of observability, missing evaluation harness, lack of guardrails, and more. In an audit, you tie two or three of these by name to the specific feature and to the design move that defends against each — “this agent has excessive agency because it sends without a confirm step; I’d add an act-with-confirm gate,” or “this is brittle automation — no human fallback — so I’d add the escape hatch.” That pattern-to-mechanism-to-design-move chain is what separates a strong AI critique from a list of vibes.

Worked redesign: a customer-support chatbot

Take a common, flawed pattern: a support chatbot that is the default channel, answers confidently, has a tiny “AI may be inaccurate” footer, a sparkle avatar, no visible confidence, no human handoff, and a thumbs up/down. Run the rubric. Mental model (L1): the confident, coworker-style avatar inflates expectations; no capability framing, no refusal frontier. Errors (L2): the dominant family is hallucination on policy questions (the Air Canada family) and there is no graceful failure — a wrong policy answer is delivered with the same confidence as a right one. Trust (L3): fluent, authoritative copy drives over-reliance; no citations to the actual policy docs. Control (L4): no escalation, so the only control is abandonment; binary thumbs throw away the “why.” Responsibility/a11y (L5): sparkle instead of a real AI label; no provenance; confidence (when shown) would be color-only.
code
1SUPPORT CHATBOT: BEFORE  ->  REDESIGN (mapped to principles)23  BEFORE                          REDESIGN4  ------------------------------  --------------------------------------------5  Sparkle avatar, "ask me         AI label + capability chips: "Answers from6   anything"                       OUR policy docs. Can be wrong -- I'll cite.7                                   I'll hand you to a human anytime."   (L1, L5)89  Confident policy answers,       Per-claim citation to the policy doc; on a10   no sources                      policy/billing question, "confirm this11                                   applies to you" step.                (L2, L3)1213  Tiny "AI may be inaccurate"     In-context, proportional note ONLY on14   footer (always on, ignored)     low-confidence/high-stakes answers.  (L2)1516  No handoff (default channel)    Escape hatch surfaced on low confidence,17                                   detected frustration, or any policy/$ topic:18                                   [ Talk to a human ]                  (L2, L4)1920  Thumbs up/down only             Keep thumbs, add "this was wrong because..."21                                   + edit-in-place; never costs quota.  (L4)2223  Confidence as a colored dot     Confidence as readable text; regenerate +24   (if any)                        escalate are keyboard-operable, aria-live25                                   on low-confidence updates.           (L5)
Notice the redesign is not “add more AI.” It’s expectation-setting (capability chips, refusal frontier), error UX (citations, graceful failure, proportional disclaimer), control (escape hatch, edit-in-place feedback), and responsibility (AI label, provenance, accessible controls) — the five through-lines, each tied to a concrete move and a named principle. That mapping is what an interviewer or portfolio reviewer is listening for: not “I’d make it nicer,” but “I’d close the When-Wrong gap with X, the over-reliance gap with Y, justified by Z.” Interview angle. In an app critique, narrate the audit out loud — name the family, the violated guideline, the redesign, and the tradeoff — rather than describing screens.

Presenting it: the five-stage format & the rubric

AI-design loops converge on a five-stage format for the design-challenge / whiteboard round (Alvin Wan, Exponent, Hello Interview): (1) clarify scope — ask about the user, the data, the success metric, the stakes before sketching; (2) define a minimal baseline; (3) iterate the design; (4) discuss post-deploy monitoring and out-of-distribution cases; (5) behaviorals. The single highest-ROI habit: write down the clarifying questions before you draw a screen. Candidates who ask for the user, the metric, and the stakes in the first 90 seconds score consistently higher than those who open with a UI.
Behind every round sits a shared rubric. The five canonical ML-system-design dimensions (Halina, corroborated by Exponent) are theoretical knowledge, hands-on practical, infrastructure understanding, technical leadership, and product leadership — but for an AI designer the weight shifts to the AI-UX rows: hallucination literacy (false positive vs false negative, named patterns), error UX (expectation-setting + ceding control), AI-UX pattern fluency (name the pattern — “apple picking,” “accordion editing,” ghost text — and place it in the flow), and story arc (defend the “why,” not describe the “what”). Seniority raises the bar predictably: mid shows breadth; senior adds concrete past depth and identifies the unique challenge; staff+ proposes innovative solutions proactively and leads most of the conversation.
  1. 01Clarify first: who is the user, what data/sources back the AI, what’s the success metric, what are the stakes? — before any screen.
  2. 02Open every “design an AI feature” answer with “is this false-positive-heavy or false-negative-heavy?” — it’s the first decision on every AI feature.
  3. 03Name patterns out loud: ghost text, apple picking, accordion editing, regenerate, citation chips, escape hatch — and place each in the user flow.
  4. 04Treat errors as the default case, not the edge: design the When-Wrong state, the disclaimer copy, and the escalation trigger explicitly.
  5. 05Discuss monitoring + drift: how would you know in production the model got worse, and what UX signal (feedback, escalation rate) catches it?
  6. 06Defend the why: for each choice, the principle (PAIR/HAX/NN-g/IBM) and the tradeoff — not “it looks cleaner.”

What a standout AI portfolio piece shows

The 2-minute portfolio screen is real — most reviewers triage in ~2 minutes — so lead each case study with a one-sentence user problem and a one-sentence outcome delta, then narrate 2–4 minutes of decisions. The five elements of a strong AI piece (consensus from Bestfolios + the error-UX literature): (1) a clearly stated human problem (not a feature first); (2) visible decision-making — iterations, alternatives considered, options rejected with reasons; (3) a named AI role — what was off-the-shelf vs custom vs designed, and where attribution lives; (4) error-state coverage — low-confidence patterns, fallbacks, the disclaimer copy you wrote, the escalation path (weak portfolios show only the happy path); (5) a measurable delta — usage, satisfaction, accuracy, or a credible proxy.
Frame your capstone redesign as exactly such a piece: the support-chatbot teardown is a portfolio case study if you (a) state the user’s pain (got wrong policy info, no way out), (b) show before/after with the principle behind each change, (c) name the AI’s role, (d) show the error states and escape hatch — not just the happy path, and (e) attach a delta (e.g. recovered-task rate, fewer escalated complaints, comprehension lift in testing). Interview angle. Common portfolio follow-ups — “where did the model fail and how did you handle it?”, “what did you remove?”, “what was the strongest evidence it works?” — are all answerable from a piece built this way, and unanswerable from a happy-path-only one.
A strong AI critique never describes the screen — it names the error family, the violated guideline, the redesign, and the tradeoff. A strong AI portfolio never shows only the happy path — it shows the When-Wrong state, the AI’s role, what you removed, and what changed.

Interview & portfolio prep

This whole lesson is the prep, so the drill is to run the act: critique a real AI feature live and present a redesign. These are the questions that recur in the app-critique, design-challenge, and portfolio rounds — answer each leading with the principle, then the design move.
  1. 01“Critique this AI app.” → run the four-pillar review + timeline walk; name violated HAX guidelines and the dominant error family; defend the why, agree/disagree with their choices.
  2. 02“Design an AI feature for X.” → clarify (user/data/metric/stakes) → FP-or-FN-heavy → baseline → iterate with named patterns → monitoring; not screen-first.
  3. 03“How would you handle the model being wrong?” → error family + expectation-setting (proportional disclaimer) + ceding control (escape hatch) + a named agentic failure pattern.
  4. 04“What separates a strong AI portfolio piece?” → human problem first, visible decisions, named AI role, error-state coverage, measurable delta — not a polished happy-path demo.
  5. 05“Where did the model fail and how did you handle it?” → answer from real error states you designed (low-confidence, fallback, escalation), with the copy you wrote.
  6. 06“What did you remove?” → an over-heavy control, an ambiguous sparkle, an ignored banner — show you cut, not just added.
  7. 07“How would you monitor this in production?” → feedback signal + escalation/abandon rate + drift watch; tie a UX signal to a model-health signal.
  8. 08“What was the strongest evidence your design works?” → a calibration/appropriate-reliance or recovered-task metric, ideally before/after, not “users said they liked it.”
Going deeper, expect the interviewer to escalate by level: at senior they want concrete depth from a real project (“tell me about a specific time the model failed and what you shipped”); at staff+ they want you to lead — propose the audit framework unprompted, prioritize by consequence, and connect the redesign to a north-star metric. The strongest close to any critique is the tradeoff you didn’t take: “I’d add a confirm step on policy answers, accepting the friction, because the cost of a wrong billing commitment outweighs the speed” — naming what you gave up signals senior judgment, which is exactly what separates a strong AI designer from someone re-skinning patterns.
articleHow to succeed at AI design interviews (the five-stage format)Alvin WanarticleHow Designers Are Actually Using AI in Their Portfolios (8 examples)BestfoliosdocsDesign patterns — HAX Toolkit (solutions to recurring human-AI problems)Microsoft HAX ToolkitarticleHow to pass a 2-minute UX portfolio screeningSuelyn Yu

Checkpoint

You’re given an app-critique prompt: “Critique this AI note-summarizer.” You have 30 minutes. What’s the strongest way to open?

AStart listing everything you’d visually redesign — colors, spacing, the summary card layoutBRun a structured audit out loud — clarify what the summarizer is for and its stakes, then walk the four-pillar review and the When-Wrong gap, defending why each issue mattersCPraise what works so the critique feels balanced, then stopDAsk the interviewer which features they want you to focus on and wait
Sign up free to answer and see why

Checkpoint

In a “design an AI feature for X” round, you’re asked to design AI-suggested replies for a healthcare messaging app. What should you establish before sketching anything?

AThe color palette and component library you’ll useBWhich message templates look nicestCThe user, the data/sources grounding the suggestions, the success metric, the stakes, and whether errors are false-positive- or false-negative-heavy — because in healthcare a wrong suggestion is high-consequenceDHow to make the suggestions appear as fast as possible
Sign up free to answer and see why

Checkpoint

Auditing the support chatbot from this lesson, you can only ship two fixes this quarter. Which pair best reflects triage-by-consequence?

AReplace the sparkle with a proper AI label, and add a confidence color-dotBAdd a human escape hatch on policy/billing and low-confidence answers, and add per-claim citations to the policy docs — because a confidently wrong policy answer with no way out is the Air Canada-class riskCImprove the thumbs up/down styling and animate the typing indicatorDRewrite the welcome copy and add more example prompts
Sign up free to answer and see why

Checkpoint

A reviewer skims your AI portfolio case study for 2 minutes and moves on without engaging. The piece opens with a hero shot of the polished feature and a list of tools you used. What’s the most likely fix?

AAdd more high-fidelity hero images so it looks even more polishedBLead with a one-sentence user problem and a one-sentence outcome delta, then show decisions, the AI’s role, and the error states — so the 2-minute screen sees substance, not logosCList every feature of the product in detailDMove the case study behind a password to seem exclusive
Sign up free to answer and see why

Checkpoint

At the end of a design-challenge round, the interviewer asks “what would you change if you had another month?” What kind of answer signals senior judgment?

A“Nothing — I think the design is complete as is.”B“I’d add more visual polish and animations.”C“I’d add a cognitive-forcing step on the high-stakes action and instrument the escalation rate to catch drift — and I’d accept the added friction there because a wrong commitment outweighs the speed”D“I’d run it past more stakeholders for sign-off.”
Sign up free to answer and see why

Could you audit a real AI feature with the four-pillar review, present a principle-mapped redesign, and run the five-stage design-challenge format and portfolio walkthrough end to end?

New to itGetting thereConfident

Takeaways

  • Audit method: the four-pillar review (PAIR mental model, HAX 18-guideline grep, NN/g anti-patterns, IBM/Carbon primitives) + the timeline walk across the four HAX phases.
  • The audit output is a redesign backlog triaged by consequence: fix the When-Wrong and control gaps first when money/rights/health/time are at stake.
  • A strong redesign maps each change to a principle and a tradeoff — expectation-setting, error UX, control, responsibility — never “add more AI” or “make it nicer.”
  • Run the five-stage design-challenge format: clarify (user/data/metric/stakes) before sketching; open with “FP-heavy or FN-heavy?”; name patterns; discuss monitoring.
  • A standout AI portfolio piece: human problem first, visible decisions, named AI role, error-state coverage (not just happy path), measurable delta — and it survives the 2-minute screen.
  • Senior signal = defend the why and name the tradeoff you didn’t take; weak signal = describe screens, dive into wireframes before clarifying, show only the happy path.

You’ve completed Designing Human-AI Interaction — now audit a real feature you use and turn the redesign into your portfolio piece.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.