Lesson 4 of 6 · 48 min

User control over automation

Automation creates a “commit moment” heavier than manual work — running an AI command feels like delegating, not authoring. The three control primitives (undo, edit, override), classifying every AI action on a commitment axis, choosing the right level of automation, and the feedback loops that keep humans in charge — grounded in Copilot’s three-tier acceptance and the Smart Reply / Magic Editor control failures.

Delegating feels different from doing

When a user runs an AI command, something psychological flips: they feel they delegated, not authored. Automation creates a “commit moment” that is heavier than manual work, and if the design doesn’t hand authorship back, the user lands in one of two bad states — over-trusting the output because undoing it feels like effort, or under-trusting the whole feature because they can’t steer it. HAX makes control its own pillar (“Give the user a sense of control” is a standalone guideline) and stacks two more in the “When Wrong” phase — “Support efficient correction” and “Support flexible correction.” IBM names it the “Design for Co-Creation” principle: give users “controls that enable them to influence the generative process and work collaboratively with the AI.” This lesson is how you keep the human in charge.
Three primitives restore authorship after any AI action: undo (revert it), edit (revise it in place), and override (replace it with the user’s own input). The senior framing is that you owe the user at least one of these for every AI action, and the choice of which — and how heavy to make it — depends entirely on one variable: how reversible the action is. Get that classification wrong and you either bury a confirm step on a destructive op, or you slap a heavyweight modal on an autocomplete. Both are control-failures.
Designing the user experience of AI productsMind the Product

The commitment axis — match control weight to reversibility

Classify every AI action on a commitment axis before you design its controls. Ephemeral actions (autocomplete-style) can be inferred with keystroke shortcuts — Tab to accept, Esc to dismiss — because the cost of a wrong guess is one keystroke. Durable actions (a file written, an email sent, a record deleted, a charge made) need an explicit confirm step plus undo, because the cost of a wrong guess is real and possibly irreversible. The mistake that quietly causes most control complaints is treating all AI actions as one commitment class — applying autocomplete-grade inference to a side-effecting action, or autocomplete to nothing-grade friction to a keystroke.
code
1CONTROL GRANULARITY -> match it to COMMITMENT WEIGHT23  Action class          Examples                    Standard control4  --------------------  --------------------------  ---------------------------5  Ephemeral completion  ghost-text autocomplete     Tab / Esc shortcuts6  Visible-but-editable  AI draft, photo edit, image reject / edit / accept trio7                                                      + "revert to original" visible8  Invoked-once          chat reply, summary         regenerate, edit before send9  Agentic / side-effect send email, file PR, delete confirm step + global off-switch10                                                      + per-action undo11  Background automation smart sort, prioritization   per-item control over the12                                                      training signal; suppression1314  ANTI-PATTERNS:  a long-lived modal on a sub-second completion (too heavy)15                  an auto-applied irreversible action with no undo  (too light)
A working rule for the durable end: for destructive or side-effecting ops (delete, send, charge), require a typed or explicit confirm-step commit — never bury it in “just press enter.” And surface undo within the natural read path of the output, not behind a hamburger menu — IBM’s Carbon for AI ships an explicit “revert to AI / revert to original” toggle that bids against manual edits, so the user can always get back to either state. Interview angle. “How do you decide what needs a confirmation step?” The strong answer is the commitment axis: map every action by reversibility, gate the irreversible ones, infer the ephemeral ones — don’t apply one friction level to everything.

Choosing the level of automation

Before you design the controls, decide how much to automate — the level of automation is itself a design choice, and the right level is set by the cost of a wrong action and the user’s desire to stay in the loop. Think of a ladder: suggest (ghost text — the human does the work, AI proposes), draft (AI produces, human reviews/edits before commit), act-with-confirm (AI proposes an action, human approves the side effect), and autonomous (AI acts, human audits after). The senior instinct is that more automation is not more value — it’s more commitment weight on the user, so you climb the ladder only as far as the stakes and reversibility allow. A draft-level email assistant that needs one click to send often beats an autonomous one that sends on its own, because the click is where the user reclaims authorship.
Two design moves make higher automation survivable. First, the provenance chip that survives the artifact: when AI output is copied into a doc, an email, or a legal filing, a verified = true | false | unknown marker should travel with it — the Mata v. Avianca → Reeves arc (2023–2026) is a chain of escalating sanctions where the underlying “this was generated” signal never survived the copy-paste into Word, so the fabricated citations looked like the lawyer’s own work. Second, conversation repair: in a long agentic session, give the user a turn-level correction (“no, I meant the Q3 deck, not Q2”) that the agent carries forward, rather than forcing a restart. Without repair, errors drift across a long session and the only control left is to abandon it. Interview angle. “How much should this feature automate?” A strong answer names the automation ladder and picks a rung by stakes/reversibility, rather than defaulting to “as autonomous as possible.”

Case study: GitHub Copilot’s three-tier acceptance

GitHub Copilot is the canonical example of both efficient and flexible correction on the same suggestion — the exact pair Amershi named in the HAX guidelines, shipping at scale. The suggestion renders as dimmed ghost text at the cursor. Tab accepts the whole thing (efficient). Cmd/Ctrl+Right Arrow accepts it word by word (flexible — take the good half, stop before the bad). Esc dismisses. And Next Edit Suggestions (NES) “predicts both the location of the next edit you’ll want to make and what that edit should be.” Three acceptance tiers on one ephemeral action, each matched to how much of the guess the user actually wants. That granularity — partial acceptance, not just all-or-nothing — is what makes ghost text feel like collaboration instead of a take-it-or-leave-it gamble.

Case study: when the control exists but the UX hides it

Two failures show that control isn’t just having a toggle — it’s making it findable. Gmail users complain they “literally cannot disable Smart Reply” without digging through settings — the control technically exists but is invisible, so the feature feels imposed. Google Photos’ Magic Editor was reportedly pulled from the Pixel 8 Pro editor entirely — removing control rather than exposing it. Both failures happen at the policy layer, not the engineering layer: a settings toggle or a non-destructive variant existed, but the UX made it impossible to find or took it away. The lesson: a control a user can’t locate is, for design purposes, a control that doesn’t exist.
code
1CONTROL THAT EXISTS  vs  CONTROL THE USER CAN ACTUALLY REACH23  Hidden (feels imposed):4    Settings > Account > General > Advanced > "Smart features" (one checkbox,5    buried 4 levels deep, controls 6 unrelated features at once)67  Reachable (feels like the user's choice):8    Right on the feature, at the moment it acts:9       [ Smart Reply suggestions ]   ( on )  ->  tap to turn off here10       "Don't suggest replies in this thread"   [ per-thread ]11       "Turn off everywhere"                     [ -> one tap to the setting ]1213  PRINCIPLE: surface the control where the feature acts, scoped to the14  action, not buried in a catch-all settings page.
For any agentic feature, expose a real off-switch and, ideally, a “model” / personality toggle (e.g. “helpful” vs “concise”) — if a user can’t swap how the agent behaves, they can’t correct course; they can only abandon. And for background automation (smart sort, prioritization), the control that matters is over the training signal and suppression — “don’t do this again,” “stop prioritizing like this” — plus transparency about what was changed. Opacity about what an automation silently did is itself a control failure. Interview angle. “A user says your AI feature feels like it’s doing things to them, not for them — diagnose it.” The strong answer points at hidden or missing control and proposes surfacing it at the point of action, scoped to the action.

Feedback as control: beyond thumbs up/down

Feedback affordances are a control surface — but binary thumbs are the weakest one. Microsoft’s own “Beyond thumbs up and thumbs down” argues binary feedback “limits your ability to understand why a model succeeds or fails.” The principle: every feedback pattern needs three vectors — accept, correct, explain — and binary yes/no throws away the explanation. The richer patterns: edit-and-resend (the correction lives in the user’s artifact, like Notion AI’s inline accept/edit/decline), cite-and-challenge (per-claim ground truth, Perplexity’s wrong-sources), structured failure report (reason + severity), and conversation repair (turn-level correction across a long session).
A concrete failure shows why feedback design is control design: ChatGPT users reported that clicking thumbs-down auto-regenerated a new answer that counted against their rate limit — so negative feedback cost them quota, discouraging the very signal the product needs. The fix is a control-design fix: decouple feedback from rate limits, and prefer “edit” over “regenerate” once a user has shown engagement. Regenerate and variants (“Regenerate / Rerun / Try again,” plus alternatives) are themselves core control primitives — they let the user reject without abandoning. Interview angle. “Thumbs up/down — good enough?” A strong answer says it’s a noisy aggregate signal that users feel goes nowhere; pair it with edit-in-artifact and a reason capture so feedback both improves the model and gives the user real steering.
A control the user can’t find is a control that doesn’t exist. Surface undo, edit, and override where the feature acts and scoped to what it did — not four levels deep in a catch-all settings page.

Interview & portfolio prep

Control shows up in the design-challenge round (“design the controls for this agent”), the app critique (“what can the user do when they disagree?”), and portfolio walkthroughs (“what did you let users override, and why?”). The signal: you classify actions by commitment, you match control weight to reversibility, and you make controls findable. Drill these.
  1. 01“What are the core control primitives for an AI action?” → undo (revert), edit (revise in place), override (replace with manual input); you owe at least one per action.
  2. 02“How do you decide what needs a confirm step?” → the commitment axis: infer ephemeral actions (Tab/Esc), gate durable/irreversible ones (confirm + undo); never one friction level for all.
  3. 03“Walk me through Copilot’s acceptance model.” → ghost text + Tab (whole, efficient) + Cmd/Ctrl+Right (word-by-word, flexible) + Esc (dismiss) + NES; both efficient and flexible correction on one action.
  4. 04“A feature feels imposed on users — diagnose.” → control exists but is hidden or missing (Smart Reply buried; Magic Editor removed); surface it at the point of action, scoped to the action.
  5. 05“Is thumbs up/down enough feedback?” → no; binary throws away the ‘why’. Need accept/correct/explain — edit-in-artifact (Notion), reason capture, cite-and-challenge; decouple from rate limits.
  6. 06“When is ghost text right vs a full output pane?” → ghost text for low-stakes, high-frequency completions; a reject/edit/accept pane the moment the output is committed, sent, or quoted.
  7. 07“How do you give control over background automation?” → control the training signal and suppression (‘don’t do this again’) + transparency about what changed; opacity is itself a control failure.
  8. 08“What control does an agent need?” → a global off-switch, per-action undo, a confirm step on side effects, and ideally a personality/‘model’ toggle so users can steer, not just abandon.
Follow-ups press on tradeoffs: “Doesn’t a confirm step add friction users hate?” (yes — which is why you scope it to irreversible actions; the friction is the feature there, and you keep it off the ephemeral majority). “What did you remove?” is a real portfolio question — Bestfolios notes strong portfolios show what the designer cut, including over-heavy controls. A standout portfolio piece on control maps the actions on a commitment axis, shows the undo/edit/override you provided per class, demonstrates a findable control at the point of action, and ideally ties it to a measurable delta (fewer accidental destructive actions, higher correction-instead-of-abandon rate).
docsInline suggestions from GitHub Copilot in VS Code (ghost text, Tab, word-by-word)Visual Studio CodearticleBeyond thumbs up and thumbs down — human-centered evaluation feedback designData Science at MicrosoftdocsAI UX Patterns — Regenerate (reject without abandoning)Shape of AIdocsDesign AI feedback loops (system + user feedback as a virtuous loop)Google PAIR

Checkpoint

Your AI assistant can send emails on the user’s behalf. The current design auto-sends as soon as the draft looks complete, with an undo toast that disappears after 5 seconds. What’s the strongest critique?

AIt’s fine — an undo toast is a recovery affordance, which is what HAX requiresBSending is durable and side-effecting, so it needs an explicit confirm-step commit before it goes out — inferring “send” the way you’d infer an autocomplete is applying the wrong commitment classCThe undo toast should stay on screen for 30 seconds instead of 5DAdd a confidence score to the draft so users know whether to trust it
Sign up free to answer and see why

Checkpoint

Users love your code-completion ghost text but say accepting a suggestion is “all or nothing” — they often want just the first line. Best design change?

AAdd partial acceptance — e.g. accept word-by-word or line-by-line with a modifier key — so users can take the good part and stop, matching Copilot’s efficient + flexible correctionBShow a modal preview of the full suggestion before they accept anythingCLower the suggestion length so it never exceeds one lineDRequire users to retype the part they want
Sign up free to answer and see why

Checkpoint

Your background “smart prioritization” quietly reorders the user’s task list. Users report it feels like the app is “doing things to them.” Best control move?

AMake the reordering more accurate so users stop noticing itBTurn the feature off by default for everyoneCMake what changed transparent and give per-item control over the signal — show that it reordered, why, and let users say “don’t prioritize like this” — since opacity about a silent automation is itself a control failureDAdd a thumbs up/down on the whole task list
Sign up free to answer and see why

Checkpoint

Analytics show users rarely give thumbs-down on wrong answers, even when they clearly disagree. You discover thumbs-down triggers an auto-regenerate that counts against their usage limit. What’s the fix?

ARemove the feedback buttons since users don’t use themBDecouple feedback from the rate limit (and from forced regeneration) so giving negative feedback costs the user nothing — then add a lightweight reason capture so the signal carries the “why”CAuto-regenerate on thumbs-up instead of thumbs-downDHide the usage limit so users don’t notice the cost
Sign up free to answer and see why

Checkpoint

A user complains your AI agent “does the wrong thing and I can’t make it behave differently — I just turned it off.” Which design gap does this most directly point to?

AThe model needs more training dataBMissing course-correction controls — no override or personality/“model” toggle — so the only option left is abandonment; the fix is to let users steer behavior, not just switch the agent offCThe off-switch is too easy to findDThe agent should be fully autonomous with no controls
Sign up free to answer and see why

Could you classify an AI feature’s actions on a commitment axis, design undo/edit/override per class, make the controls findable, and defend feedback-as-control in an interview?

New to itGetting thereConfident

Takeaways

  • Automation creates a heavier commit moment; you owe the user at least one of undo / edit / override per AI action to hand authorship back.
  • Classify actions on a commitment axis: infer the ephemeral (Tab/Esc), gate the durable/irreversible (confirm step + undo). One friction level for all is the core mistake.
  • Copilot ships efficient AND flexible correction on one action — Tab (whole), Cmd/Ctrl+Right (word-by-word), Esc — which is why ghost text became the grammar of inline AI.
  • A control the user can’t find doesn’t exist: Smart Reply buried, Magic Editor removed. Surface controls at the point of action, scoped to the action.
  • Feedback is control: binary thumbs throw away the ‘why’; need accept/correct/explain, and never penalize negative feedback (don’t make thumbs-down cost quota).
  • Agents need a global off-switch, per-action undo, a confirm step on side effects, and a personality toggle — so users can steer, not just abandon.

Next: responsible & accessible AI design — bias, accessibility, and inclusive design for probabilistic UIs.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.