Put it together: design an end-to-end AI interaction with a real prompt strategy and a human-in-the-loop review/approve flow where the human is also the evaluator. The five HITL patterns, confidence-based escalation, the review-gate numbers, provisional states as a system, and the portfolio piece that clears every AI-design rubric at once.
HITL is a UX strategy, not a safety net
The capstone reframing: human-in-the-loop is no longer a back-office failsafe — for a senior AI designer it is a primary interaction surface, and the choices about checkpoints and provisional states directly drive trust. The brief: design an end-to-end AI interaction — say, an assistant that drafts customer replies — with a real prompt strategy and a review/approve flow. The highest-leverage insight you’ll apply: HITL pays off most when the user is also the evaluator, and every override becomes an input to the next prompt cycle. The button that approves is also the button that improves the model.
This pulls every prior lesson into one artifact. The prompt (L1) sets the contract and voice; context (L2) grounds the draft in the customer’s history with citations; the copy (L3) handles refusals and the empty state when there’s nothing to draft; the success rubric (L4) defines what a good draft is and how you’ll measure override/edit rate; the prototype (L5) makes the review flow feel real with streaming and provisional states. The review/approve flow is the connective tissue. Interview angle. This is exactly the “design an end-to-end AI feature” case round — and the differentiator is whether your HITL design is a thoughtful interaction system or a single “Approve” button bolted on.
The scoping move that earns the first points: name the AI’s role before the UX. In the drafting assistant, the model classifies the customer’s intent and proposes a reply; the human approves or edits; the system explains why it drafted what it did (the cited sources). Stating that division out loud lets you then enumerate the non-deterministic failure surfaces — a hallucinated policy, a wrong tone for an angry customer, a refund promise it can’t authorize — and design around each. That sequence (scope the role → enumerate the failures → design the surfaces) is precisely what the case rubrics reward, and skipping it is what produces a generic chatbot.
Aufait UX names five patterns that belong in every AI product playbook. Treat them as the component set for your review/approve flow — and recent “society-in-the-loop” research reinforces that these are “core UX design strategies,” not just technical safeguards.
01Confirmation checkpoints — gate system changes, financial actions, and access control behind explicit human approval.
02Provisional AI states — “suggested,” “draft,” “pending review”: nothing reads as final until a human signs off.
03Transparent status dashboards — show what’s processing, what’s automated, and where human input is required.
04Confidence-based escalation — routine/low-risk proceeds automatically; uncertain/high-impact routes to a human.
05Context-rich alerts — include reason, impact, and urgency so reviewers aren’t buried in alert fatigue.
code
1THE FIVE HITL PATTERNS (your review/approve component set)23 1. Confirmation Checkpoints gate system changes, financial actions,4 access control -- require explicit approval5 2. Provisional AI States "suggested" / "draft" / "pending review" --6 nothing reads as final until a human says so7 3. Transparent Status show what the AI is processing, what's8 Dashboards automated, where human input is required9 4. Confidence-Based routine + low-risk -> auto; uncertain or10 Escalation high-impact -> routed to a human11 5. Context-Rich Alerts include reason + impact + urgency12 (so the human isn't drowning in alert fatigue)
Confidence-based escalation is the pattern that bridges autonomy and trust, and it’s the one to feature: “routine, low-risk actions proceed automatically; uncertain or high-impact cases are sent to users.” A UX-readable confidence score feeds a routing decision, and the interface must clearly distinguish auto-approved from human-required outcomes. The senior lever is decision traceability — recording the human’s changes to AI output with clear reasons turns the HITL system into a labeled dataset that improves the next prompt cycle. That is the “human is also the evaluator” loop made concrete: the override reason is training signal. Interview angle. Anthropic’s loop asks “what user risks would you consider?” — confidence-based escalation routing high-impact cases to a human is a precise answer.
The highest-leverage HITL design makes the user the evaluator: every approval, edit, and rejection — captured with its reason — is both the gate that ships safe output and the signal that improves the next one. One surface, two jobs; design both.
The review workflow — stages, gates, and the numbers
For document- and content-centric products, Glean’s playbook gives the operational shape and the payoff. A structured workflow with role-based stages (writer/model → editor → SME → approver) and review gates for factual and brand-voice checks delivers 40-60% faster approval cycles and cuts revision rounds from 5-7 down to 2-3. Two of Glean’s concepts map straight onto the provisional-state pattern: grounded inputs (drafts are explicitly grounded in company knowledge) and permission-aware access (only authorized materials surface). The stakes are why this matters: 54% of CMOs flag brand-voice drift as a top governance risk, and 42% of companies abandoned most gen-AI initiatives in the past year — the review flow is what keeps a feature on the right side of that line.
code
1A REVIEW/APPROVE FLOW (stages -> artifact -> pass criterion)23 Stage Reviewer Decision artifact Pass criterion4 ----------- ------------- ------------------------ ------------------------5 Draft writer/model provisional state "draft" all claims grounded in6 permitted sources7 Brand voice editor inline diff vs style guide no banned phrases; tone8 score above threshold9 Factuality SME citation graph every claim cited;10 SME override logged11 Approve approver signed manifest reason + impact + urgency12 logged1314 Each logged override/reason -> training signal for the next prompt cycle.
Notice that every row produces a logged decision with a reason — that is the data loop again, and it’s what separates a review flow that just gates output from one that compounds quality over time. The AIUXPlayground encodes the same idea at conversation level with repair contracts: for every AI action, define in advance whether repair means verify, retry with deltas, or escalate to a human. Your capstone should name the repair path for each step of the review flow, not leave it implicit. Interview angle. “Walk me through what happens when the human rejects the draft” is the follow-up that tests whether you designed the recovery and the learning loop, or just the happy approval.
A review flow with no logged reasons is a turnstile; a review flow that captures why each correction was made is a flywheel. The role-based gates that cut revision rounds from 5-7 to 2-3 only compound if every decision leaves a trace the next prompt can learn from.
Putting the end-to-end interaction together
Here is the capstone assembled, as a checklist you could defend in a case round. Prompt strategy: a slotted system prompt (identity/scope/voice/output-contract) with the customer-reply voice pinned and a refusal branch for out-of-policy asks. Context: retrieve the customer’s recent tickets and the relevant help-center articles, fenced from instructions, surfaced as inline citations on the draft. Generation UX: stream the draft into a provisional “suggested reply” state; never auto-send. HITL: confidence-based escalation — high-confidence, low-risk replies offer one-tap send; low-confidence or refund/legal-touching replies force human review with a context-rich alert. Learning loop: every edit and rejection is logged with a reason and fed back. Success metrics: accepted-without-edit rate, override rate vs a baseline, and a stratified hallucination check.
The strongest AI feature design is a loop, not a funnel: the human approves, the human corrects, and the correction — captured with its reason — makes the next draft better. Design the button that improves the model, not just the button that ships the output.
A final synthesis tension worth naming in an interview: OpenAI’s craft bar (“insane emphasis on design quality and craft,” one iconic image you can talk about for five minutes) sits against Anthropic’s “visual polish matters less than reasoning.” They are not in conflict — OpenAI measures shipped craftsmanship, Anthropic measures design reasoning under interrogation. Randy Hunt’s “flexible stance” resolves it: prove you can switch postures based on what the product needs. A capstone that ships a polished, defensible end-to-end AI interaction — and whose designer can explain every trade-off and failure state — clears both bars at once.
Provisional states as a system, not a label
The single most load-bearing pattern in the whole flow is the provisional state, and juniors treat it as a badge (“draft”) instead of a system. Done well, it’s a full lifecycle: an output is born suggested, becomes a draft the human can edit, enters pending review when routed by confidence, and only then becomes approved (or rejected, which triggers a repair path). Glean’s “grounded inputs” and “permission-aware access” map directly onto this — drafts are explicitly grounded in company knowledge and only surface authorized materials, so the provisional state also carries where this came from and who may see it.
code
1THE PROVISIONAL-STATE LIFECYCLE (not just a "draft" badge)23 suggested -> draft -> pending review -> approved4 | | | |5 model human edits routed here by acts / sends6 proposes in place LOW confidence (logged + reasoned)7 |8 +--> rejected --> repair path9 (verify / retry-with-deltas / escalate)1011 Each state also carries provenance (grounded source) + permission (who sees it).
Beyond the single user: HITL as a governance surface
For higher-stakes products, the loop widens. Recent “society-in-the-loop” research frames HITL not “merely as technical safeguards” but as “core UX design strategies” — the review flow becomes a governance surface answerable to more than the one user clicking approve. Practically, that means your decision logs aren’t just training signal; they’re an audit trail (who approved what, on what evidence, with what reason), and your confidence-based routing isn’t just convenience; it’s a policy about which actions a machine may take unsupervised. The senior designer treats the HITL surface as the place where the product’s accountability lives, and designs it to be inspectable after the fact, not just usable in the moment.
Interview & portfolio prep
This capstone is, almost exactly, the AI-design case round and the portfolio centerpiece. The published rubrics nest: Zapier/Metaview’s AI-fluency level (Resistant → Capable → Adoptive → Transformative) is the baseline; Aakash’s five dimensions grade the case; Anthropic’s five (incl. “design for safety in real-world AI contexts”) grade the reasoning; OpenAI’s pillars grade the shipped artifact; Hunt grades the portfolio narrative. You must score across all of them. Lead with the user value being traded, name the failure surfaces before the feature, and end with one open question of your own.
01“Design an end-to-end AI feature.” → prompt strategy + grounded context + provisional generation + confidence-based HITL + a logged learning loop + trust/quality metrics.
02“What user risks would you consider?” → route uncertain/high-impact cases to a human (confidence-based escalation); gate financial/access actions with confirmation checkpoints.
03“What happens when the human rejects the draft?” → a defined repair path (verify / retry with deltas / escalate) and the rejection logged with a reason as training signal.
04“How does HITL improve the product over time?” → decision traceability: every human correction + reason becomes a labeled example feeding the next prompt cycle.
05“What’s a strong portfolio piece here?” → shows the seam (human craft vs model), the failure/empty/refusal states, the HITL system, and the trade-offs and outcomes.
06“OpenAI craft vs Anthropic reasoning — which matters?” → both: shipped craftsmanship + defensible reasoning; the flexible stance (handcraft vs vibe-code, with rationale) resolves it.
07“How do you avoid alert fatigue in a review flow?” → context-rich alerts with reason + impact + urgency, and confidence routing so humans only see what needs them.
Going deeper, prepare for the layered probes that decide the loop: the resistance probe (“what would change your mind about using AI here?” — show principled judgment, not reflexive avoidance), the trade-off probe (“latency vs quality / autonomy vs escape-hatch — defend your line,” have three pre-built trade-offs ready), and the workflow probe on every tool you mention (“what you asked, what came back, what was wrong, what you changed”). The strongest candidates answer two-to-three layers of “why” beneath each opening answer and leave the team a clear next-step thread — treating feedback as signal, never as attack.
You’re designing the review/approve flow for an AI that drafts customer replies. Some replies are routine, some touch refunds and legal terms. What’s the senior HITL design?
AAuto-send everything to maximize speed, with an undo windowBRequire manual approval on every single replyCConfidence-based escalation: one-tap send for high-confidence low-risk replies; force human review for low-confidence or refund/legal-touching ones
Leadership wants the review flow to also make the AI “get better over time.” Which design choice delivers that?
ALog every human edit and rejection with a reason, and feed those reasons back into the prompt-update cycleBAdd a thumbs-up/down with no comment fieldCPeriodically swap to a newer base model
Reviewers complain they’re “drowning in approval requests.” Which two patterns most directly fix this?
ABigger fonts on the alerts and a daily digest emailBConfidence-based escalation (only uncertain/high-impact reach a human) + context-rich alerts (reason + impact + urgency)CAuto-approve anything older than an hour
In a case round you present your end-to-end AI feature. Which version of the design best matches what the rubrics reward?
AA polished chat UI with a “powered by AI” badge and a single Approve buttonBScoped AI role + enumerated failure surfaces + grounded/cited drafts + provisional/refusal/empty states + confidence routing + logged learning loop + trust metricsCA beautiful hero screen you can talk about for five minutes, with no failure states shown
An interviewer asks: “What happens when a human rejects the AI’s draft in your flow?” What’s the strongest answer?
A“It goes back to the queue and the user writes their own reply from scratch.”B“There’s a defined repair path — verify, retry with deltas, or escalate — and the rejection is logged with a reason that feeds the next prompt cycle.”C“The model tries again automatically until the human accepts something.”
A strong portfolio piece clears every rubric at once: shipped craft + defensible reasoning + visible failure states + the human/AI seam.
You’ve built the full arc — prompts, context, copy, success criteria, prototyping, and an end-to-end HITL interaction. Take it into a real portfolio piece.