Lesson 1 of 6 · 47 min

How prompts shape UX

A prompt is not a magic incantation — it is the most concentrated piece of product behavior a designer ever ships. How one clause changes what users feel, the four-part anatomy you design slot by slot, real leaked system prompts (Claude, v0, Notion) as case studies, and the interview probe that opens every AI-design loop.

Why the prompt is a design surface, not a config string

A senior reframing up front: in a model-mediated product, the system prompt is simultaneously UI copy, brand voice, and executable instruction. Every line you write is interpreted by a probabilistic model and turned into behavior the user feels. That is why Anthropic describes the system prompt as “a massive framework that defines how the model works, what tools it has access to.” The designer who treats it like a magic incantation ships drift; the one who treats it like a design-system token — owned, versioned, testable — ships a product. This lesson is the ground floor: you move from “writing copy” to “writing behavior shaped by copy.”
Here is the move that makes it concrete. A single guardrail clause in Claude’s system prompt — “Claude never starts its response by saying a question or idea or observation was good, great, fascinating, profound, excellent, or any other positive adjective” — changes the perceived personality of the entire product. No screen changed. No component changed. One sentence of prompt copy removed the sycophantic opener millions of users would otherwise read on every turn. Interview angle. When an interviewer asks “where does the brand voice of an AI product live?”, the strong answer names the system prompt as the highest-leverage surface — and can point to a clause like that one.
AI Is Your New Design MaterialJosh Clark / AIGA

The four-part anatomy — design each slot, name each owner

Across the four major prompting guides (dair-ai, Anthropic, OpenAI, Microsoft) the same skeleton recurs. dair-ai decomposes a prompt into four elements — Instruction, Context, Input Data, Output Indicator; OpenAI’s developer-message model maps it to Identity, Instructions, Examples, Context. Treat this as the base unit of AI UX: every clause has a function, an owner (brand, product, safety, or engineering), and a stable id so it can be changed, tested, and reverted — exactly like a Figma component.
The four guides each emphasize a different facet, and a senior designer borrows from all of them rather than picking one. Worth knowing how they divide labor, because interviewers cite them and the strongest designers mutate them into one operating model:
  1. 01dair-ai → the cognitive vocabulary: elements of a prompt, the technique zoo (few-shot, CoT, ReAct), and context engineering as a discipline.
  2. 02OpenAI → structure: the Identity / Instructions / Examples / Context developer-message skeleton, plus RAG and prompt-caching affordances.
  3. 03Microsoft Foundry → process: clear instructions, repeat the rule at the end, break the task down, specify the output schema.
  4. 04Anthropic → the highest-trust behavior bar: refusal as a UX moment, honesty/anti-sycophancy clauses, and persona design.
code
1THE PROMPT IS A HIERARCHY OF DESIGN TOKENS -- not one free-text blob.23  SLOT              WHAT IT CONTROLS                 LIKELY OWNER4  ---------------   ------------------------------   -----------------5  Identity          who the assistant is             brand + product6  Instructions      scope: what it can / cannot do   product + safety7  Voice / character how it expresses itself          brand (design)8  Examples          format + edge-case behavior      design + eng9  Context           retrieved data, date, user role  engineering10  Output contract   shape the renderer will parse    design + eng1112  Rule: if a clause has no owner and no eval, it is a bug waiting to ship.
The canonical dair-ai illustration is the classifier pattern: Classify the text into neutral, negative, or positive. Text: I think the food was okay. Sentiment: — the trailing word “Sentiment:” is the Output Indicator, and the colon plus whitespace primes the model to continue with exactly one label. That tiny piece of formatting is design: it constrains the response shape so the downstream UI never has to parse a paragraph. dair-ai and Microsoft both land on the same practical band for demonstrations — 3 to 5 examples wrapped in delimiters — which a designer should curate the way they would a screenshot gallery: cover the edge cases, keep the format identical, prefer diverse variants over the one canonical happy path.
Microsoft and Harvard’s “Anatomy of a Prompt” distill the same skeleton to TIC — Task, Instructions, Context. The senior move is to make TIC a Figma-frame-shaped template in your design repository: every prompt spec labels each slot, names its owner, and links to its eval criteria. That turns prompts from one-off prose into a maintained design system, the same way a component library beats ad-hoc screens. When a clause regresses behavior, you change the named region and revert it — you are not editing a wall of text and hoping.

Case study: Vercel v0 — the system prompt IS the product surface

v0 is a product almost entirely defined by its prompt. Its identity sits in a <v0_info> block (“You are v0, an AI assistant created by Vercel to be helpful, harmless, and honest” … “emulate the world’s most proficient developers”). Its output morphology — React, Next.js, Mermaid diagrams, HTML, Markdown — is governed by <v0_codeBlock_types> paired with explicit metadata directives on each fenced block (project, file, and type attributes). The look-and-feel users experience is not separate from the prompt; the prompt is what causes the components to render. When a v0 output looks on-brand and pastes cleanly into a Next.js app, that is a prompt-design win, not a model win.
A future-facing line from Anthropic’s frontend-aesthetics guidance sharpens the point for designers: it warns the model away from “generic, ‘on distribution’ outputs” that produce “what users call the ‘AI slop’ aesthetic. Avoid this: make creative, distinctive frontends that surprise and delight.” That is a taste directive expressed as prompt copy — the same instinct a designer brings to a brief, encoded so a model executes it. Interview angle. “Show me where design taste enters an AI product that has no traditional UI.” Point here: taste lives in the output-contract and voice clauses of the prompt.
The prompt is the product. v0’s entire interface — what it builds, how it formats, how on-brand it feels — is produced by structured prompt clauses, not by a separate design layer bolted on afterward.

Behavior-shaping clauses: the levers and their UX side effects

Designers must know the standard techniques because each is a behavioral lever with a predictable UX tradeoff — not an engineering detail to hand off. The four guides cluster them consistently; what a designer cares about is the column on the right: what does this do to the experience?
code
1PROMPT TECHNIQUE -> WHAT IT CHANGES FOR THE USER23  Technique            UX effect                         The cost you design around4  ------------------   -------------------------------   -----------------------------5  Zero-shot            terse, fast, no setup             brittle on nuance / edge cases6  Few-shot (3-5 ex.)   locks format + tone               wrong examples bias every reply7  Chain-of-thought     more correct on hard tasks        slower, verbose, leaks reasoning8  ReAct (think+act)    visible thought/action trace      a loading-state design decision9  Structured output    deterministic, parseable shape    sterilizes expressive writing10  Role / system        sets identity, scope, tone        drifts without boundary examples1112  A designer picks the lever by the experience they want, then owns the side effect.
Take ReAct concretely. dair-ai’s pattern produces a visible thought / action / observation trajectory (“Thought 1: I need to search… Action 1: Search[…] Observation 1: …”). Whether the product surfaces those steps as transparent reasoning, hides them behind a “thinking…” loading state, or summarizes them into a one-line status — that is a pure design decision, made at the prompt layer, not an engineering afterthought. The AI UX Playground catalogs this exact surface as “status steps,” one of 30 named patterns. Interview angle. “The agent takes 8 seconds — what do you put on screen?” The strong answer ties the loading experience to whether you expose, summarize, or hide the ReAct trace.

Case study: Notion AI & Duolingo — leaked prompts as voice specs

Two more production prompts that leaked and read as voice specifications. Notion AI’s system prompt pins it to a workspace-assistant register and constrains output to Notion-flavored Markdown — the copy and the renderer are designed together. Duolingo’s “Lily” character prompt encodes a deadpan, sarcastic persona so consistently that users recognized the voice from a single leaked block. The lesson for designers: a leaked prompt is a portfolio-grade artifact — it shows, line by line, how a team turned a brand personality into executable behavior. Studying them is the fastest way to calibrate your own.
Read leaked system prompts the way a film student reads screenplays. Notion’s, Duolingo’s Lily, Claude’s — each is a finished example of a brand personality compiled into instructions a model executes. The voice you feel in the product is right there in the text.
There is a deeper theory under this. Anthropic’s Persona Selection Model research frames an AI assistant as “an enacted human-like persona” learned in pre-training, so every prompt cue activates a character. The designer implication is to run a “psychological implication” check on your clauses: not just “is this behavior good or bad?” but “what does this behavior imply about who the assistant is?” Deliberately designing a positive archetype dilutes the gravitational pull of the HAL-9000 / Terminator tropes the model has also absorbed. This is character design with a model as the actor.

The process ladder: a designer iterates a prompt empirically

Prompting is a lab, not a vibe workshop. OpenAI’s process ladder is the order to climb: “Start with zero-shot, then few-shot, neither of them worked, then fine-tune.” A designer translates that to: write the simplest instruction, add 3-5 examples only if the format/voice won’t hold, and reserve fine-tuning for a fixed style or skill — not for volatile facts (that’s retrieval, Lesson 2). Anthropic’s own loop adds a self-correction chain — generate a draft → review against the criteria → refine — which is exactly how a designer iterates a layout, just with text outputs and a rubric instead of pixels and a critique.
code
1THE PROMPT ITERATION LADDER (climb only as far as you must)23  1. Zero-shot         one clear instruction. cheapest. try this first.4       |               (fails on nuance / inconsistent format)5  2. Few-shot          add 3-5 delimited examples to lock format + voice.6       |               (still inconsistent on edge cases)7  3. Repeat + structure  restate key rules at the END; add an output schema;8       |                 break the task into steps (Microsoft Foundry).9  4. Fine-tune         only for a FIXED style/skill, never volatile facts.1011  Each rung is a design decision with a cost; stop at the first that passes the rubric.

Interview prep

AI-design loops open by testing whether you treat the prompt as a design material with real leverage — or as a mysterious black box. Interviewers (and portfolio reviewers) probe four things: can you name where voice/behavior lives, can you decompose a prompt into ownable slots, can you point to real product examples, and can you reason about a clause’s second-order UX effects. Lead with the mechanism, then the experience it produces. Interview angle. Aakash Gupta’s AI-product-design rubric explicitly penalizes “surface-level” AI mentions and rewards demonstrated “AI technical fluency” — naming real prompt structures is how you clear that bar.
  1. 01“Where does an AI product’s brand voice live?” → the system prompt’s identity + voice clauses; cite Claude’s anti-sycophancy line as a one-clause personality change.
  2. 02“Walk me through the anatomy of a prompt.” → Identity / Instructions / Examples / Context (+ output contract); each a slot with an owner and an eval.
  3. 03“Give a real example of the prompt being the product.” → v0: XML-tagged blocks govern what renders and how on-brand it feels.
  4. 04“Few-shot vs zero-shot — when, as a designer?” → few-shot to lock format/tone (3-5 curated examples); zero-shot for fast, simple, low-nuance tasks.
  5. 05“The agent is slow — what goes on screen?” → a ReAct-trace decision: expose, summarize, or hide thought/action/observation as status steps.
  6. 06“How is designing a prompt different from writing UI copy?” → it’s executable + probabilistic; you’re authoring behavior, and you must test it empirically.
  7. 07“Why structure a prompt with XML/delimiters?” → it makes slots reviewable, swappable, and parseable; the output contract keeps the renderer from breaking.
  8. 08“How do you keep a prompt from drifting on a model update?” → version it, gate changes with an eval set, and pin owners per clause — treat it like a design token.
Going deeper, expect follow-ups that separate “read a thread” from “shipped one”: “you changed one clause and the whole tone shifted — why?” (the model generalizes a voice cue across every turn; small copy, large blast radius); “how would you A/B a voice change?” (rubric-scored sample of outputs, not a vibe check — Lesson 4); and “who signs off on the system prompt?” (brand owns voice clauses, safety owns refusals, engineering owns context — name the matrix). In every case, attach the clause to the surface it controls.
articleHighlights from the Claude 4 system prompt (annotated)Simon WillisonarticleThe full prompt of v0.dev — structured output as product surfacebaoyu.iodocsPrompt engineering overview (anatomy, roles, structure)AnthropicpaperThe persona selection model — why AI assistants enact a characterAnthropic

Checkpoint

A PM says “the assistant feels too fawning — it praises every message before answering.” You own the design fix. What is the highest-leverage move?

AAdd a sentence to the system prompt that bans opening with positive adjectives, then eval a sample of outputsBRedesign the chat bubble styling to feel more neutralCAdd a “tone” dropdown so users can pick a less enthusiastic mode
Sign up free to answer and see why

Checkpoint

You inherit a 600-word system prompt that is one undifferentiated paragraph. The team can’t tell why a recent edit broke the output format. What’s the senior first move?

ARewrite it shorter so it’s easier to readBAsk engineering to add automated tests and leave the prompt as-isCDecompose it into labeled slots (identity, instructions, voice, examples, output contract) with an owner per slot
Sign up free to answer and see why

Checkpoint

Your AI feature must return answers the UI renders as a strict card (title, 3 bullets, one CTA). Outputs keep arriving as free-form prose that breaks the card. Best prompt-design fix?

AIncrease the model’s temperature so it’s more creative with structureBSpecify an explicit output contract (schema/format + an Output Indicator) and give 3-5 examples in the exact card shapeCAdd a post-processor that tries to reformat prose into a card
Sign up free to answer and see why

Checkpoint

In an AI-design interview you’re asked: “Show me where design taste enters a product like v0 that has no conventional UI to art-direct.” Strongest answer?

ATaste enters in the marketing site and onboarding, not the model outputBIt doesn’t — model output quality is purely an engineering/eval concernCIn the voice and output-contract clauses — e.g. an anti-“AI slop” aesthetic directive that steers the model toward distinctive, on-brand output
Sign up free to answer and see why

Checkpoint

A reviewer asks how you’d ship a brand-voice change to the system prompt without regressing behavior for millions of users. Best process?

AVersion the prompt, run a rubric-scored eval on a held-out sample before/after, and gate the rollout on itBPush the change live and watch support tickets for a few daysCAsk the model to review its own new prompt and approve it
Sign up free to answer and see why

Could you decompose a real system prompt into ownable design slots and name how one clause changes the user experience?

New to itGetting thereConfident

Takeaways

  • The system prompt is UI copy + brand voice + executable behavior at once — the highest-leverage design surface in an AI product.
  • Design it slot by slot: identity, instructions, voice, examples, output contract — each with an owner and an eval.
  • One clause has a huge blast radius (Claude’s anti-sycophancy line) — small copy, product-wide personality change.
  • v0 proves the prompt IS the product surface; taste enters through voice + output-contract clauses, not a separate layer.
  • Pick prompt techniques by the experience you want (ReAct trace = a loading-state decision); own the side effect.
  • Replace vague adjectives with behavioral clauses + examples, and version the prompt like a design token.

Next: what context, retrieval, and tools change about the experience — and the trust/latency surfaces each one adds.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.