A prompt is not a magic incantation — it is the most concentrated piece of product behavior a designer ever ships. How one clause changes what users feel, the four-part anatomy you design slot by slot, real leaked system prompts (Claude, v0, Notion) as case studies, and the interview probe that opens every AI-design loop.
Why the prompt is a design surface, not a config string
A senior reframing up front: in a model-mediated product, the system prompt is simultaneously UI copy, brand voice, and executable instruction. Every line you write is interpreted by a probabilistic model and turned into behavior the user feels. That is why Anthropic describes the system prompt as “a massive framework that defines how the model works, what tools it has access to.” The designer who treats it like a magic incantation ships drift; the one who treats it like a design-system token — owned, versioned, testable — ships a product. This lesson is the ground floor: you move from “writing copy” to “writing behavior shaped by copy.”
Here is the move that makes it concrete. A single guardrail clause in Claude’s system prompt — “Claude never starts its response by saying a question or idea or observation was good, great, fascinating, profound, excellent, or any other positive adjective” — changes the perceived personality of the entire product. No screen changed. No component changed. One sentence of prompt copy removed the sycophantic opener millions of users would otherwise read on every turn. Interview angle. When an interviewer asks “where does the brand voice of an AI product live?”, the strong answer names the system prompt as the highest-leverage surface — and can point to a clause like that one.
The four-part anatomy — design each slot, name each owner
Across the four major prompting guides (dair-ai, Anthropic, OpenAI, Microsoft) the same skeleton recurs. dair-ai decomposes a prompt into four elements — Instruction, Context, Input Data, Output Indicator; OpenAI’s developer-message model maps it to Identity, Instructions, Examples, Context. Treat this as the base unit of AI UX: every clause has a function, an owner (brand, product, safety, or engineering), and a stable id so it can be changed, tested, and reverted — exactly like a Figma component.
The four guides each emphasize a different facet, and a senior designer borrows from all of them rather than picking one. Worth knowing how they divide labor, because interviewers cite them and the strongest designers mutate them into one operating model:
01dair-ai → the cognitive vocabulary: elements of a prompt, the technique zoo (few-shot, CoT, ReAct), and context engineering as a discipline.
02OpenAI → structure: the Identity / Instructions / Examples / Context developer-message skeleton, plus RAG and prompt-caching affordances.
03Microsoft Foundry → process: clear instructions, repeat the rule at the end, break the task down, specify the output schema.
04Anthropic → the highest-trust behavior bar: refusal as a UX moment, honesty/anti-sycophancy clauses, and persona design.
code
1THE PROMPT IS A HIERARCHY OF DESIGN TOKENS -- not one free-text blob.23 SLOT WHAT IT CONTROLS LIKELY OWNER4 --------------- ------------------------------ -----------------5 Identity who the assistant is brand + product6 Instructions scope: what it can / cannot do product + safety7 Voice / character how it expresses itself brand (design)8 Examples format + edge-case behavior design + eng9 Context retrieved data, date, user role engineering10 Output contract shape the renderer will parse design + eng1112 Rule: if a clause has no owner and no eval, it is a bug waiting to ship.
The canonical dair-ai illustration is the classifier pattern: Classify the text into neutral, negative, or positive. Text: I think the food was okay. Sentiment: — the trailing word “Sentiment:” is the Output Indicator, and the colon plus whitespace primes the model to continue with exactly one label. That tiny piece of formatting is design: it constrains the response shape so the downstream UI never has to parse a paragraph. dair-ai and Microsoft both land on the same practical band for demonstrations — 3 to 5 examples wrapped in delimiters — which a designer should curate the way they would a screenshot gallery: cover the edge cases, keep the format identical, prefer diverse variants over the one canonical happy path.
Microsoft and Harvard’s “Anatomy of a Prompt” distill the same skeleton to TIC — Task, Instructions, Context. The senior move is to make TIC a Figma-frame-shaped template in your design repository: every prompt spec labels each slot, names its owner, and links to its eval criteria. That turns prompts from one-off prose into a maintained design system, the same way a component library beats ad-hoc screens. When a clause regresses behavior, you change the named region and revert it — you are not editing a wall of text and hoping.
Case study: Vercel v0 — the system prompt IS the product surface
v0 is a product almost entirely defined by its prompt. Its identity sits in a <v0_info> block (“You are v0, an AI assistant created by Vercel to be helpful, harmless, and honest” … “emulate the world’s most proficient developers”). Its output morphology — React, Next.js, Mermaid diagrams, HTML, Markdown — is governed by <v0_codeBlock_types> paired with explicit metadata directives on each fenced block (project, file, and type attributes). The look-and-feel users experience is not separate from the prompt; the prompt is what causes the components to render. When a v0 output looks on-brand and pastes cleanly into a Next.js app, that is a prompt-design win, not a model win.
A future-facing line from Anthropic’s frontend-aesthetics guidance sharpens the point for designers: it warns the model away from “generic, ‘on distribution’ outputs” that produce “what users call the ‘AI slop’ aesthetic. Avoid this: make creative, distinctive frontends that surprise and delight.” That is a taste directive expressed as prompt copy — the same instinct a designer brings to a brief, encoded so a model executes it. Interview angle. “Show me where design taste enters an AI product that has no traditional UI.” Point here: taste lives in the output-contract and voice clauses of the prompt.
The prompt is the product. v0’s entire interface — what it builds, how it formats, how on-brand it feels — is produced by structured prompt clauses, not by a separate design layer bolted on afterward.
Behavior-shaping clauses: the levers and their UX side effects
Designers must know the standard techniques because each is a behavioral lever with a predictable UX tradeoff — not an engineering detail to hand off. The four guides cluster them consistently; what a designer cares about is the column on the right: what does this do to the experience?
code
1PROMPT TECHNIQUE -> WHAT IT CHANGES FOR THE USER23 Technique UX effect The cost you design around4 ------------------ ------------------------------- -----------------------------5 Zero-shot terse, fast, no setup brittle on nuance / edge cases6 Few-shot (3-5 ex.) locks format + tone wrong examples bias every reply7 Chain-of-thought more correct on hard tasks slower, verbose, leaks reasoning8 ReAct (think+act) visible thought/action trace a loading-state design decision9 Structured output deterministic, parseable shape sterilizes expressive writing10 Role / system sets identity, scope, tone drifts without boundary examples1112 A designer picks the lever by the experience they want, then owns the side effect.
Take ReAct concretely. dair-ai’s pattern produces a visible thought / action / observation trajectory (“Thought 1: I need to search… Action 1: Search[…] Observation 1: …”). Whether the product surfaces those steps as transparent reasoning, hides them behind a “thinking…” loading state, or summarizes them into a one-line status — that is a pure design decision, made at the prompt layer, not an engineering afterthought. The AI UX Playground catalogs this exact surface as “status steps,” one of 30 named patterns. Interview angle. “The agent takes 8 seconds — what do you put on screen?” The strong answer ties the loading experience to whether you expose, summarize, or hide the ReAct trace.
Case study: Notion AI & Duolingo — leaked prompts as voice specs
Two more production prompts that leaked and read as voice specifications. Notion AI’s system prompt pins it to a workspace-assistant register and constrains output to Notion-flavored Markdown — the copy and the renderer are designed together. Duolingo’s “Lily” character prompt encodes a deadpan, sarcastic persona so consistently that users recognized the voice from a single leaked block. The lesson for designers: a leaked prompt is a portfolio-grade artifact — it shows, line by line, how a team turned a brand personality into executable behavior. Studying them is the fastest way to calibrate your own.
Read leaked system prompts the way a film student reads screenplays. Notion’s, Duolingo’s Lily, Claude’s — each is a finished example of a brand personality compiled into instructions a model executes. The voice you feel in the product is right there in the text.
There is a deeper theory under this. Anthropic’s Persona Selection Model research frames an AI assistant as “an enacted human-like persona” learned in pre-training, so every prompt cue activates a character. The designer implication is to run a “psychological implication” check on your clauses: not just “is this behavior good or bad?” but “what does this behavior imply about who the assistant is?” Deliberately designing a positive archetype dilutes the gravitational pull of the HAL-9000 / Terminator tropes the model has also absorbed. This is character design with a model as the actor.
The process ladder: a designer iterates a prompt empirically
Prompting is a lab, not a vibe workshop. OpenAI’s process ladder is the order to climb: “Start with zero-shot, then few-shot, neither of them worked, then fine-tune.” A designer translates that to: write the simplest instruction, add 3-5 examples only if the format/voice won’t hold, and reserve fine-tuning for a fixed style or skill — not for volatile facts (that’s retrieval, Lesson 2). Anthropic’s own loop adds a self-correction chain — generate a draft → review against the criteria → refine — which is exactly how a designer iterates a layout, just with text outputs and a rubric instead of pixels and a critique.
code
1THE PROMPT ITERATION LADDER (climb only as far as you must)23 1. Zero-shot one clear instruction. cheapest. try this first.4 | (fails on nuance / inconsistent format)5 2. Few-shot add 3-5 delimited examples to lock format + voice.6 | (still inconsistent on edge cases)7 3. Repeat + structure restate key rules at the END; add an output schema;8 | break the task into steps (Microsoft Foundry).9 4. Fine-tune only for a FIXED style/skill, never volatile facts.1011 Each rung is a design decision with a cost; stop at the first that passes the rubric.
Interview prep
AI-design loops open by testing whether you treat the prompt as a design material with real leverage — or as a mysterious black box. Interviewers (and portfolio reviewers) probe four things: can you name where voice/behavior lives, can you decompose a prompt into ownable slots, can you point to real product examples, and can you reason about a clause’s second-order UX effects. Lead with the mechanism, then the experience it produces. Interview angle. Aakash Gupta’s AI-product-design rubric explicitly penalizes “surface-level” AI mentions and rewards demonstrated “AI technical fluency” — naming real prompt structures is how you clear that bar.
01“Where does an AI product’s brand voice live?” → the system prompt’s identity + voice clauses; cite Claude’s anti-sycophancy line as a one-clause personality change.
02“Walk me through the anatomy of a prompt.” → Identity / Instructions / Examples / Context (+ output contract); each a slot with an owner and an eval.
03“Give a real example of the prompt being the product.” → v0: XML-tagged blocks govern what renders and how on-brand it feels.
04“Few-shot vs zero-shot — when, as a designer?” → few-shot to lock format/tone (3-5 curated examples); zero-shot for fast, simple, low-nuance tasks.
05“The agent is slow — what goes on screen?” → a ReAct-trace decision: expose, summarize, or hide thought/action/observation as status steps.
06“How is designing a prompt different from writing UI copy?” → it’s executable + probabilistic; you’re authoring behavior, and you must test it empirically.
07“Why structure a prompt with XML/delimiters?” → it makes slots reviewable, swappable, and parseable; the output contract keeps the renderer from breaking.
08“How do you keep a prompt from drifting on a model update?” → version it, gate changes with an eval set, and pin owners per clause — treat it like a design token.
Going deeper, expect follow-ups that separate “read a thread” from “shipped one”: “you changed one clause and the whole tone shifted — why?” (the model generalizes a voice cue across every turn; small copy, large blast radius); “how would you A/B a voice change?” (rubric-scored sample of outputs, not a vibe check — Lesson 4); and “who signs off on the system prompt?” (brand owns voice clauses, safety owns refusals, engineering owns context — name the matrix). In every case, attach the clause to the surface it controls.
A PM says “the assistant feels too fawning — it praises every message before answering.” You own the design fix. What is the highest-leverage move?
AAdd a sentence to the system prompt that bans opening with positive adjectives, then eval a sample of outputsBRedesign the chat bubble styling to feel more neutralCAdd a “tone” dropdown so users can pick a less enthusiastic mode
You inherit a 600-word system prompt that is one undifferentiated paragraph. The team can’t tell why a recent edit broke the output format. What’s the senior first move?
ARewrite it shorter so it’s easier to readBAsk engineering to add automated tests and leave the prompt as-isCDecompose it into labeled slots (identity, instructions, voice, examples, output contract) with an owner per slot
Your AI feature must return answers the UI renders as a strict card (title, 3 bullets, one CTA). Outputs keep arriving as free-form prose that breaks the card. Best prompt-design fix?
AIncrease the model’s temperature so it’s more creative with structureBSpecify an explicit output contract (schema/format + an Output Indicator) and give 3-5 examples in the exact card shapeCAdd a post-processor that tries to reformat prose into a card
In an AI-design interview you’re asked: “Show me where design taste enters a product like v0 that has no conventional UI to art-direct.” Strongest answer?
ATaste enters in the marketing site and onboarding, not the model outputBIt doesn’t — model output quality is purely an engineering/eval concernCIn the voice and output-contract clauses — e.g. an anti-“AI slop” aesthetic directive that steers the model toward distinctive, on-brand output
A reviewer asks how you’d ship a brand-voice change to the system prompt without regressing behavior for millions of users. Best process?
AVersion the prompt, run a rubric-scored eval on a held-out sample before/after, and gate the rollout on itBPush the change live and watch support tickets for a few daysCAsk the model to review its own new prompt and approve it