Lesson 3 of 6 · 48 min

Designing AI copy & system messages

The system message is the most concentrated brand voice a designer ever writes — and refusals, error states, and empty states are where AI products win or lose trust. Real system-message text and the UX it produces, the refusal-as-design-pattern (substitute, not scold), voice constraints that leak from the delivery pipeline, and why Klarna had to walk back AI-first support.

Refusals and empty states are brand moments, not errors

A senior reframing: in an AI product, the moments that most shape trust are the edges — when the model declines, when retrieval comes up empty, when something fails. These are not 500 pages to apologize for; they are the highest-signal brand-voice moments you will design. Anthropic’s own refusal pattern proves it: rather than reproducing copyrighted lyrics, Claude is instructed to say “Rather than reproducing lyrics from ‘Let It Go’ (which is copyrighted material), I’d be happy to create an original ice princess poem.” That is a refusal that pivots into the user’s underlying goal — a design pattern, not a dead end.
The system message is where this voice is authored, and it is the most concentrated piece of brand copy a designer ever ships — one block read on every single turn. OpenAI’s developer-message pattern gives the four labeled regions to fill: Identity, Instructions, Examples, Context. Microsoft’s Foundry guide hands you copy-ready skeletons; the point is that these are editable design artifacts, and the wording you choose is the wording the product speaks. Interview angle. OpenAI’s “app critique” round asks for your “opinion on motion and copy in the app” — bring the same lens to a system message: where is the copy confident when the model is not?
The system message is read on every single turn — it is the most-shipped piece of copy in the entire product. Treat it with the care you’d give a hero headline, because it is one, repeated thousands of times a day.
Prompt Workshop — designing system messages with ClaudeZack Witten / AI Engineer

What goes in the system message — two real skeletons

Microsoft publishes two system messages a designer can copy structurally. They are deliberately plain, which is the lesson: clear scope and an explicit recovery branch beat clever prose. Read the “I don’t know” line as a designed refusal-cum-recovery, and the “recommend users go to the IRS website” line as a UI handoff — the assistant telling the user where to go next, exactly like a good human support rep.
code
1TWO PRODUCTION SYSTEM MESSAGES (structurally copyable)23  TAX ASSISTANT4    Assistant is an intelligent chatbot to help users with tax questions.5    Instructions:6    - Only answer questions related to taxes.7    - If unsure, you can say "I don't know" or "I'm not sure," and8      recommend users go to the IRS website for more information.910  SENTIMENT ANALYZER11    You're an assistant that analyzes sentiment from speech data.12    Users paste a string; you respond with an assessment of the speaker.13    Rate on a scale of 1-10 (10 highest). Explain why this rating was given.1415  Notice: scope ("only taxes"), a refusal branch ("I don't know"), a16  handoff ("go to the IRS website"), and an output contract (1-10 + reason).
That “I don’t know / go to the IRS website” branch is a design decision with real consequences. A model with no permission to say “I don’t know” will confabulate rather than admit a gap — so the absence of a refusal branch is itself a UX choice (the wrong one) that ships hallucinations. Designing the explicit out, and the handoff that follows it, is how you turn a limitation into a helpful moment. Interview angle. Anthropic’s loop asks you to “design a feature that helps users understand AI limitations” — a well-designed “I’m not sure, here’s where to look” response is a direct answer.
Notice the four moves packed into that tiny tax-assistant message, because they generalize to any system message you write: a scope clause (only taxes) that defines the feature’s lane; a refusal branch (“I don’t know”) that licenses honesty; a handoff (the IRS website) that tells the user where to go next; and an output contract (in the sentiment example, “1-10 + reason”) that shapes what the renderer receives. A system message that’s missing any of these has a corresponding hole in the UX — no scope means off-topic drift, no refusal branch means confabulation, no handoff means a dead end, no output contract means a broken render.

Refusal patterns: substitute, don’t scold

Anthropic’s system prompt encodes two refusal archetypes worth stealing. The first is the substantive refusal with no moralizing: Claude “does not provide information that could be used to make chemical or biological or nuclear weapons… even if the person seems to have a good reason for asking for it” — strict scope, no lecture. The second is the rule that makes refusals not feel preachy: “If Claude cannot or will not help the human with something, it does not say why or what it could lead to, since this comes across as preachy and annoying.” The design principle: decline cleanly, then pivot to what you can do.
code
1REFUSAL UX -- the same decline, two very different experiences23  SCOLD (rejected pattern)4    "I can't help with that. Generating that content would be5     irresponsible and potentially harmful, and I won't do it."6    -> reads as preachy, annoying; user feels lectured78  SUBSTITUTE (designed pattern)9    "I can't reproduce the copyrighted lyrics -- but I'd be happy to10     write you an original ice-princess poem in that style."11    -> declines cleanly, then advances the user's underlying goal1213  Rule: a refusal is a redirect. Name the boundary once, then offer the path.
There is a deliberate override here that designers must understand, because it’s a judgment call, not a law. The “don’t explain why” rule is right for a general assistant — but a designer working in a regulated or wellness vertical may want the opposite: a transparent explanation of why the assistant is declining (“I’m not able to give medical dosing advice; here’s why and who can”). That is a defensible override, chosen for the context, not a violation. Knowing when to invert a default like this — and being able to defend it — is exactly the trade-off reasoning interview rubrics reward.

Voice constraints leak from the whole delivery pipeline

Voice is not only “tone” — it is shaped by the entire delivery pipeline, and the constraints leak into the prompt. Anthropic’s best-practices include the line “Your response will be read aloud by a text-to-speech engine, so never use ellipses since the text-to-speech engine will not know how to pronounce them.” and, for parser-fed surfaces, “Format your response in plain text only. Do not use LaTeX… write all math using standard text characters.” A designer must read the downstream — TTS, screen readers, the renderer — as constraints on what counts as “expressive.” The same rule that’s right for a voice agent is wrong for an artifact-rendering app.
A surface-specific example from Claude’s prompt: “Claude responds in sentences or paragraphs and should not use lists in chit chat.” That rule makes chat feel human — but a designer building a structured, scannable artifact UI would invert it and demand lists. The voice guideline is the UX, and it must be tuned per surface: the conversational surface and the document surface want opposite formatting. This is why a single global “voice” spec fails; you write voice per surface, against its delivery pipeline.
code
1VOICE IS PER-SURFACE -- the same brand, different rules by pipeline23  Surface              Pipeline constraint        Formatting rule4  ------------------   ------------------------   --------------------------5  Voice agent (TTS)    spoken aloud               no ellipses, no markdown,6                                                  spell out symbols7  Chat                 read on screen, casual     sentences/paragraphs,8                                                  no lists in chit-chat9  Artifact / doc       scanned, structured        lists + headings ENCOURAGED10  Parser-fed field     downstream code reads it   plain text, strict schema,11                                                  no LaTeX / markup1213  One brand voice, four rule sets -- the surface and its pipeline decide.

Empty & error states: the copy you write for when it breaks

Because a probabilistic feature fails differently from a deterministic one, it needs more edge-state copy, not less — and this is the copy juniors forget. Four states recur and each needs designed words: the empty state (nothing to draft / no results yet — say what the feature does and how to start, not a blank box), the low-confidence state (the model is unsure — hedge honestly and offer a next step rather than guessing), the refusal state (out of scope — decline and redirect), and the error/repair state (a tool failed — retry, partial result, or escalate). The AI UX Playground’s rule is the discipline to adopt: define the repair path in advance for every feature — verify, retry with deltas, or escalate to a human.
code
1THE EDGE-STATE COPY CHECKLIST (write these, not just the happy path)23  State            What the copy must do                Anti-pattern4  --------------   ----------------------------------   ----------------------5  Empty            say what it does + how to start      a blank box / spinner6  Low-confidence   hedge honestly, offer a next step    confident guess7  Refusal          decline in one line, then redirect   scold + dead end8  Error / repair   retry / partial / escalate -- named  raw stack trace910  Senior tell: a candidate who designed all four shipped a SYSTEM, not a screen.
Anyone can write the happy-path reply. The AI designer’s craft shows in the empty state, the “I’m not sure,” the clean refusal, and the repair — because that is the copy users actually hit when the model is uncertain or wrong.

Case study: Klarna’s AI-first support retreat

Klarna launched AI-first customer service hard, then in 2025 had to rebalance toward a human-hybrid model — a widely-cited cautionary tale. The root cause practitioners point to: insufficient evaluation coverage of the edge cases the original rubrics never anticipated, so the refusal/recovery paths underperformed exactly where real customers needed them. For designers the lesson is concrete and uncomfortable: the retry mechanism — human takeover, fallback URL, escalation phrase — is a design surface, and Klarna’s experience degraded because that surface was underspecified. The fix is UX Content Collective’s “design responsively” discipline: look at a sample of real outputs, find the problematic areas, and add criteria to catch them.
The Glean content-review data quantifies the upside of getting the human-handoff right: a structured workflow with role-based review gates delivers 40-60% faster approval cycles and cuts revision rounds from 5-7 down to 2-3. And the stakes are real — Glean cites that 42% of companies abandoned most of their gen-AI initiatives in the past year (up from 17% the year before), and 54% of CMOs flag “brand voice drift” as a top governance risk. The through-line: the edges (refusal, handoff, voice consistency) are not polish — they decide whether the product survives contact with real users. Interview angle. Klarna is the case study to cite when asked “tell me about an AI UX that failed and why.”
Klarna’s walk-back from AI-first support wasn’t a model failure — it was a design failure at the edges. The refusal, the handoff, and the recovery were underspecified, and that is exactly where real customers showed up. Design the edges first.

Interview prep

Copy/voice rounds — and the portfolio “critique this AI copy / system message” archetype — test whether you treat refusals and error/empty states as designed brand moments, and whether you can defend voice as a per-surface trade-off. The strong critique uses the same lens you’d use on a system prompt: where does the copy incentivize the wrong behavior, where is it confident when the model isn’t, where does it scold instead of redirect. Lead with the user’s feeling, then the clause that produces it.
  1. 01“How should an AI feature refuse?” → name the boundary once, no moralizing, then substitute the user’s underlying goal (Claude’s “Let It Go” → original poem).
  2. 02“Why design the ‘I don’t know’ branch?” → without permission to admit a gap, the model confabulates; the explicit out + handoff prevents hallucinations.
  3. 03“Is warm, encouraging voice always right?” → no; it’s a per-product/per-surface call — anti-sycophancy for honesty-first, warmth for a companion.
  4. 04“Why would the same product format chat and an artifact differently?” → chat wants sentences (human feel); a scannable artifact wants lists — voice is per surface.
  5. 05“What constrains voice beyond tone?” → the delivery pipeline: TTS (no ellipses), parsers (plain text/no LaTeX), screen readers — design against the downstream.
  6. 06“Tell me about an AI UX that failed.” → Klarna’s AI-first support retreat: underspecified refusal/recovery + thin eval coverage of edge cases → human-hybrid walk-back.
  7. 07“Critique this system message.” → check scope, refusal branch, handoff, output contract, and where copy is over-confident relative to the model.
  8. 08“Where do you put the human handoff?” → as a designed escalation surface (takeover / fallback URL / escalation phrase); role-based review gates cut cycles 40-60%.
Going deeper, expect the portfolio follow-ups: “walk me through the empty state for this feature” (most candidates only designed the happy path — the empty/low-confidence state is the tell); “how would you test that your refusals don’t feel preachy?” (sample real outputs, score against a voice rubric, iterate — Lesson 4); and “show the seam between your copy and the model’s output” (which words are your system message vs the model’s free generation, and how you constrained the latter). A strong portfolio piece shows the refusal and error states, not just the polished demo.
articleKlarna customer service: from AI-first to human-hybridPromptLayervideoDesigning with Claude: from prompt to productionAnthropic / ClaudedocsPrompting best practices (voice constraints, refusals)AnthropicarticleHow to implement an AI content-review workflow (review gates, numbers)Glean

Checkpoint

Your assistant must decline requests outside its scope (e.g. legal advice in a fitness app). A reviewer says the current refusals “feel like a lecture.” Best redesign?

AAdd more detail explaining exactly why each request is declinedBDecline in one clause without moralizing, then redirect to what the app can help withCRemove the refusals so it never says no
Sign up free to answer and see why

Checkpoint

The same model powers both a voice assistant (TTS) and a docs-style artifact view. One voice spec is causing odd output in each. What’s the senior take?

APick the format that looks best in the artifact view and use it everywhere for consistencyBLower the temperature so the formatting stabilizesCAuthor voice per surface against its pipeline — no ellipses/no lists for TTS, lists and structure for the artifact
Sign up free to answer and see why

Checkpoint

Leadership wants to launch an AI-only support bot with no human escalation “to cut costs.” Drawing on Klarna, what do you advocate for as the designer?

ADesign the refusal/recovery paths and a human-handoff surface up front, with eval coverage of edge cases before scalingBShip AI-only and add human escalation later if complaints come inCArgue AI should never handle support
Sign up free to answer and see why

Checkpoint

A portfolio reviewer asks to see how your AI feature behaves “when things go wrong.” Your case study currently shows only successful answers. What does this reveal — and the fix?

ANothing — a clean happy path is the strongest way to show the designBIt reveals the piece designed a screen, not a system; add the refusal, empty/low-confidence, and recovery states with the decisions behind themCAdd more visual polish to the successful screens
Sign up free to answer and see why

Checkpoint

You’re critiquing a system message in an interview. Which observation best demonstrates senior judgment?

A“The copy is grammatically clean and the tone is friendly.”B“The instructions read nicely out loud.”C“There’s no refusal branch and no handoff, so it’ll confabulate on gaps; and the copy promises certainty the model can’t guarantee.”
Sign up free to answer and see why

Could you write a refusal that redirects, design the empty/error states, and defend a voice choice as a per-surface trade-off?

New to itGetting thereConfident

Takeaways

  • The system message is the most concentrated brand voice you ship — read on every turn; design it as Identity / Instructions / Examples / Context.
  • A refusal is a redirect: name the boundary once, skip the lecture, substitute the user’s real goal (explain “why” only in regulated/wellness contexts).
  • Design the “I don’t know” branch + handoff, or the model confabulates on gaps.
  • Voice is per-surface and shaped by the pipeline: TTS, parsers, and screen readers constrain what counts as expressive; chat ≠ artifact.
  • Klarna’s AI-first retreat = underspecified recovery + thin edge-case evals; the human handoff is a design surface.
  • Strong portfolios show the refusal, empty, and error states — happy-path-only reads as “designed a screen, not a system.”

Next: how to define and measure success for a probabilistic feature — rubrics, the metrics that matter, and LLM-as-judge.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.