AI prototyping is three layers, not one tool — Figma for shape, Framer/Magic Patterns for on-brand demo, code+LLM (v0/Bolt/Lovable) for behavioral truth. The Wizard-of-Oz move for validating an interaction before a model exists, when to really wire the model, the streaming + skeleton patterns that make probabilistic UIs feel responsive, and how Notion’s design team actually prototypes.
Three layers, three different truths
AI prototyping is not a single medium, and the senior mistake is collapsing it into one. There are three layers, each answering a different question about a model-mediated interaction. Figma is the discovery layer — does the team share one shape? Framer / Magic Patterns is the polished-demo layer — does it look and ship on-brand, with real CMS and a clickable URL? Code + LLM (v0, Bolt.new, Lovable) is the truth layer — does it actually run and feel right against a real model? The move is to start upstream for alignment, drop to Framer for visual stakes, then drop again to code+LLM when latency, hallucinations, or edge cases start to matter to the design decision.
Why bother dropping all the way to code? Because, as Notion product designer Brian Lovin puts it, design tools “cannot capture how it will feel like to use that thing.” A static Figma of an AI feature hides exactly the properties that define it — the latency, the variance, the wrong answers. The fidelity ladder is a way to spend the cheapest effort that still answers the question on the table. Interview angle. OpenAI’s designer loop expects “latency + scalability awareness” as a craft pillar and follows up with “did that design lead to latency issues?” — you can only answer that if you’ve felt the feature run, not just drawn it.
1THE PROTOTYPING STACK -- pick the layer by the question23 Layer Best question answered Tools Graduate when4 ------------------ --------------------------- ------------------ ------------------5 Figma "do we share one shape?" Make, Make-interact team aligned6 Framer / Magic Pat. "does it look + ship?" Framer AI, Wireframr stakeholder demo7 Code + LLM "does it run + feel right?" v0, Bolt, Lovable latency/edge cases89 Cost free/low ------------- mid ----------------- high (compute + time)10 Risk slow redo ----------- design debt ---------- LIVE hallucinations11 Horizon hours --------------- days ----------------- weeks / next 10 iters
Each layer has real mechanics worth knowing. Figma Make lets you bring a frame from a Figma library so AI output stays on-brand, and its Make-interactions wire promptable flows between frames in a few clicks (consistent layer naming improves the output, and you’re expected to edit or remove the AI’s baseline interactions). Framer AI generates layouts/sections/CMS and even reviews the site for contrast, typos, missing alt text, and SEO before launch; its 2025 Wireframer takes the opposite bet — structure, not style — to remove blank-canvas friction while you keep aesthetic control. v0 defaults to shadcn/ui specifically because those primitives are “designed with the right primitives and patterns to help models generate real, brand-aligned interfaces” — and via MCP, a component registry becomes “the mechanism in which we humans and machines contextualize and use a design system.”
The Adobe design team maps the same trip and reports the speed: with Cursor, v0, Replit, Bolt, and Figma Make in the kit, a designer can go from idea to interface in a single week. Their prompting craft note is the transferable one — feed the model visual context (screenshots of existing designs, inspirational images, links to the design files) so its output starts on-brand rather than as generic Tailwind. That’s the difference between code+LLM as a slot machine and code+LLM as a design instrument: the designer supplies the taste in the prompt, and the tool executes it at speed.
Start upstream in Figma for alignment, drop to Framer for visual stakes, drop again to code+LLM for behavioral truth — and tell the team which layer each prototype is at. A Figma flow that looks done is answering “do we share a shape,” not “will this feel good to use.”
Case study: Notion’s shared “prototype playground”
Brian Lovin built Notion’s design team a shared Next.js app that doubles as the prototyping surface — “organized by designer name and provides shared components for Notion-style UI elements,” so everyone iterates against the same baseline instead of forking one-off prototypes. Reusable skill files capture recurring solutions (e.g. a skill that searches for icons and synonyms to prevent the AI hallucinating an icon that doesn’t exist), and slash commands like /figma chain real work — MCP install, design extraction, code implementation, verification. Lovin’s rule crystallizes the philosophy: “when Claude asks you to do something, teach it to do that thing itself,” so manual intervention shrinks over time.
The transferable lesson is about shared infrastructure: prototyping speed compounds when one designer’s clever prompt becomes every designer’s clever prompt. Framer encodes the same instinct with Branches — because AI edits are non-deterministic, you isolate each prompt iteration so a bad generation can’t corrupt your working version. A hidden tension to manage: branches are an under-appreciated source of design drift (variants diverge, and the prompt/model that produced a merged result gets hard to recall), so the senior practice is to stamp each branch with the prompt and model version that produced it.
code
1NOTION'S SHARED PROTOTYPE PLAYGROUND (the pattern to copy)23 one Next.js app, organized by designer name4 + shared components for Notion-style UI elements5 + skill files = reusable solutions (e.g. icon-search to stop6 the AI hallucinating an icon that doesn't exist)7 + slash commands chain real work:8 /figma -> MCP install -> extract design -> implement -> verify910 Lovin's rule: "when Claude asks you to do something,11 teach it to do that thing itself."12 -> manual intervention shrinks every time; the team's prompt library grows.
Faking vs really wiring an LLM — graded fidelity
The binary “fake or wire?” is the wrong framing; the right one is graded fidelity — the cheapest level that still answers the question. Fake-the-LLM (Wizard of Oz) belongs at shape-validation. NN/g’s five-step setup is the canonical playbook: build the prototype in the fastest medium, pick a response method (closed = wizard chooses from a set list, open = wizard types anything, hybrid = both), write a protocol with roles, brief and rehearse the wizard, and pilot before live sessions. The method predates chatbots — Zappos used manual order fulfillment as its WOZ MVP to test the value prop before automating anything. A useful tip: have the wizard pose as a “notetaker” so the user keeps believing in the machine during the session.
code
1MATCH FIDELITY TO THE QUESTION ON THE TABLE23 Question you're answering Cheapest fidelity Move up when4 ------------------------------------ ------------------- ----------------------5 Do users understand the affordance? paper + closed WOZ critics stop noting it6 Does the interaction feel natural? open WOZ (human on wizard latency >7 Slack/Discord) model latency8 Will it ship w/o embarrassing halluc.? real LLM, streaming 1 in 5 testers hits a9 in Figma or v0 bad output10 Will it survive production traffic? real LLM + evals eval regression > 1%11 + HITL
Wire-the-real-LLM belongs at the edge-case stage, because that’s where Figma stops being honest. Practitioner Oliver Engel’s Pitch.ai public-speaking assistant is the textbook WOZ: he scoped with a Kano survey, ran moderated sessions where on-screen cue cards pretended to prompt the assistant, and built a deliberately minimal rig — local server with live-reloading, an HTML/CSS/JS interface, Processing scripts as digital props (an OpenCV face tracker, an oscilloscope voice waveform), and a “calculating results” loader that hid his manual updates. Every choice was a designer’s choice, so every failure became a design signal, not a model failure. The WOZ beauty: “the machine is believable” without any model at all — the only prerequisite is honest framing with participants afterward.
Wizard of Oz turns every failure into a design signal instead of a model excuse — because every response was a human’s choice. That’s its superpower: you learn whether the interaction works before you’ve spent a dollar on inference.
The two patterns that make a probabilistic UI feel responsive
Once you wire a real model, two patterns are essentially required, and they solve different latency budgets (so you usually want both). Streaming renders tokens incrementally so the user sees output begin almost immediately — it reduces perceived wait for long generations. Skeleton prediction shows the layout/structure of the answer before any tokens arrive — it sets expectations when total time is highly variable. Add provisional states (“draft,” “suggested,” “pending review”) so the user knows nothing is final yet. The Adobe design team adds a prompting craft note: include screenshots of designs, inspirational images, and links to design files in the prompt so the model has visual context — a step pure WOZ can’t replicate.
01Composer — slash commands, multimodal intent, follow-up chips: where the user forms the request.
02Conversation — thread branching, scroll-to-bottom, memory management: how turns accumulate.
03Response — streaming, skeleton prediction, status steps, response refinement: how output arrives.
04Context governance — citation tooltips, confidence indicators: how the user verifies and trusts.
05Trust calibration — failure disclosure, human-in-the-loop checkpoints, repair contracts: how it recovers.
The AI UX Playground’s Chat UX framework is the coverage checklist to iterate against — 30 patterns across Composer (slash commands, multimodal intent, follow-up chips), Conversation (thread branching, memory), Response (streaming, skeleton prediction, status steps), Context Governance (citation tooltips, confidence indicators), and Trust Calibration (failure disclosure, HITL checkpoints, repair contracts). Its standing orders are senior-level: “start with composer and turn-taking before tuning visual polish,” and “for every chat feature, define the repair path in advance — verify, retry with deltas, or escalate to a human.” One more rule for testing AI prototypes from NN/g: do not use AI to moderate usability tests — the moderator stays human; AI is for transcription, tagging, synthesis, and recruitment.
Generative UI: pick the shape before you prototype
A frontier decision that quietly determines what your prototype even is: when the model can generate interface, not just text, you must pick the generative-UI shape first. CopilotKit names three. Static — the model fills fixed, pre-built components (predictable, on-brand; best for high-traffic surfaces). Open-ended — the model emits raw HTML/iframes (maximum flexibility; fastest for prototyping, riskiest in production). Declarative — the model returns a structured spec the app renders (cross-framework, controlled). The same prompt can produce three wildly different products, so a senior designer chooses the shape before prototyping, not after — otherwise you’re evaluating an artifact whose category you never decided.
code
1GENERATIVE-UI SHAPES -- decide BEFORE you prototype23 Shape Model produces Best for Risk4 ------------ ----------------------- ------------------- ------------------5 Static fills fixed components high-traffic surfaces least flexible6 Open-ended raw HTML / iframes fast prototyping unpredictable in prod7 Declarative a structured spec to cross-framework needs a render layer8 render910 Same prompt, three different products -- the shape is a product decision.
Interview prep
Prototyping shows up two ways in AI-design hiring: the live whiteboard (“prototype an AI feature that helps you discover X”) and the portfolio walkthrough, where the strongest signal is now how you prototyped, not just what. Randy Hunt’s “strongest” portfolio label is Flexible stance — “sometimes I handcraft, sometimes I vibe-code; here’s why, when, and the results.” Zapier/Metaview’s workflow probe is the one to rehearse for every tool you mention: “what you requested, what the AI returned, what was incorrect, and what you changed.” Lead with the question you were answering, then the fidelity you chose to answer it.
01“How do you prototype an AI feature?” → name the three layers (Figma shape / Framer demo / code+LLM truth) and graduate down only when the cheaper layer can’t decide it.
02“Fake or wire the model?” → graded fidelity: WOZ (closed→open) for shape, real streaming model once latency/hallucinations/edge cases are the question.
03“Walk me through a Wizard-of-Oz study.” → NN/g’s 5 steps; closed vs open vs hybrid response; wizard posing as notetaker; pilot before live (Pitch.ai as the example).
04“What makes a probabilistic UI feel responsive?” → streaming (perceived wait) + skeleton prediction (variable total time) + provisional states; they solve different budgets.
05“Walk me through what the AI generated and what you changed.” → the workflow probe: what you asked, what came back, what was wrong, what you kept — per tool.
06“Why shadcn/ui or a component registry with v0?” → on-brand generation; via MCP the design system becomes how humans and models share components.
07“How do you keep non-deterministic AI edits from corrupting your work?” → branch per iteration (Framer/Notion) and stamp each branch with its prompt + model version.
08“Can you use AI to run the usability test?” → no — keep a human moderator; AI is for transcription, tagging, synthesis, recruitment (NN/g).
Going deeper, expect: “you have three AI tools — how do you choose?” (Hunt’s “tool curiosity” pattern — you tried several and recorded the trade-offs, not just used the most popular); “this prototype is janky — is that okay?” (depends on the layer and question; “loose and unaligned” vibe-code with no taste is the explicitly weak portfolio tag, but a scrappy WOZ rig is fine because it’s answering shape); and “how would you test this without real users yet?” (define each scenario as a deterministic prompt the model replays and capture time-to-first-token, hallucination rate, refusal rate, edit-distance from a gold answer). Show the seam between your craft and the model’s output.
You need to learn whether users will understand a novel “ask-then-approve” interaction for an AI agent. No model is built yet. Cheapest prototype that answers it?
AA Wizard-of-Oz prototype with a human supplying responses (closed first), tested in moderated sessionsBWait for engineering to wire the real model, then testCA high-fidelity Framer site with polished visuals
A wired prototype generates a long answer; users stare at a spinner and some abandon before it finishes. Which combination best fixes the perceived experience?
ASwap to a bigger model so the answer is betterBStream tokens as they generate AND show a skeleton of the answer’s structure firstCAdd a longer, more elaborate loading animation
In a portfolio review you’re asked to “walk through what the AI generated and what you did with it” for a v0 prototype. What demonstrates senior practice?
A“I prompted it and it gave me the final screens, so I shipped those.”B“I didn’t use AI for this one to keep it clean.”C“Here’s what I asked, what came back, what was wrong (off-brand spacing, a hallucinated icon), and what I kept vs rebuilt by hand — and why.”
Your team iterates an AI feature with Framer/AI edits and keeps losing good versions to non-deterministic regenerations. Best practice?
ABranch each prompt iteration and stamp every branch with the prompt + model version that produced itBStop using AI edits and hand-build everythingCKeep one file and undo when a generation is bad
A reviewer pushes: “this prototype looks janky.” Which response reflects the senior framing of fidelity?
A“All AI prototypes look rough — that’s expected.”B“It’s a Wizard-of-Oz rig answering whether the interaction reads — scrappy is correct here; the on-brand build comes once shape is validated.”C“I’ll add polish to every screen before the next review.”
Could you pick the right prototyping layer + fidelity for a given question, run a WOZ study, and add streaming/skeleton to a wired prototype?
New to itGetting thereConfident
Takeaways
Prototype in three layers: Figma (shape) → Framer/Magic Patterns (on-brand demo) → code+LLM (behavioral truth); name which layer answers the question.
Use graded fidelity: Wizard of Oz (closed→open) validates interaction shape before any model; wire the real model when latency/hallucinations/edge cases are the question.
Streaming + skeleton prediction solve different latency budgets — use both; add provisional states so nothing reads as final.
Notion’s playground = shared infra turns one designer’s prompt into everyone’s; branch per AI edit and stamp it with prompt + model version.
Keep a human moderator for usability tests; AI is for transcription, tagging, synthesis, recruitment.
In the portfolio, show the flexible stance and the seam — handcraft vs vibe-code with rationale; janky-final is the weak tag.
Next: the capstone — design an end-to-end AI interaction with a prompt strategy and a human-in-the-loop review/approve flow.