Provenance is a three-mode model, not a binary “AI-generated” badge. Authored / suggested / executed each generate different design primitives, the user decision and rollback cost differ for each, and the sparkle icon is becoming a cross-product standard whose semantics are not yet locked. Adobe Content Credentials, IBM Carbon’s AI Label, Cloudscape, ServiceNow Horizon — the shipped systems and what to steal.
Lesson 3 · Provenance
Authored, suggested, or executed
One badge, three different decisions
Most products treat provenance as a single category — “AI-generated” — which collapses three fundamentally different user decisions into one badge. A DALL·E image the user pastes verbatim, a drafted paragraph the user will edit, and an email the agent already sent are not the same thing, and the user’s next action differs for each: present-or-reject the whole artifact, edit-and-accept a partial one, or confirm-or-undo an action that already committed. Microsoft, Google PAIR, and ServiceNow Horizon all converge on the need for distinct visual rules per mode, because of the ELIZA effect NN/g documents — users attribute human-like certainty to anything an AI produces if the mode is unclear. This lesson is the three-mode model and the patterns that make each legible.
Define the three modes precisely, because the whole pattern set hangs off them. Authored by AI: the AI produced the output with no human edit in the loop; the decision is present or reject the whole artifact; the surface needs a persistent badge and ideally provenance metadata (C2PA / Content Credentials). AI-suggested: the AI produced a partial artifact — a draft, a ranking, a completion — the human is expected to revise; the decision is edit and accept; the surface needs a variant-aware container with a revert-to-AI affordance. AI-executed: the AI took an action (sent, filed, deleted), often outside the current view; the decision is confirm or undo; the surface needs a confirmation step plus an audit trail. This taxonomy aligns with PAIR’s Mental Models chapter: set realistic expectations and avoid unintended deception by communicating what the system actually did.
code
1THE THREE-MODE PROVENANCE MODEL23 mode user decision primary HAX typical visual4 --------- ------------------------ ------------- ------------------------5 Authored present or reject the G1, G2, G18 persistent "Generated by6 whole artifact AI" badge + C2PA metadata7 Suggested edit, accept, or reject G6, G9 variant container with a8 the partial output revert-to-AI button9 Executed confirm or undo an G10, G16, G17 confirmation step + audit10 action that committed log + per-action indicator1112 Mode matters MORE than the quality of the output. One badge for all three13 collapses three design answers into one weak gesture.
The strongest production reference for the authored mode is Adobe Firefly Content Credentials. When 100% of the pixels are AI-generated, Firefly attaches C2PA-compliant Content Credentials: issuer (Adobe), date, AI tool, app/device, the editing action (“Created” / “Other edits”), and a visual thumbnail when generated on the web. End users inspect via the Adobe Content Authenticity Inspect tool. Adobe also co-founded the C2PA coalition. The pattern is passport-style provenance: the asset carries its own portable, tamper-evident audit trail, so an auditor can answer “who made this and with what model” without an employee on a phone. Interview angle. For any enterprise generative tool, a <ProvenanceInspector> that surfaces a C2PA manifest is no longer optional — naming it signals you know the authored mode has an industry standard, not just a badge.
Suggested: the variant-aware container (IBM Carbon for AI)
The suggested mode is where design systems do their most interesting work, and IBM Carbon for AI is the cleanest shipped example. Carbon ships an AI Label component that is both an indicator of AI instances and the trigger for an explainability popover (the “first layer of explainability”). Crucially, any AI-aware component “can toggle between the AI variant and the default variant depending on the user’s interaction,” and the user can “switch back to the initially AI-generated content via a revert to AI button.” That is the mechanism: every AI state has a paired default state, and a one-tap path back to the suggestion. Carbon styles all of this with a light metaphor — brightness, glow, gradients — so AI surfaces read predictably across products.
Executed: confirmation, audit, and confidence-scaled friction
The executed mode is the highest-risk because the user may not be looking when the action commits. Microsoft’s Copilot design team built this as a system of trust signals, not one badge: a visual layer (brand colors/accents that show when Copilot produced content), an information layer (an “AI badge” to learn about the underlying tech via a click), and a behavioural layer — intentional points of friction, e.g. asking whether the user has “fact checked” the content before sharing, plus citations in every result. The design rule that ties it together is confidence-scaled friction: more friction on executed-mode and low-confidence outputs, less on suggested-mode. A destructive executed action earns a confirmation modal and an audit-log entry; a low-stakes suggestion earns a cheap dismiss.
Interview angle. “How would you signal that an AI agent did something on the user’s behalf?” The weak answer is “a toast that says done.” The strong answer separates the three signals PAIR pre-allocates to distinct slots: a capability message at entry, a confidence display at the point of prediction, a disclosure badge on the artifact, and a change notification over time — and never lets one tooltip do all four jobs. For an executed action specifically: a pre-action confirmation scaled to blast radius, a per-action indicator, and an audit trail with undo.
Citation: provenance for claims, not just artifacts
Provenance also operates at the claim level — where did this specific sentence come from? Shape of AI’s Citations pattern catalogs how shipped products do it, and the spread is the design vocabulary you should know by name: Perplexity shows multi-source inline references with favicon + title for lightweight verification; Granola uses direct quotation plus a hover “peek” at the transcript; Adobe Acrobat inlines highlights in the summary panel that click through to the source; Intercom Fin cites the policy/source behind each answer; Sana pops citations over highlighted spans; Notion AI cites nested company docs for long-tail exploration. The differentiator these share: a citation UI that scrolls the user back into the source resists the “AI as oracle” trap and turns the model from a replacement into a research assistant.
code
1CITATION DESIGN CONSIDERATIONS (Shape of AI) -- match the cue to the claim23 axis factual claim discovery / exploration4 ------------ -------------------------- --------------------------5 specificity point to the EXACT passage point to a source cluster6 placement inline cue at the sentence a panel beside long-form7 mode hover-preview + click-thru browsable reference list8 broken refs show them, never hide let users rescope/drop weak910 Rule of thumb: cite at the level of the user's query, and never let the11 citation UI run faster (more confident) than the claim it backs.
Two failure modes to design against. First, broken citations must be visible, not hidden — a dead or mismatched reference that fails silently is worse than no citation, because it manufactures false trust. Second, the citation must not outrun the claim: a confident favicon row attached to a hedged or wrong sentence is a trust mismatch. And the system should let users rescope references — filter by domain, drop a weak source, add a file — without restarting the whole query, which Shape of AI flags as a maturity marker. Interview angle. “How do you stop users from over-trusting an AI answer?” → claim-level citations sized to the query, broken refs surfaced, and a path back into the source — not a bigger disclaimer.
Cross-surface consistency: the sparkle as a shared language
The most striking development of 2024–2026 is convergent: five independent design systems adopted a near-universal glyph — the sparkle — and a near-universal label, “Generated by AI.” ServiceNow Horizon reserves an “AI sparkle” to indicate AI use, paired with an “AI tagline” that names which skill produced the content. AWS Cloudscape prescribes the verbatim label “Generated by AI,” forbids reusing the sparkle for non-generative (classical-ML) output, and forbids per-row labels inside a group (one label per group). IBM Carbon uses the light metaphor. Ripple/Watermark reserves an “ai-sparkles” icon for AI actions only. Google’s own research (Pozos & Schmidt) documents how a single sparkle became “a symbol for Google AI.” The mechanism is convergent selection pressure: users confused by inconsistent glyphs pushed each system, separately, toward a stable iconography.
The practitioner takeaway is unusually clear: pick one icon and one label for your product and commit; the cost of coining a new motif is higher than joining the consensus. But heed the counterevidence — Shape of AI’s Iconography pattern notes the sparkle’s meaning “is not yet consistent” (it appears “ambiently or as generic AI” in some products), and Cloudscape’s rule against reusing it for classical ML exists precisely because the semantics are not locked. Interview angle. A candidate who says “I’d use a sparkle” is fine; one who says “I’d reserve the sparkle exclusively for generative AI, pair it with a specific verb label, and never use it for classical-ML predictions” is demonstrating they’ve read the shipped systems.
The version-control gap — where designers can lead
One pillar is conspicuously immature: model versioning as a trust signal. Today most products surface a model name only at the picker (“Gemini 2.5 Pro,” “Claude Sonnet”), with no in-product language for surfacing a version change in flight. HAX is explicit — G18 “notify users about changes,” G14 “update and adapt cautiously” — and PAIR’s “re-boarding” recommendation says: if a feature changes significantly, trigger a small re-onboarding. The working rule the research distills: new model behind the same behaviour = silent; new model that changes behaviour = re-board. A concrete artifact to propose: a model chip (model name + last-updated date) inside every AI surface’s settings, plus a one-line “what changed” when the underlying model shifts between sessions. This is the pillar where an AI designer can most visibly lead, because no canonical pattern exists yet.
An agent drafts a reply, the user edits it heavily, but then wants the original AI version back. Which provenance pattern does the design system owe this surface?
AA persistent “AI-generated” watermark on the final textBA variant-aware container with a revert-to-AI button (Carbon for AI pattern)CA confirmation modal before the reply is editable
Your enterprise image generator outputs assets that downstream auditors must be able to trace to “who made this with what model.” What is the right primitive?
AA larger “AI” badge rendered onto the corner of every imageBA tooltip on hover explaining the image was AI-generatedCC2PA Content Credentials carried with the asset, surfaced via a ProvenanceInspector (Adobe Firefly pattern)
A team wants one tooltip to carry capability info, confidence, the AI disclosure, and change notifications. Why is that a problem?
AThese four signals belong in different slots and on different clocks — capability at entry, confidence at the prediction, disclosure on the artifact, change over timeBNothing — consolidating into one tooltip is cleaner and reduces UI clutterCTooltips are inaccessible, so the content should move to a modal instead
You are standardising an AI icon across the product. Cloudscape’s rule is most relevant — which choice reflects it?
AUse the sparkle broadly for anything “smart,” including classical-ML predictions, for visual consistencyBReserve the sparkle exclusively for generative AI output, pair it with a specific verb label, and don’t apply it to classical-ML resultsCInvent a new bespoke AI glyph so your product stands out from the sparkle consensus
Between sessions, the underlying model is swapped for one that changes the tone and structure of outputs noticeably. What does the provenance/versioning pattern call for?
ANothing visible — model versions are an implementation detail users do not needBA full re-onboarding flow every time any model is updatedCA lightweight re-board: surface a model chip (name + last-updated) and a one-line “what changed,” since behaviour shifted
Provenance is the densest part of an AI-design portfolio review because it touches trust, attribution, and risk at once. Lead with the three-mode model, then show you know the shipped references (Adobe C2PA, Carbon’s revert-to-AI, Cloudscape’s label rules, Horizon’s two axes). Each answer should name a mode and a mechanism.
01“What are the modes of AI provenance?” → authored (reject whole), suggested (edit/accept), executed (confirm/undo) — different decision and rollback cost each.
02“How do you signal AI-suggested content?” → a variant-aware container with explainability + a revert-to-AI button (Carbon), not a static badge.
03“Authored content in an enterprise tool?” → C2PA Content Credentials carried with the asset + a ProvenanceInspector; passport-style, inspectable.
04“How do you signal an executed action?” → confirmation scaled to blast radius + per-action indicator + audit log + undo; confidence-scaled friction.
05“One icon for AI?” → reserve the sparkle for generative output only, pair with a specific verb label, never reuse for classical ML (Cloudscape).
06“Disclosure wording?” → plain-language, persistent, verb-led (“Summarized with AI”), not generic “AI-powered” (Kontent.ai / Shape of AI).
07“Where does confidence live?” → at the point of prediction (PAIR: error bars / shaded ranges), not buried in a tooltip with everything else.
08“What does a strong provenance portfolio piece show?” → the same suggestion rendered across modes with distinct, consistent provenance treatments, plus the version chip.
Follow-ups probe the seams and the immature pillar. Expect “show me the executed-mode audit trail,” “what happens when the model is uncertain in suggested mode,” and the differentiator: “how would you communicate a model version change?” — where most candidates have nothing, and a model chip + behaviour-gated re-board is a standout. Note the divergence in shipped intents (Copilot’s AI badge is informational, Cloudscape’s label is attributional, Horizon’s tagline is descriptive of the skill) and say which intent you’re designing for before picking the label.
Could you take one AI suggestion and design distinct, consistent provenance for its authored, suggested, and executed modes — and propose a version-change signal?
New to itGetting thereConfident
Takeaways
Provenance is three modes — authored / suggested / executed — each with a different user decision and rollback cost, not one “AI-generated” badge.
Reserve one icon (the sparkle) for generative output only, pair it with a specific verb label, and join the consensus — don’t coin a new motif.
Model versioning is the immature pillar: silent when behaviour is unchanged, re-board (model chip + “what changed”) when it shifts — where designers can lead.
Next: consistency — keeping a design system coherent when the model output varies and is sometimes wrong.