Lesson 2 of 6 · 47 min

Product KPIs + model-quality metrics

The metric stack — product/business outcomes on top, model-quality in the middle, system reliability at the base — plus the divergence alarm that catches a model silently degrading while your North Star climbs. With named-company stacks (GitHub Copilot/Accenture, Notion) and the metrics that mislead.

Why one number lies

The defining AI-PM metrics failure is conflating “the model is good” with “the product is good,” and collapsing both into a single “AI quality” score. Galileo is explicit that generic single-number metrics “create an illusion of confidence that is unjustified.” The thing that makes AI metrics different from a normal feature’s engagement-and-retention dashboard: model quality and product outcome can move in opposite directions, silently. Your North Star can climb while the model quietly rots on a slice of users — and a flat KPI sheet will never show it. This lesson builds the layered stack that does.
The mental model every primary source converges on is a stack, not a flat list. Top layer: product / business outcomes — the only layer that measures value (task completion, adoption, retention, deflection, LTV). Middle: model-quality / behavioural metrics — hallucination rate, faithfulness, factual correctness, recall, language adherence. Base: system / reliability — p95 latency, cost per task, escalation-to-human rate, error budget. The PM’s job is to pick one dominant metric per layer and refuse to average them into a single score. 5D Vision frames the same four layers (model, system, product experience, business); Mixpanel’s categories (adoption & engagement, model monitoring & quality, business impact) line up exactly.
code
1THE AI METRIC STACK (one dominant metric per layer)23  LAYER              EXAMPLE METRICS                       CATCHES4  ---------------    ----------------------------------    --------------------------5  Business / NSM     task completion, deflection, LTV      is the product used / valued6  Product engage.    DAU/WAU, retention, time-to-task      did behaviour move7  Model quality      hallucination rate, faithfulness,     is the MODEL improving or8                     recall, language adherence            silently degrading9  System / reliab.   p95 latency, $/task, escalation rate  is the runtime trustworthy1011  + DIVERGENCE ALARM across layers: fire when NSM and model-quality drift apart.12  Sources: 5D Vision (4 layers); Mixpanel (30+ AI metrics); KORE1 (divergence).
Why a stack and not a number? Two mechanisms. First, each layer catches a different failure class: system metrics catch a slow or expensive runtime, model-quality metrics catch a regression on a slice, product metrics catch a feature nobody adopts. A single average hides all three behind each other. Second, and this is the AI-specific part, the layers can diverge — a UX change can push adoption up while a model swap drops faithfulness, and the blended “score” looks flat. KORE1’s rubric for the canonical interview question names the strong-answer decomposition outright: “a layered set: a product outcome, a model-quality metric (e.g. hallucination rate or faithfulness), and an alarm for when the two diverge.” The divergence alarm is the decisive element.
Becoming an AI PM — metrics, evals, and what AI PMs actually doLenny’s Podcast (Aman Khan, Arize)

The divergence alarm: catching a silent model regression

This is the metric a non-AI PM never has to build, and the one interviewers probe hardest. KORE1’s exact prompt: “You ship an assistant and the north-star metric goes up. How would you know the model is quietly getting worse anyway?” The failure it’s testing is real: model quality can decay on a sub-population (a new language, a new doc type, a provider model update) while your top-line — say weekly active users — keeps climbing on momentum, marketing, or the 90% of users who are fine. A flat dashboard reports green right up until the churn shows up weeks later.
The senior fix is a held-out quality gate that runs independently of the product metric. Concretely: a “Golden-100” eval on a stratified, held-out set, run weekly, that blocks model promotion if faithfulness drops more than ~5% — regardless of what DAU is doing. Pair it with online detectors: an escalation-to-human / “ask for human” rate (Notion tracks this deflection signal directly), a human-correction rate, and per-slice quality so a regression concentrated in one cohort surfaces in 24 hours, not 24 days. Interview angle. If you can’t name the divergence detector — “weekly stratified eval gating promotion on a faithfulness delta, plus per-slice escalation-rate alerts” — the grader concludes you haven’t internalised that AI products fail in a way traditional ones don’t.

GitHub Copilot at Accenture: a four-layer stack in production

GitHub’s published enterprise study with Accenture (May 2024) is the cleanest public example of a layered stack actually shipped, with real numbers. Read it as four metrics, one per layer: onboarding success (a leading fit indicator) — “96% success rate among initial users” who installed the extension and began accepting suggestions the same day; adoption — “over 80% of Accenture participants successfully adopted” the tool; per-interaction acceptance (the behaviour-as-quality proxy) — developers “accepted around 30% of suggestions”; and perceived productivity — 43% rated it “extremely easy” to use. That spread is a textbook stack: deterministic system signal, online usage behaviour, and human perception, side by side.
The PM lesson from Copilot isn’t the specific numbers, it’s the discipline of behaviour over survey: the 30% suggestion-acceptance rate is a per-interaction quality signal you can read continuously, while the 43% “extremely easy” survey is a lagging, sentiment-y backstop. The rule a senior PM applies: prefer behaviour over survey as soon as you have signal, and never let a model-accuracy number substitute for a retention number. Notion’s parallel: 80% of its AI team’s work is “evaluating feedback and traces in Braintrust” rather than ad-hoc vibe checks — its tracked metrics are hallucination detection, factual correctness, tool-usage accuracy, and recall, each laddering up from a defined “reward function.”
code
1GITHUB COPILOT @ ACCENTURE — one metric per layer (May 2024)23  LAYER            METRIC                         NUMBER     ROLE4  --------------   ----------------------------   --------   --------------------------5  Fit (leading)    same-day onboarding success    96%        early signal of fit6  Adoption         participants who adopted        >80%       did behaviour move7  Model quality    suggestion ACCEPTANCE rate     ~30%       behaviour-as-quality proxy8  Perception       rated tool "extremely easy"     43%        lagging sentiment backstop910  Takeaway: pick one per layer; prefer behaviour (acceptance) over survey11  the moment you have signal; never let accuracy stand in for retention.
A second named stack, outside developer tools, shows the same layered shape on a support product — and a cautionary edge. In its first month, Klarna’s OpenAI-powered assistant handled 2.3 million conversations — two-thirds of its customer-service chats — doing the equivalent work of 700 full-time agents, while landing on par with human agents on CSAT (the product/quality layer), cutting average resolution time from 11 minutes to under 2 (the system layer), and driving a 25% drop in repeat inquiries (a behavioural quality signal — fewer re-asks means the answer actually resolved the errand). Read it as the support-chatbot archetype instantiated: deflection as the NSM, repeat-inquiry rate and CSAT as model-quality, resolution time as system. The cautionary edge a senior PM names too: Klarna’s deflection numbers were widely reported up front, but durable quality is the harder claim — which is exactly why the divergence alarm and per-slice quality gate matter on a product like this, not just a headline deflection rate.

Metrics that mislead — and the product-type stacks that don’t

Some metrics actively lie for AI products. Raw accuracy misleads when errors are stratified (85% overall can be 100% on the easy 80% and 25% on the hard 20% — a worse product than 90% flat). Average latency misleads because the tail is what users feel — quote p95, not the mean. Thumbs-up rate misleads because of a well-documented bias against hedged answers: the peer-reviewed EMBER study (Lee et al., NAACL 2025) shows evaluators carry a negative bias toward epistemic markers — outputs that hedge with “I think” or “possibly” get marked down relative to confident ones of equal correctness — so a confidently-worded hallucination collects thumbs-up while an honest, calibrated answer is penalised. And engagement misleads as a quality proxy: more messages can mean the assistant is failing and users are re-asking. The senior move is to name, for each metric, the failure it can hide.
The antidote is a pre-stacked metric set per product archetype, so in a case round you’re not inventing from scratch. RAG assistant: task-completion (NSM) · hallucination rate + faithfulness (quality) · p95 latency (system). Recommendation: conversion of recommended item · precision@k / recall@k · cost per recommendation. Action-taking agent: successful end-to-end task rate · hallucination + action-reversibility · spend-cap adherence + kill-switch invocation rate. Support chatbot: tickets deflected · human-correction + escalation rate · p95 latency + $/resolution. Interview angle. Asked to “define metrics for [product],” open with the four-layer stack, name the dominant metric per layer for that archetype, then add the divergence alarm — and refuse to name a metric until you’ve scoped the task.
One more senior nuance: bind the model-quality metric to a downstream cost, because the same hallucination rate means different things in different products. A 3% hallucination rate is a rounding error in a brainstorming tool and a lawsuit in a benefits chatbot (see Air Canada, L5). So the metric stack is incomplete without a one-line cost-of-error statement per layer — it’s what turns “faithfulness dropped 4 points” into “escalate now” or “monitor.” This is the bridge from metrics to launch gates: a threshold is only meaningful once you’ve priced what crossing it costs.
articleResearch: Quantifying GitHub Copilot’s impact in the enterprise with AccentureGitHubarticleThe Four Types of Metrics for AI Product Managers5D VisionpaperEMBER — Are LLM-Judges Robust to Expressions of Uncertainty? (the negative bias against hedged answers)Lee et al., NAACL 2025articleKlarna’s AI assistant does the work of 700 full-time agents (first-month support metrics)Klarna

Checkpoint

Your assistant’s weekly active users are up 12% quarter-over-quarter and leadership wants to declare victory. What does a senior AI PM check first?

AWhether model quality diverged from the rising NSM — run a held-out, per-slice quality gate (e.g. faithfulness) independent of WAUBWhether the 12% is statistically significantCWhether marketing spend explains the lift
Sign up free to answer and see why

Checkpoint

A teammate proposes a single blended “AI Quality Score” (one 0–100 number) as the feature’s headline metric. Best response?

AApprove it — one number is easiest for executives to trackBApprove it but only if computed by an LLM judgeCReplace it with a stack — one dominant metric per layer (business / model-quality / system) plus a divergence alarm
Sign up free to answer and see why

Checkpoint

Your support bot’s thumbs-up rate is high, but escalations to human are also rising. What’s the most likely explanation?

AThe thumbs-up rate is wrong and should be ignoredBConfident-but-wrong answers collect thumbs-up while genuinely failing — cross-check against escalation/correction rate and a held-out faithfulness evalCUsers are happy but lonely, so they escalate to chat with a human
Sign up free to answer and see why

Checkpoint

In a case round you’re asked to define success metrics for an action-taking AI agent that can send emails and make purchases. Strongest stack?

ADAU and revenue — the same metrics as any featureBSuccessful end-to-end task rate (NSM) · hallucination + action-reversibility (quality) · spend-cap adherence + kill-switch invocation rate (system) · plus a divergence alarmCJust measure how many actions it takes per session
Sign up free to answer and see why

Checkpoint

Your dashboard shows average latency at 600ms (fine) but users complain the assistant “feels slow.” What’s the likely metric mistake?

AThe complaints are subjective and can be ignored given the 600ms averageBYou’re reporting the mean — switch to p95/p99, where a slow tail (long outputs, cold caches) is what users experienceCLatency doesn’t affect AI product perception
Sign up free to answer and see why

Interview prep

Metric-definition questions are ~21–24% of modern PM interviews, and the AI versions add a twist the grader is specifically listening for: the divergence alarm. The pattern that wins every metrics question: refuse the single number, lay out the four-layer stack in 30 seconds, name the dominant metric per layer for the product type, then describe the held-out divergence detector by name. Weak answers list DAU and revenue and stop.
  1. 01“How do AI product metrics differ from normal engagement/retention metrics?” → AI metrics layer model-quality on top of product metrics and add a divergence alarm, because the two can drift apart silently.
  2. 02“North star is up — how do you know the model isn’t quietly worse?” → an independent, held-out, per-slice quality gate (e.g. faithfulness) plus escalation/correction-rate alerts.
  3. 03“Define metrics for an AI assistant.” → task completion (NSM) · hallucination + faithfulness (quality) · p95 latency (system) · divergence alarm.
  4. 04“Which metric is most likely to mislead you?” → blended ‘accuracy’/single score; also thumbs-up (a documented bias rewards confident phrasing over hedged-but-correct answers) and average (not p95) latency.
  5. 05“How did GitHub measure Copilot’s success?” → a layered stack: 96% same-day onboarding, >80% adoption, ~30% suggestion acceptance, 43% ‘extremely easy.’
  6. 06“Behaviour metric vs survey metric?” → prefer behaviour (acceptance, completion) over survey the moment you have signal; survey is a lagging backstop.
  7. 07“How often do you re-evaluate, and on what trigger?” → continuous online detectors + a weekly held-out gate that blocks model promotion on a quality delta.
  8. 08“Tie model quality to business outcome.” → bind each quality metric to a cost-of-error so a faithfulness drop maps to escalate-vs-monitor, not just a number.
Going deeper, the follow-ups that separate levels: “if model-quality improves but the product metric drops, what do you conclude?” (the quality gain isn’t reaching users — check UX, latency, or that you optimised the wrong quality axis); “how would you detect harm on a sub-population within 24 hours?” (per-slice quality dashboards + alerting, not a quarterly review); and “what’s your threshold for halting a rollout that’s improving engagement but degrading accuracy?” (a pre-declared faithfulness / hallucination floor that overrides engagement — the bridge to launch gates in L5). Always name the metric, the slice, and the trigger.

Could you lay out the four-layer metric stack for a given AI product and name the divergence alarm that catches a silent model regression?

New to itGetting thereConfident

Takeaways

  • Use a stack — business / model-quality / system — with one dominant metric per layer; never blend into a single score.
  • The AI-specific instrument is a divergence alarm: an independent, held-out, per-slice quality gate that fires when model quality drops as the NSM climbs.
  • GitHub Copilot/Accenture is a real four-layer stack: 96% onboarding, >80% adoption, ~30% acceptance, 43% “easy.”
  • Prefer behaviour over survey the moment you have signal; thumbs-up carries a documented bias toward confidently-worded answers over hedged-but-correct ones (EMBER).
  • Quote p95/p99 latency, not the mean; stratify accuracy; bind each quality metric to a cost-of-error.
  • Pre-stack metrics per product archetype so the case round is a recall, not an invention.

Next: designing the eval — the offline / online / human mix, and LLM-as-judge with its position and verbosity traps.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.