The economics an AI PM owns, not engineering: cost-per-task as the North Star, the output/input price asymmetry, real 2026 per-million-token numbers, why a routing cascade hits ~97% of frontier quality at ~24% of cost, the latency levers, and the debugging playbook when quality or cost regresses.
Cost is a product problem now
There is a hard line in AI-PM interviews and on the job: cost and latency are product concerns you own, not an engineering handoff. The candidate who says “engineering will figure it out” is visibly weaker than the one who can route, cap, and spec the economics. The reason is structural — routing every request to a frontier model can cost $0.50–$2.00 per active user per day, which at scale is enough to consume an entire product’s budget. So a senior AI PM reasons in four numbers per model in play: $/1M input tokens, $/1M output tokens, latency-to-first-token, and the context window — and makes cost-per-completed-task, not cost-per-token, the North Star.
Start with the asymmetry that drives everything: output tokens cost far more than input tokens — roughly 5×–8× more on frontier models. The 2026 published per-million-token prices make it concrete. The mechanism, from lesson 1: reading the prompt (prefill) is parallel and cheap; writing the answer (decode) is sequential and compute-heavy. Every product lever that prunes output — concise instructions, a JSON schema, a max-tokens cap — attacks the dominant cost directly, and every lever that reuses a cached prefix attacks the input cost. The cheapest feature is a big cached prompt with a short, structured output.
code
12026 PER-MILLION-TOKEN PRICES (representative; verify before quoting)23 Model $/1M in $/1M out out/in ratio4 ----------------------- -------- --------- ------------5 Grok 4.1 0.20 0.50 2.5x6 Gemini 3 Flash 0.50 3.00 6.0x7 Claude Haiku 4.5 1.00 5.00 5.0x8 OpenAI GPT-5.2 1.75 14.00 8.0x9 Gemini 3.1 Pro 2.00 12.00 6.0x10 Claude Sonnet 4.6 3.00 15.00 5.0x11 Claude Opus 4.8 5.00 25.00 5.0x1213 Two facts to carry: (1) output costs ~5-8x input -> prune what it WRITES.14 (2) the cheap-to-frontier spread is ~50x on input -> the right model per task is15 a bigger lever than any prompt tweak. Embeddings are cheap but NOT free:16 re-embedding a corpus on every change is a classic hidden cost spike.
Do the back-of-envelope, because interviewers grade the assumptions out loud. monthly_cost ≈ requests × (in_tokens × price_in + out_tokens × price_out). A support assistant at 200k requests/day with a 1,200-token input and 250-token output on a mid-tier model (≈$2.50/$10.00 per million) runs on the order of ~$1,100/day, ~$33k/month — and the output, despite being a fifth of the tokens, drives a large share of the bill. Change one variable and the lever becomes obvious: halving output tokens cuts both cost and latency; moving the easy 60% of traffic to a cheap model collapses the average. Interview angle. “Roughly what will this cost per month?” is a standard senior screen — they are testing whether you think in tokens and state assumptions, not whether you memorised a price.
The hidden costs the naive estimate misses
Two costs hide from the naive estimate, and a senior PM surfaces them. (1) Embeddings are cheap but not free. A cheap embedding model runs on the order of ~62,500 pages per $1, a larger one ~9,615 pages per $1 — a ~6.5× gap — and the trap is treating embedding as a one-time cost. Re-embedding a corpus on every document change (or every model swap) is one of the most common surprise cost spikes; freshness has a recompute bill you schedule, not a free lunch. (2) Retries multiply spend silently. Timeouts, rate-limits, and validation failures each trigger another full call, so a feature with a 5% retry rate quietly pays ~5% more — and far more if a flaky downstream tool sends an agent into a retry loop. Interview angle. “Where are the hidden costs in a RAG feature?” → the embedding/re-index bill and retry amplification, not the headline per-token price.
The cascade: ~97% of frontier quality at ~24% of cost
The most important architectural pattern for AI economics is the router / cascade. A published 2026 framework reaches 97.25% of GPT-4’s quality at only 24.18% of the cost under time constraints — proven, not theoretical. The mechanism: a small, cheap model handles the easy 60–80% of traffic, and a larger frontier model is called only when the small model signals low confidence (cascade) or when an upfront classifier decides the request is hard (route-by-intent). This is how you square “users want frontier quality” with “frontier pricing eats the margin” — you do not pay frontier prices on the requests that did not need them.
code
1ROUTING / CASCADE: pay frontier prices only when you must23 ROUTE BY INTENT (classify first) CASCADE (escalate on low confidence)4 ------------------------------------ ----------------------------------------5 cheap classifier reads the request cheap model answers first6 simple/chitchat/lookup -> small/ if it's confident -> done (most traffic)7 rule engine if low confidence -> escalate to frontier8 nuanced/multimodal -> frontier => ~small-model cost on the easy majority910 Published result: ~97% of GPT-4 quality at ~24% of cost under time constraints.11 The PM play is not "buy a router" by default -- it's DESIGN THE FALLBACK:12 cheap first, frontier on failure; measure quality on a fixed eval set per tier.
The numbers make the case concrete. Suppose 70% of your traffic is simple (lookups, chitchat, short replies) and 30% is genuinely hard. Serving everything on a frontier model is the baseline cost; routing the 70% to a model that is ~10× cheaper drops the blended cost to roughly 0.7 × 0.1 + 0.3 × 1.0 ≈ 0.37 of baseline — a ~63% cut — while the hard 30% still gets frontier quality. That single back-of-envelope is why routing is the highest-leverage economics move, and why the published cascade result (~24% of cost at ~97% of quality) is not surprising once you see the arithmetic.
The senior framing is restraint: the play is not “buy an AI gateway” by default — it is design the fallback. Cheap model first, frontier on failure; or a classification step that routes by intent. Managed gateways (Martian, Vera, and others) sell this as a product, and they can be the right buy — but only behind the same eval discipline as everything else, because a router that silently sends hard requests to the cheap model degrades quality invisibly. The metric that protects you is quality-on-a-fixed-eval-set per tier, so you can see exactly what the cheap path costs you in correctness.
A second economic lever lives in the embeddings layer for any RAG feature: dimension is a recurring bill, not a footnote. Bigger embedding vectors retrieve marginally better but cost more to compute, store, and serve, and — because re-embedding is periodic — that cost recurs every time the corpus changes. The PM instinct is to treat embedding-model and dimension choice as a cost decision with a quality tradeoff, and to schedule re-embedding (and budget for it) rather than discovering it as a surprise line item. This is why “just embed everything with the largest model” is rarely the right default at scale.
Latency: what users feel, and the levers
Latency splits into two numbers a PM should separate. Time-to-first-token (TTFT) is the “is it thinking?” gap; it tracks the prompt length (prefill). Total time tracks the output length (one decode step per token). They have different fixes: shrink or cache the prompt to cut TTFT; shrink the output, stream it, or pick a faster-decoding/smaller model to cut total time. The most important product move is streaming — showing tokens as they generate masks latency and dramatically improves perceived speed, which, per Notion, is perceived as quality. Reasoning models add a twist: their hidden “thinking” tokens add real latency and cost, so a reasoning-effort knob (low/medium/high) is a cost/speed dial you set per use case.
Context length is a quieter latency-and-cost knob. Every extra retrieved chunk adds input tokens (more prefill, more TTFT) and can dilute retrieval quality, so “add more context” is rarely free on either axis. And there is a cost lever that lives in the prompt: prompt caching. When the system prompt and few-shot examples are stable, caching that prefix cuts its input cost by up to ~90% and reduces TTFT — which is why you put the stable, expensive parts of the prompt first. Anthropic’s guidance even frames fine-tuning partly as a latency/cost play: distillation lets you move logic out of the prompt (paid every call) into the weights (paid once), shortening prompts permanently.
One more economics judgment shows up in strategy rounds: “should we make the current model cheaper or invest in the next, more capable one?” The weak answer is “better models win.” The strong answer reasons about the whole P&L: what does each option do to margin (a cheaper model widens it now), to competitive pressure (Claude and Gemini are pricing against you, so a capability lead may be short-lived), and to the model timeline (frontier prices fall fast — today’s expensive model is next quarter’s mid-tier). Capability lift is not the same as user-outcome lift or business value; the senior instinct is to tie the choice to retention and unit economics, not to a benchmark.
Routing every interaction to a frontier model can cost $0.50–$2.00 per active user per day — at scale, enough to consume the entire business budget. The pattern that wins is a routing graph that sends nuanced, multimodal prompts to the frontier model and simple lookups to a cheap tier. — the AI PM’s menu.
Streaming deserves its own line because it is the cheapest perceived-latency win you have. Total generation time does not change, but showing tokens as they appear converts a 4-second silent wait into a response that feels immediate and lets the user start reading — the same “latency is perceived as quality” lever Notion leaned on. The product cost is real engineering (server-sent events, partial-render UI, a cancel/regenerate control), but the payoff is large and it stacks with the real cost cuts (caching, routing, output caps). The senior framing: cut actual latency where you can, and mask the rest with streaming.
Debugging a regression: cost or quality moved
Two interview staples double as on-the-job playbooks. “Your chatbot’s quality dropped 15% last week — debug it.” The weak answer is “check the logs.” The strong answer inventories the real causes in order: a prompt change shipped without an eval, a silent model update by the provider (pin versions!), data drift in inputs or the retrieval corpus, an infra rollback or config change, rate-limiting or truncation under load, and a retrieval regression (a re-index or parser change). Each is checkable, and the meta-lesson is that you cannot debug what you did not instrument — which is why pinned versions plus an eval suite are the prerequisite, not the afterthought.
“Your costs spiked — find it.” Same discipline: output tokens crept up (a verbosity change, reasoning-effort raised, a prompt that now invites longer answers), prompt-cache hit-rate fell (the stable prefix changed, so every call pays full input), traffic mix shifted toward the frontier tier (the router is sending more to the expensive path), retries multiplied under errors, or an embedding re-index ran (the classic hidden RAG spike). The dashboard a senior PM insists on, alongside standard product analytics: cost-per-completed-task, token usage by tier, cache hit-rate, and p50/p95 latency. Interview angle. Naming cost-per-task (not cost-per-request) and the cache hit-rate as first-class metrics is a strong senior signal.
Interview prep
Economics rounds test whether you own cost, latency, and quality as product metrics. Interviewers probe four things: can you estimate cost in tokens with stated assumptions, do you reach for routing/caching/output-capping rather than “engineering will handle it,” can you separate TTFT from total latency, and do you debug a regression by named cause. Speak in $/user/day, $/task, and percentages on an eval — not in vibes.
01“Estimate this feature’s monthly cost.” → requests × (in×price_in + out×price_out); output ~5–8× input; state model and token assumptions out loud.
02“Make it cheaper without tanking quality.” → cap/shorten output, cache the stable prefix, and route the easy majority to a cheap model — a cascade hits ~97% quality at ~24% cost.
03“Why care about tokens as a PM?” → they set cost (billed per token), latency (scales with output), and the context limit — they are the feature’s unit economics.
04“Frontier model for everything — what’s wrong?” → ~$0.50–$2.00/user/day at scale can eat the budget; route by intent and reserve frontier for nuanced/multimodal requests.
05“Cut latency.” → separate TTFT (shrink/cache the prompt) from total time (shrink/stream output, smaller model); streaming masks the wait.
06“Chatbot quality dropped 15% — debug.” → prompt change, silent model update, data drift, infra rollback, rate-limiting, retrieval regression — not “check the logs.”
07“Costs spiked — find it.” → longer outputs, cache-hit-rate drop, traffic shifting to frontier, retries, or an embedding re-index; watch cost-per-task and cache hit-rate.
08“GPT-4 cheaper or invest in GPT-5?” → reason about margins, rival pressure (Claude/Gemini), and model timelines — not “better models win.”
Going deeper, the follow-ups test architectural and economic literacy together: “where would you put a cache?” (the stable system-prompt/few-shot prefix, and embeddings of unchanged docs); “what’s your routing decision for this traffic?” (classify, send simple/lookup to small or a rule engine, nuanced to frontier); “how does a reasoning model change the math?” (hidden thinking tokens add cost and latency; use the effort knob, route per request); and “what do you monitor to know cost is regressing?” (cost-per-task, tokens-by-tier, cache hit-rate, p95). The through-line: every economic answer ends in a metric and a fallback design, never “engineering owns that.”
A consumer feature routes every request — from “hi” to complex multi-doc analysis — to a frontier model, and the per-user-per-day cost is threatening the business case. Highest-leverage first move?
ANegotiate a volume discount with the provider and keep routing everything to frontierBAdd a router/cascade: a cheap model (or rule engine) handles the easy majority, escalating to frontier only on low confidence or nuanced intentCIncrease the context window so each call does more work
Users complain the assistant “feels slow to start” even though the full answer arrives reasonably fast. The prompt is large; the output is short. What’s the targeted fix?
ASwitch to a model with a larger context windowBThe slow start is TTFT, driven by the large prompt — shrink and cache the stable prefix (and stream output) to cut time-to-first-tokenCCap the output tokens lower
Overnight, answer quality drops ~15% with no deploy from your team. What’s the senior debugging order?
AInventory real causes — silent provider model update (pin versions), data/retrieval drift, infra rollback, rate-limiting/truncation, a prompt change — and check against your eval setBJust check the logs and wait to see if it recoversCImmediately fine-tune the model to recover quality
A RAG feature’s monthly bill suddenly doubles with flat traffic and unchanged prompts. Which hidden cause should you check first?
AOutput tokens must have doubledBA full embedding re-index ran (re-embedding the corpus) — the classic hidden RAG cost spike — so check ingestion/re-index jobs and cache hit-rateCThe provider raised prices, nothing to do
An exec asks you to set the North Star metric for an AI feature’s economics. What do you propose, and why not cost-per-token?
ACost-per-token, because it’s what the provider billsBNumber of API calls, since it’s easy to trackCCost-per-completed-task (total spend ÷ successful tasks), supported by tokens-by-tier, cache hit-rate, and p95 latency
Could you estimate a feature’s cost in tokens, design a routing/caching plan to hit a budget, separate TTFT from total latency, and debug a cost or quality regression by cause?
New to itGetting thereConfident
Takeaways
Cost and latency are PM-owned product metrics; the North Star is cost-per-completed-task, not cost-per-token.
Output tokens cost ~5–8× input — prune what the model writes (caps, schemas, concision) and cache the stable prefix (~90% off input).
A routing cascade reaches ~97% of frontier quality at ~24% of cost — cheap model first, frontier on failure; design the fallback, don’t reflexively buy a gateway.
TTFT (shrink/cache the prompt) and total time (shrink/stream the output, smaller model) are different dials; streaming masks latency.
Debug regressions by named cause (silent model update, drift, prompt change, cache-miss, re-index, retries) — pin versions and instrument cost-per-task.
Reason about model choice with margins and rival pressure; the cheap-to-frontier spread (~50× on input) is a bigger lever than any prompt tweak.
Next: the capstone — map model capabilities to a concrete product feature set, end to end.