Latency is UX and cost is the bill — both are engineered. SLOs by use case, the prefill/decode latency model, streaming as perception, and the full lever stack: prompt caching, semantic caching, batching, routing/cascades, small fine-tunes — plus the FinOps interview round.
The two numbers that kill features in production
A feature that delights in a demo can die on contact with real traffic for two boring reasons: it’s too slow (users abandon) or too expensive (the bill outruns the value). Both are engineered, not fixed properties of “the model” — and the levers are stackable and mostly cheap. Notion cut p50 from ~2s to 350ms and frames it bluntly: “latency is perceived as quality.” This lesson is the mental model (where the time and money actually go) plus the ordered lever stack, ending with the FinOps round interviewers run on senior AI engineers.
Two numbers decide whether a feature survives: cost (tokens × price, across every call) and latency (prefill + decode + any retrieval). The single most important fact, carried over from L1: these decompose differently. Cost is dominated by output tokens (priced ~3–5× input) and total volume; latency splits into time-to-first-token (prefill, ∝ prompt length) and total time (∝ output length). You optimise them with different levers, and conflating “make it cheaper” with “make it faster” is the most common junior mistake.
1WHERE THE TIME AND MONEY GO (one call)23 latency = prefill(input_tokens) + output_tokens x inter-token-latency4 \__ sets TTFT (prompt len) __/ \__ sets TOTAL time (output len) __/56 cost = input_tokens x price_in + output_tokens x price_out7 (cacheable! 0.1x on a hit) (output ~3-5x input; never cached)89 Levers map onto the terms:10 shrink/cache PROMPT -> better TTFT + cheaper input11 shrink OUTPUT -> better total time + cheaper output (the big cost lever)12 cheaper MODEL/route -> lower price_in & price_out on the easy majority
Set latency SLOs from human impact, not your GPU’s median
The perception thresholds are well established: TTFT under ~200ms feels instant, under ~500ms responsive, beyond ~1s users notice, beyond ~2s they abandon. Different surfaces need different budgets:
code
1Latency SLOs by use case (TTFT = time to first token, ITL = inter-token)23 Use case TTFT (p99) ITL Stream?4 -------------------- ---------- ----- -------5 Voice agent ~150 ms 30 ms yes6 Interactive chat ~300 ms 50 ms yes7 Inline code complete ~500 ms 25 ms yes8 RAG-augmented chat 1500 ms+ 80 ms yes (prefill budget covers retrieval)9 Batch / async agent n/a - no1011 Source: Spheron LLM inference SLO guide
Interview angle. “What latency SLO would you set for this feature?” The senior move is to name the percentile and the surface, not quote a single average. Budget on p95/p99 (the tail is what users feel and remember), specify TTFT separately from total/ITL (a voice agent lives or dies on TTFT; a batch summariser doesn’t care), and tie the number to a human threshold (“>2s and they abandon, so p95 TTFT < 1s with streaming”). Averages hide the tail that actually drives churn — saying “p50 is 600ms” when p99 is 8s is the answer that loses the round.
Streaming: a perception lever, not a wall-clock one
Server-sent events cut perceived latency by up to ~80% by showing the first token immediately — but they don’t reduce total completion time. Default streaming for chat and inline completions. Two traps: with agents/tools, buffer tool-call argument deltas and validate structured output only at the end (never per chunk); with reasoning models, you pay for hidden thinking tokens even when streaming, the most common source of surprise bills. And honour client cancellation server-side, or you keep paying for a completion nobody is reading.
Lever 1 — prompt caching: the first lever, not the last
Provider prompt caching is where most teams leave the biggest, cheapest win on the table. OpenAI does it automatically on stable prefixes ≥ ~1k tokens for up to ~80% latency and ~90% input-cost reduction; Anthropic uses explicit breakpoints (reads at 0.1× base, writes at 1.25×/2×). The mechanism is the L1 prefill/KV-cache idea reused: a cache hit lets the provider skip recomputing the prefill for the shared prefix, so you save both the compute (latency) and the input-token charge. The one rule that decides whether you get it at all: keep dynamic content (timestamps, user ids, request ids) after the breakpoint — a single per-call token inside the cached prefix sends your hit-rate to zero.
code
1PROMPT CACHE: order the prompt so the prefix is byte-stable23 BAD (hit-rate ~0%) GOOD (hit-rate high)4 ----------------------------- -----------------------------5 "Today is 2026-06-23. You are..." "You are a support assistant..." <- cached prefix6 "...long stable instructions..." "...long stable instructions..." (>=1k tokens)7 "User 8841 asks: ..." --- cache breakpoint ---8 "Today is 2026-06-23. User 8841: ..." <- dynamic tail9 ^ timestamp+user id INSIDE prefix ^ all per-call content AFTER the prefix10 -> prefix differs every call -> prefix identical -> 0.1x input + ~80% faster TTFT
Levers 2–3 — semantic response caching & batching
(2) Semantic response caching serves a stored answer when a new query is semantically close to a past one (embedding similarity ≥ a threshold). Reported hit rates: 18–60% in RAG and ~20% in open Q&A at a 95% similarity threshold — and it skips the model entirely on a hit, so it cuts both cost and latency to near zero for repeat questions. Layer it after provider caching, and never serve stateful or freshness-sensitive answers from it (a tool call that mutates state, “what’s my balance,” “today’s incidents”). The threshold is the dangerous knob: too low and you serve a confidently-wrong cached answer to a subtly-different question. (3) Batching (OpenAI/Anthropic batch APIs) gives ~50% off in exchange for a 24h SLA — so it’s for offline work only: bulk embeddings, eval runs, backfills, nightly summarisation. Never on an interactive path.
Lever 4 — routing & cascades: pay frontier prices only for the hard tail
Most traffic is easy; a minority is hard. Routing & cascades exploit that: FrugalGPT (Chen et al.) matches GPT-4 quality at up to 98% lower cost by trying a cheap model first and escalating only when a scorer says the answer is weak; RouteLLM (lm-sys) cuts cost ~85% while keeping ~95% of GPT-4 quality with a learned router that picks the model before the call. Two patterns, one idea: a cascade runs cheap→expensive sequentially (you sometimes pay twice but never under-serve), a router classifies up front (one bill, but a misroute under-serves). What to watch at scale: the escalation rate (creeps up → your savings evaporate and you may now pay two bills), router/classifier drift (the traffic mix shifts and the router’s thresholds go stale), and the operational cost of running and evaluating two models instead of one.
Lever 5 — small fine-tunes & self-hosting for the narrow case
When the output distribution is narrow and the volume is high, a small fine-tuned model beats calling a frontier API. Character.ai’s Kaiju serves at roughly 13× lower cost than a frontier API by combining multi-query attention, sliding-window attention, cross-layer KV sharing, and int8 quantisation — engineering the inference stack, not just the weights. The tradeoff is real: you take on fine-tuning data curation, eval, hosting, GPU capacity planning, and the risk that the narrow model fails on the long tail the frontier model handled for free. So this is the last lever — reach for it only when caching, routing, and a cheaper API model have been exhausted and the workload is genuinely narrow and high-volume.
python
1# Worked monthly estimate — the math that should precede a model choice2reqs_per_day = 200_0003in_tokens = 1_200 # stable 1,000-token system prefix + ~200 dynamic4out_tokens = 2505price_in = 2.50 / 1_000_000 # $/input token (mid-tier model)6price_out = 10.00 / 1_000_000 # $/output token78base = reqs_per_day * 30 * (in_tokens*price_in + out_tokens*price_out)9# prompt caching: ~1,000 of the 1,200 input tokens are a cached prefix at 0.1x10cached_in = (1_000*price_in*0.1) + (200*price_in)11cached = reqs_per_day * 30 * (cached_in + out_tokens*price_out)12print(round(base), round(cached)) # ~$33k vs ~$20k/mo -- caching pays the rent
code
1THE LEVER STACK (apply cheapest-effort first)23 Lever Typical win Cost / catch4 -------------------- ----------------------- -----------------------------5 1 Prompt caching ~90% input, ~80% TTFT free; protect the stable prefix6 2 Semantic resp cache 18-60% (RAG) hit rate wrong-answer risk if threshold low;7 never for stateful/fresh queries8 3 Batch API ~50% off 24h SLA -> offline only9 4 Routing / cascade up to 98% (cost), 85% escalation-rate creep, drift,10 keeping ~95% quality two models to run + eval11 5 Small fine-tune / ~13x cheaper (Kaiju) data+hosting+eval burden; fails12 self-host on the long tail; narrow case only
Case studies & scale: the cost cliffs that only appear at volume
Notion serves 100M+ users by putting a smaller fine-tuned model on dedicated infra on the hot path instead of a frontier model, cutting chat p50 from ~2s to ~350ms — “latency is perceived as quality.” GitHub Copilot caps the assembled prompt near 6,000 characters specifically to keep fast models inside a tight inline-completion latency budget — a hard ceiling that also caps per-request cost. Perplexity sustains ~200M queries/day at p50 358ms / p95 <800ms by doing cheap retrieval first and running the expensive reranker only on the top candidates — the cascade idea applied to retrieval. The failure mode that only shows up at scale is the cost cliff: a feature that’s cheap at 1k req/day quietly becomes a five-figure monthly line item at 1M req/day, and the usual culprits are an un-cached prefix, an output that grew over time, a reasoning model left on the high-volume path, and a router whose escalation rate crept up unnoticed. The defence is the same as L1’s silent-regression defence applied to spend: a per-request cost/latency span and a dashboard, so the cliff shows up as a trend line, not a surprise invoice.
The cheapest LLM call is the one you don’t make (a cache hit), the second-cheapest is the one a small model handles (a route), and the most expensive is the frontier-model call you make by default because nobody measured the bill.
Interview prep
Cost/latency is where senior AI-engineering interviews get quantitative: expect a back-of-envelope estimate, “make this 5× cheaper,” and “set the SLOs.” They’re testing whether you think in tokens and percentiles, know the lever stack and its order, and can name the catch on each lever. Always state assumptions out loud (model, token counts, price/million, traffic) — the reasoning is what’s graded, not the exact dollar figure.
01“Estimate this feature’s monthly cost.” → requests × (in×price_in + out×price_out); output is ~3–5× input and usually dominates; state model, token counts, and price/million out loud.
02“What drives latency, and how do you cut it?” → TTFT = prefill (∝ prompt) → shrink/cache the prompt; total = output length → shrink output, stream, faster model. Different levers for each.
03“Make this 5× cheaper without tanking quality.” → cache the stable prefix (~90% input), route the easy majority to a cheap model, shrink output, batch any offline work — measure escalation rate.
04“Why is prompt caching the first lever?” → it skips prefill recompute on a hit (~90% input cost, ~80% TTFT), it’s automatic on OpenAI, and it compounds with every other lever — protect the prefix from dynamic tokens.
05“When does semantic caching help vs hurt?” → high-duplicate, non-stateful traffic (FAQ/RAG, 18–60% hit) helps; creative or stateful/fresh queries hurt, and a low similarity threshold serves confidently-wrong answers.
06“Cascade vs router?” → cascade tries cheap→expensive (never under-serves, sometimes pays twice); router classifies up front (one bill, a misroute under-serves) — both need the escalation/route rate monitored.
07“When is a fine-tune / self-host worth it?” → narrow output distribution + high volume (Kaiju ~13× cheaper); not worth the data/hosting/eval burden for broad or low-volume tasks.
08“What SLO would you set?” → name the percentile (p95/p99, not average) and the surface; TTFT separate from total/ITL; tie it to a human threshold (>2s → abandon).
09“Does streaming make it cheaper or faster?” → neither — it improves perceived latency; total time and cost are unchanged, and reasoning-model thinking tokens still bill while streaming.
10“How do you avoid a surprise bill at scale?” → a per-request cost/latency span + dashboard and spend caps at the gateway, so a cost cliff shows as a trend, not an invoice.
Push it deeper. Likely follow-ups: “Your cache hit-rate is 5% — debug it.” (a dynamic token sits inside the cached prefix; move timestamps/ids after the breakpoint). “Your cascade’s escalation rate climbed from 20% to 60% — what happened and what now?” (traffic got harder or the scorer drifted; re-tune the threshold or retrain the router — and note you may now be paying two bills, killing the savings). “Reasoning model or standard for a high-volume extraction?” (standard — reasoning’s hidden tokens blow cost and latency on a shallow, high-volume task; reserve reasoning for the hard, low-volume tail and route per request). “The bill doubled but traffic is flat — where do you look first?” (output-length creep, a model swap, a cache regression, or a router change — the per-request cost span tells you which).
Your prompt-cache hit-rate is near 0% despite a large, seemingly stable system prompt. Most likely cause?
AThe provider disabled caching for your accountBA dynamic value (timestamp, user id, request id) sits inside the cached prefix, so the prefix differs every requestCYour prompt is over 1,000 tokens, which is too long to cache
Which workload is semantic response caching most likely to help?
AAn open-ended creative-writing assistantBA support/FAQ bot over a stable knowledge base with repetitive questionsCA tool-calling agent that mutates account state
A feature sends a 1,200-token prompt (1,000 of it a stable system prefix) and returns ~250 tokens, on a frontier model, 200k req/day. Where is the biggest cost win?
ATrim the 200 dynamic input tokens down to 100BEnable prompt caching on the 1,000-token stable prefix (0.1× input) and route the easy majority to a cheaper modelCRaise max_tokens so answers finish in one call
A high-volume extraction endpoint (shallow task, strict structured output) is on a reasoning model “for accuracy,” and both cost and p99 latency are too high. Best move?
AMove the shallow, high-volume majority to a standard (non-reasoning) model and reserve the reasoning model for the genuinely hard tail, routing per requestBTurn on streaming so the reasoning model feels fasterCRaise the temperature to speed up generation
You add a cheap→expensive cascade and savings look great at launch, but two months later spend is back up though traffic is flat. Most likely cause and the right instrumentation?
AThe cheap model got more expensive — switch providersBThe escalation rate crept up (harder traffic or scorer drift), so more calls hit the expensive model — track escalation rate and per-request cost as trend lines, then re-tune the threshold/routerCPrompt caching stopped working — disable it
Could you set percentile SLOs, run a cost back-of-envelope, pick the right levers in order for a given traffic shape, and field the FinOps interview round?
Not yetMostlyConfident
Takeaways
Decompose first: TTFT = prefill (∝ prompt), total time = output length; cost is output-dominated (~3–5× input) — different levers for each term.
Set TTFT/ITL SLOs per surface on a percentile (p95/p99), not an average; budget for the tail users feel.
Streaming cuts perceived latency, not total time or cost — and is a trap for agents and reasoning models.
Prompt caching is the first, near-free lever (~90% input / ~80% TTFT); protect the stable prefix from dynamic tokens.
Semantic caching (high-duplicate, non-stateful), batching (offline, ~50%), routing/cascades (up to 98%/85%), and small fine-tunes (Kaiju ~13×) are the next levers — each with a catch to monitor.
Watch for the cost cliff at scale with a per-request cost/latency span and spend caps; escalation-rate creep and output growth are the usual culprits.
Finally: assemble all five lessons into one tested, structured, reliable microservice.