Lesson 3 of 6 · 48 min

Serving, latency & cost at scale

The four serving modes and how to pick one; the LLM serving stack that actually wins — continuous batching, PagedAttention, prefix caching — with the real multipliers; GPU economics and the two-number latency SLO (TTFT and ITL); model routing; and the autoscaling lag that bounds it all.

Where the design meets a clock and a bill

Bucket 6 — serving and operations — is where hand-waving costs the most points. The research is explicit: interviewers grade whether you can break a 300ms end-to-end budget into ~20ms inference + ~50ms feature fetch + ~230ms network/buffer, sketch the cost in dollars-per-1k-QPS, and name the failure modes that only appear in production. This lesson is the serving plane: the four modes, the LLM stack with real throughput multipliers, GPU economics, the two-number latency SLO, model routing, and the autoscaling lag that none of it escapes.
Chip Huyen’s taxonomy gives four serving modes, and naming the right one is the first move. Batch prediction (cron + materialized indexes) is cheapest and trivially correct but stale by minutes-to-hours — fine for churn scores or daily reports. Online synchronous (request-response) gives the freshest features and the simplest model but a source-of-truth join at request time is slow under burst. Online asynchronous (request + queue) absorbs spikes at the cost of 1–10s latency. Streaming continuously updates features into a continuously scoring model. Classify on two axes — latency-to-decision and freshness — and the mode falls out.
vLLM: Easy, Fast, and Cheap LLM Serving for EveryoneWoosuk Kwon & Xiaoxuan Liu (PyTorch / UC Berkeley)

Latency budgets are first-order — and they come in two numbers for LLMs

For classical-ML interactive systems, the budget is one p99: real-time recsys target sub-100ms p99 for the whole pipeline; news feeds under 200ms; fraud under 250–300ms end-to-end with ~20ms for inference. Convert the prompt’s “how fresh?” into a millisecond budget before picking a model — that ordering is itself a senior tell. For LLMs the budget splits in two, and conflating them is the most common mistake. Time-to-first-token (TTFT) is the cost of prefill — compute-bound on prompt length. Inter-token latency (ITL) is the cost of decode — memory-bandwidth-bound, because the GPU must read the entire KV cache for every single token it generates.
code
1LLM LATENCY SLO -- two numbers, two levers (interactive chat)23  METRIC   PHASE     BOUND                         P99 TARGET   FIRST LEVER4  ------   -------   ---------------------------   ----------   -------------------5  TTFT     prefill   compute-bound on prompt len   300 ms       shrink/cache prompt6                       (500ms = perceived lag,                    (prefix cache)7                        800ms = abandonment)8  ITL      decode    memory-bandwidth over KV      50 ms        shrink KV footprint9                       cache (read whole cache                    (quantize, smaller10                        per token)                                 model, spec-decode)1112  "Why not a bigger GPU?" -> ITL is bandwidth-bound over the KV cache; a bigger13  GPU helps prefill and concurrency, NOT per-token decode latency.
Interview angle. The curveball “why not just use a bigger GPU?” is testing exactly this. The answer: a bigger GPU helps prefill (more compute) and concurrency (more requests in flight), but per-token decode latency is bounded below by the time to read the KV cache once — it’s memory-bandwidth-limited, not compute-limited. So you cut ITL by shrinking the KV footprint (quantization, a smaller model, speculative decoding), not by buying a faster chip. Being able to say “TTFT and ITL are optimised with different levers” is the senior LLM-serving signal.

The LLM serving stack that wins: continuous batching + PagedAttention

The single largest throughput lever in self-hosted LLM serving is swapping naive request-level batching for continuous batching with PagedAttention (vLLM, SGLang, TGI). The numbers are dramatic and worth memorising. Anyscale’s OPT-13B benchmark: continuous batching alone gives 8× throughput over naive batching; with PagedAttention’s memory optimizations, up to 23×. vLLM’s own numbers: 14–24× over HuggingFace Transformers and 2.2–3.5× over TGI. PagedAttention treats the KV cache like a virtual-memory page table (fixed-size blocks, ~16 tokens each), cutting KV-cache memory waste from 60–80% down to under 4% — which let LMSYS cut its serving GPU count in half.
This connects directly to a failure mode that confuses people: KV-cache fragmentation OOMs even when total GPU memory looks free. Naive Transformers serving reserves a contiguous block sized for each request’s maximum sequence length; long-context requests fragment that reserved space into unusable holes, so new requests OOM despite apparent free memory. GMI Cloud lists “No KV-Cache Optimization” and “No Continuous Batching” among its top-10 production killers for exactly this reason. Interview angle. “Your inference cluster OOMs under load but nvidia-smi shows free memory — why?” → KV-cache fragmentation; the fix is paged attention, not more GPUs.

Prefix caching: the single largest latency/cost lever

For any workload with repeated prompt structure — chat, agents, RAG — prefix caching is the highest-ROI optimization. Anthropic’s prompt caching delivers up to 90% cost savings and 85% latency reduction: a cache read costs 0.1× base input, versus 1.25× for a 5-minute cache write (or 2× for a 1-hour TTL). The catch every senior must know: cache hits require an exact prefix match, including whitespace and formatting. So you engineer the prompt template deliberately — put the stable part (system prompt + retrieved context) first, the variable part (the user’s turn) last — so the same prefix precedes the variable suffix and ~80% of requests land on the cache-hit path.
code
1PROMPT CACHING ECONOMICS (Anthropic-style, base = 1.0x input)23  MODE                     MULTIPLIER    NOTE4  ----------------------   -----------   ----------------------------------5  cache read (hit)         0.1x          up to 90% cheaper, 85% faster6  cache write, 5-min TTL   1.25x         pay once to populate7  cache write, 1-hr TTL    2.0x          for long-lived stable prefixes8  standard input           1.0x          baseline (no cache)910  RULE: stable prefix FIRST (system + retrieved context), variable LAST11        (user turn). Exact-match including whitespace -> design the template.1213  Also: output tokens cost ~3-5x input. Cheapest call = big CACHED prompt + short output.
A worked cost example from the research makes it concrete. A single 8× H100 node runs ~$1,800–$4,800/day on-demand before any traffic (H100 on-demand averages ~$3.15/hr across 46 providers, spot drops it ~10×). A 70B model at 16 req/s average / 60 req/s peak costs ~$34k/yr per replica. But if the median prompt is 2k tokens, prefix caching cuts prefill cost ~10× for the prefix, dropping steady-state GPU utilization from ~70% to ~25% — letting one node serve 2–3× the volume and halving amortized $/1k-tokens. This is the kind of dollars-per-QPS math interviewers want to see.

Model routing: the cheapest model that meets quality

Routing sets up the dual problem: pick the cheapest model that still clears the quality bar. vLLM’s Semantic Router frames this as Mixture-of-Models — one orchestrator routes each query across heterogeneous backends using domain, embedding, and other signals. The clever bit is the Modular LoRA architecture: instead of N full forward passes to classify, it does one base pass plus N lightweight LoRA adapters, keeping classification cost roughly constant as the model menu grows. Martian’s RouterBench formalises the cost/quality tradeoff. Interview angle. A router that always picks the cheap model is wrong 5–10% of the time; one that always picks the smart model is 3–10× more expensive — so you instrument router decisions separately from outcomes.

Autoscaling lag: the constraint nothing escapes

GPU autoscaling is fundamentally different from CPU autoscaling because of state and pre-allocation. Google’s GKE guide is explicit: do not scale on GPU memory for vLLM/TGI — they preallocate the KV cache, so memory-used “only works for scaling up, and won’t scale down.” Do not scale on GPU utilization alone either (it measures active time, not work). Scale on queue depth or batch size tied to a latency target — start the queue-depth target at 3–5 and raise it until the SLO is met. This is a precise, citable answer to “how would you autoscale this?”
The hard floor is cold start. An 8× H100 vLLM replica takes tens of seconds to load weights, capture CUDA graphs, and warm up — and on serverless GPU platforms it can be much worse: a public benchmark measured a vLLM cold start at 460 seconds (7m40s), unusable for interactive workloads. Since HPA/KEDA poll on a 15–60s interval and a new pod needs 10–90s to become ready, you cannot scale from zero inside a one-minute spike. The mitigation to name: a min-replica buffer sized to absorb the burst (predictive baseline for the diurnal cycle, elastic reactive for the residual), plus an atomic weight cache to dedupe downloads across replicas.
code
1GPU AUTOSCALING -- what to scale on (GKE guidance)23  SIGNAL                 GOOD FOR INFERENCE?   IMPLICATION4  --------------------   -------------------   ---------------------------------5  GPU duty cycle (UTIL)  NO                    measures active time, not work6  GPU memory (FB_USED)   scale-UP only         vLLM/TGI preallocate -> never7                                                 scales down8  queue depth (server)   YES                   start target 3-5, raise to hit SLO9  batch size (server)    YES                   for latency-sensitive workloads1011  COLD START: 8xH100 vLLM = tens of seconds (serverless seen at 460s!).12  Poll interval 15-60s + pod-ready 10-90s => keep a MIN-REPLICA BUFFER.

Interview prep

Serving questions reward precise numbers and the prefill/decode mental model. Lead with the latency budget per stage, then the lever that moves it. Answer each in 60–90 seconds.
  1. 01“Break down a 300ms fraud budget.” → ~20ms inference + ~50ms feature-store fetch + ~230ms network/buffer; state per-stage budgets, not one number.
  2. 02“TTFT vs ITL?” → TTFT = prefill, compute-bound on prompt (cut by caching/shrinking prompt); ITL = decode, bandwidth-bound over the KV cache (cut by smaller KV / spec-decode).
  3. 03“Why continuous batching?” → 8× over naive, up to 23× with PagedAttention (Anyscale OPT-13B); it’s an 8× GPU-capex difference on the same workload.
  4. 04“Cluster OOMs but memory looks free — why?” → KV-cache fragmentation from contiguous max-length reservations; fix is paged attention, not more GPUs.
  5. 05“Why prefix caching?” → ~90% cost / ~85% latency on repeated prefixes; exact-match, so put the stable part first and the user turn last.
  6. 06“How do you autoscale GPU inference?” → on queue depth / batch size tied to a latency SLO; never on GPU memory (preallocated) or utilization alone.
  7. 07“Why not a bigger context model?” → cost + super-linear attention + quality decay; a medium model + reranker often beats it at ~20% of cost.
  8. 08“What’s your cost per 1k QPS?” → tokens × price (output ~3–5× input), GPU $/hr × replicas; show the prefix-cache savings (70%→25% utilization).
Going deeper, the follow-ups that probe ops maturity: “Black Friday / breaking-news 10× spike — what changes?” (min-replica buffer pre-warmed to P95, queue-depth autoscaling, shadow-mode any model upgrade, and degrade gracefully — serve a cheaper model rather than time out); “FP16 or FP8 on Hopper?” (FP8 is ~2× the throughput; running FP16 where FP8 works is GMI’s pitfall #5 — “wrong precision for the workload”); “single-region or multi-region?” (single-region is a total-downtime risk on a datacenter outage — name active-active failover); and “how do you bound a runaway cost?” (request-level logging + cost alerting — two more of the top-10 pitfalls; a silent bug with no cost alert is a real war story).
articlevLLM: Easy, Fast, and Cheap LLM Serving with PagedAttentionvLLM (Kwon et al., UC Berkeley)articleAchieve 23x LLM Inference Throughput (continuous batching benchmark)AnyscaledocsBest practices for autoscaling LLM inference (queue depth, not GPU memory)Google Cloud (GKE)docsPrompt caching — exact-prefix economics and designClaude API Docs

Checkpoint

Your interactive LLM chat has acceptable TTFT but users complain text “types out” too slowly. Which change targets the right bottleneck?

AShrink or prefix-cache the prompt to speed up prefillBAdd more replicas behind the load balancerCReduce the KV-cache footprint — quantize, use a smaller model, or speculative decoding — because ITL is memory-bandwidth-bound over the KV cacheDIncrease the context window so the model has more room
Sign up free to answer and see why

Checkpoint

Your self-hosted LLM cluster starts rejecting requests with OOM errors under load, but nvidia-smi shows several GB free per GPU. Most likely cause and fix?

AThe model weights are too large; switch to a smaller modelBKV-cache fragmentation — naive serving reserves contiguous max-length blocks per request, leaving unusable holes; switch to PagedAttention (vLLM/SGLang/TGI)CA memory leak in the tokenizer; restart the pods on a scheduleDInsufficient replicas; scale on GPU memory utilization
Sign up free to answer and see why

Checkpoint

A RAG chat assistant reuses a 1,500-token system prompt and retrieved context on every call. How do you cut cost and latency most effectively?

ASwitch to a smaller model for everythingBLower the temperature to reduce token generationCEngineer the prompt so the stable prefix (system + retrieved context) comes first and the variable user turn last, then enable prefix caching for ~90% cost / ~85% latency on the prefixDIncrease max output tokens so fewer follow-up calls are needed
Sign up free to answer and see why

Checkpoint

You’re configuring autoscaling for a vLLM deployment on GKE. Which scaling signal should you use?

AGPU memory utilization (DCGM_FI_DEV_FB_USED)BGPU duty cycle / utilizationCServer queue depth (or batch size) tied to a target latency, starting the queue target around 3–5 and tuning to hit the SLODRequests per second with a fixed replica count
Sign up free to answer and see why

Checkpoint

An interviewer asks: “to halve per-token decode latency, would buying H200s instead of H100s help?” Strongest answer?

AYes — a faster GPU always lowers latency proportionallyBOnly partially — a bigger/faster GPU mainly helps prefill and concurrency; ITL is memory-bandwidth-bound over the KV cache, so the bigger lever is shrinking the KV footprint (quantization, smaller model, speculative decoding)CNo — GPU choice never affects latencyDYes — bigger memory means a smaller KV cache
Sign up free to answer and see why

Could you defend a per-stage latency budget, the LLM serving stack, GPU economics, and an autoscaling plan in a live round?

Not yetGetting thereConfident

Takeaways

  • Convert “how fresh?” into a per-stage millisecond budget before picking a model — and break the end-to-end SLO into named stages.
  • For LLMs, TTFT (prefill, compute-bound) and ITL (decode, bandwidth-bound over the KV cache) need different levers; a bigger GPU doesn’t fix ITL.
  • Continuous batching + PagedAttention is 8–23× over naive; KV-cache fragmentation causes OOMs with free memory — fix is paged attention.
  • Prefix caching is ~90% cost / ~85% latency on repeated prefixes, but exact-match — order the template stable-first, variable-last.
  • Autoscale GPU inference on queue depth / batch size tied to a latency SLO, never on GPU memory (preallocated) or utilization; keep a min-replica buffer for cold starts.
  • Model routing picks the cheapest model that clears the bar; a medium model + reranker often beats a giant-context call at ~20% of the cost.

Next: monitoring, drift & retraining — online vs offline metrics, drift detection, the retraining loop, rollback, and an incident scenario.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.

Serving, latency & cost at scale · ML & LLM System Design…