The four serving modes and how to pick one; the LLM serving stack that actually wins — continuous batching, PagedAttention, prefix caching — with the real multipliers; GPU economics and the two-number latency SLO (TTFT and ITL); model routing; and the autoscaling lag that bounds it all.
Where the design meets a clock and a bill
Bucket 6 — serving and operations — is where hand-waving costs the most points. The research is explicit: interviewers grade whether you can break a 300ms end-to-end budget into ~20ms inference + ~50ms feature fetch + ~230ms network/buffer, sketch the cost in dollars-per-1k-QPS, and name the failure modes that only appear in production. This lesson is the serving plane: the four modes, the LLM stack with real throughput multipliers, GPU economics, the two-number latency SLO, model routing, and the autoscaling lag that none of it escapes.
Chip Huyen’s taxonomy gives four serving modes, and naming the right one is the first move. Batch prediction (cron + materialized indexes) is cheapest and trivially correct but stale by minutes-to-hours — fine for churn scores or daily reports. Online synchronous (request-response) gives the freshest features and the simplest model but a source-of-truth join at request time is slow under burst. Online asynchronous (request + queue) absorbs spikes at the cost of 1–10s latency. Streaming continuously updates features into a continuously scoring model. Classify on two axes — latency-to-decision and freshness — and the mode falls out.
Latency budgets are first-order — and they come in two numbers for LLMs
For classical-ML interactive systems, the budget is one p99: real-time recsys target sub-100ms p99 for the whole pipeline; news feeds under 200ms; fraud under 250–300ms end-to-end with ~20ms for inference. Convert the prompt’s “how fresh?” into a millisecond budget before picking a model — that ordering is itself a senior tell. For LLMs the budget splits in two, and conflating them is the most common mistake. Time-to-first-token (TTFT) is the cost of prefill — compute-bound on prompt length. Inter-token latency (ITL) is the cost of decode — memory-bandwidth-bound, because the GPU must read the entire KV cache for every single token it generates.
code
1LLM LATENCY SLO -- two numbers, two levers (interactive chat)23 METRIC PHASE BOUND P99 TARGET FIRST LEVER4 ------ ------- --------------------------- ---------- -------------------5 TTFT prefill compute-bound on prompt len 300 ms shrink/cache prompt6 (500ms = perceived lag, (prefix cache)7 800ms = abandonment)8 ITL decode memory-bandwidth over KV 50 ms shrink KV footprint9 cache (read whole cache (quantize, smaller10 per token) model, spec-decode)1112 "Why not a bigger GPU?" -> ITL is bandwidth-bound over the KV cache; a bigger13 GPU helps prefill and concurrency, NOT per-token decode latency.
Interview angle. The curveball “why not just use a bigger GPU?” is testing exactly this. The answer: a bigger GPU helps prefill (more compute) and concurrency (more requests in flight), but per-token decode latency is bounded below by the time to read the KV cache once — it’s memory-bandwidth-limited, not compute-limited. So you cut ITL by shrinking the KV footprint (quantization, a smaller model, speculative decoding), not by buying a faster chip. Being able to say “TTFT and ITL are optimised with different levers” is the senior LLM-serving signal.
The LLM serving stack that wins: continuous batching + PagedAttention
The single largest throughput lever in self-hosted LLM serving is swapping naive request-level batching for continuous batching with PagedAttention (vLLM, SGLang, TGI). The numbers are dramatic and worth memorising. Anyscale’s OPT-13B benchmark: continuous batching alone gives 8× throughput over naive batching; with PagedAttention’s memory optimizations, up to 23×. vLLM’s own numbers: 14–24× over HuggingFace Transformers and 2.2–3.5× over TGI. PagedAttention treats the KV cache like a virtual-memory page table (fixed-size blocks, ~16 tokens each), cutting KV-cache memory waste from 60–80% down to under 4% — which let LMSYS cut its serving GPU count in half.
This connects directly to a failure mode that confuses people: KV-cache fragmentation OOMs even when total GPU memory looks free. Naive Transformers serving reserves a contiguous block sized for each request’s maximum sequence length; long-context requests fragment that reserved space into unusable holes, so new requests OOM despite apparent free memory. GMI Cloud lists “No KV-Cache Optimization” and “No Continuous Batching” among its top-10 production killers for exactly this reason. Interview angle. “Your inference cluster OOMs under load but nvidia-smi shows free memory — why?” → KV-cache fragmentation; the fix is paged attention, not more GPUs.
Prefix caching: the single largest latency/cost lever
For any workload with repeated prompt structure — chat, agents, RAG — prefix caching is the highest-ROI optimization. Anthropic’s prompt caching delivers up to 90% cost savings and 85% latency reduction: a cache read costs 0.1× base input, versus 1.25× for a 5-minute cache write (or 2× for a 1-hour TTL). The catch every senior must know: cache hits require an exact prefix match, including whitespace and formatting. So you engineer the prompt template deliberately — put the stable part (system prompt + retrieved context) first, the variable part (the user’s turn) last — so the same prefix precedes the variable suffix and ~80% of requests land on the cache-hit path.
code
1PROMPT CACHING ECONOMICS (Anthropic-style, base = 1.0x input)23 MODE MULTIPLIER NOTE4 ---------------------- ----------- ----------------------------------5 cache read (hit) 0.1x up to 90% cheaper, 85% faster6 cache write, 5-min TTL 1.25x pay once to populate7 cache write, 1-hr TTL 2.0x for long-lived stable prefixes8 standard input 1.0x baseline (no cache)910 RULE: stable prefix FIRST (system + retrieved context), variable LAST11 (user turn). Exact-match including whitespace -> design the template.1213 Also: output tokens cost ~3-5x input. Cheapest call = big CACHED prompt + short output.
A worked cost example from the research makes it concrete. A single 8× H100 node runs ~$1,800–$4,800/day on-demand before any traffic (H100 on-demand averages ~$3.15/hr across 46 providers, spot drops it ~10×). A 70B model at 16 req/s average / 60 req/s peak costs ~$34k/yr per replica. But if the median prompt is 2k tokens, prefix caching cuts prefill cost ~10× for the prefix, dropping steady-state GPU utilization from ~70% to ~25% — letting one node serve 2–3× the volume and halving amortized $/1k-tokens. This is the kind of dollars-per-QPS math interviewers want to see.
Model routing: the cheapest model that meets quality
Routing sets up the dual problem: pick the cheapest model that still clears the quality bar. vLLM’s Semantic Router frames this as Mixture-of-Models — one orchestrator routes each query across heterogeneous backends using domain, embedding, and other signals. The clever bit is the Modular LoRA architecture: instead of N full forward passes to classify, it does one base pass plus N lightweight LoRA adapters, keeping classification cost roughly constant as the model menu grows. Martian’s RouterBench formalises the cost/quality tradeoff. Interview angle. A router that always picks the cheap model is wrong 5–10% of the time; one that always picks the smart model is 3–10× more expensive — so you instrument router decisions separately from outcomes.
Autoscaling lag: the constraint nothing escapes
GPU autoscaling is fundamentally different from CPU autoscaling because of state and pre-allocation. Google’s GKE guide is explicit: do not scale on GPU memory for vLLM/TGI — they preallocate the KV cache, so memory-used “only works for scaling up, and won’t scale down.” Do not scale on GPU utilization alone either (it measures active time, not work). Scale on queue depth or batch size tied to a latency target — start the queue-depth target at 3–5 and raise it until the SLO is met. This is a precise, citable answer to “how would you autoscale this?”
The hard floor is cold start. An 8× H100 vLLM replica takes tens of seconds to load weights, capture CUDA graphs, and warm up — and on serverless GPU platforms it can be much worse: a public benchmark measured a vLLM cold start at 460 seconds (7m40s), unusable for interactive workloads. Since HPA/KEDA poll on a 15–60s interval and a new pod needs 10–90s to become ready, you cannot scale from zero inside a one-minute spike. The mitigation to name: a min-replica buffer sized to absorb the burst (predictive baseline for the diurnal cycle, elastic reactive for the residual), plus an atomic weight cache to dedupe downloads across replicas.
code
1GPU AUTOSCALING -- what to scale on (GKE guidance)23 SIGNAL GOOD FOR INFERENCE? IMPLICATION4 -------------------- ------------------- ---------------------------------5 GPU duty cycle (UTIL) NO measures active time, not work6 GPU memory (FB_USED) scale-UP only vLLM/TGI preallocate -> never7 scales down8 queue depth (server) YES start target 3-5, raise to hit SLO9 batch size (server) YES for latency-sensitive workloads1011 COLD START: 8xH100 vLLM = tens of seconds (serverless seen at 460s!).12 Poll interval 15-60s + pod-ready 10-90s => keep a MIN-REPLICA BUFFER.
Interview prep
Serving questions reward precise numbers and the prefill/decode mental model. Lead with the latency budget per stage, then the lever that moves it. Answer each in 60–90 seconds.
01“Break down a 300ms fraud budget.” → ~20ms inference + ~50ms feature-store fetch + ~230ms network/buffer; state per-stage budgets, not one number.
02“TTFT vs ITL?” → TTFT = prefill, compute-bound on prompt (cut by caching/shrinking prompt); ITL = decode, bandwidth-bound over the KV cache (cut by smaller KV / spec-decode).
03“Why continuous batching?” → 8× over naive, up to 23× with PagedAttention (Anyscale OPT-13B); it’s an 8× GPU-capex difference on the same workload.
04“Cluster OOMs but memory looks free — why?” → KV-cache fragmentation from contiguous max-length reservations; fix is paged attention, not more GPUs.
05“Why prefix caching?” → ~90% cost / ~85% latency on repeated prefixes; exact-match, so put the stable part first and the user turn last.
06“How do you autoscale GPU inference?” → on queue depth / batch size tied to a latency SLO; never on GPU memory (preallocated) or utilization alone.
07“Why not a bigger context model?” → cost + super-linear attention + quality decay; a medium model + reranker often beats it at ~20% of cost.
08“What’s your cost per 1k QPS?” → tokens × price (output ~3–5× input), GPU $/hr × replicas; show the prefix-cache savings (70%→25% utilization).
Going deeper, the follow-ups that probe ops maturity: “Black Friday / breaking-news 10× spike — what changes?” (min-replica buffer pre-warmed to P95, queue-depth autoscaling, shadow-mode any model upgrade, and degrade gracefully — serve a cheaper model rather than time out); “FP16 or FP8 on Hopper?” (FP8 is ~2× the throughput; running FP16 where FP8 works is GMI’s pitfall #5 — “wrong precision for the workload”); “single-region or multi-region?” (single-region is a total-downtime risk on a datacenter outage — name active-active failover); and “how do you bound a runaway cost?” (request-level logging + cost alerting — two more of the top-10 pitfalls; a silent bug with no cost alert is a real war story).
Your interactive LLM chat has acceptable TTFT but users complain text “types out” too slowly. Which change targets the right bottleneck?
AShrink or prefix-cache the prompt to speed up prefillBAdd more replicas behind the load balancerCReduce the KV-cache footprint — quantize, use a smaller model, or speculative decoding — because ITL is memory-bandwidth-bound over the KV cacheDIncrease the context window so the model has more room
Your self-hosted LLM cluster starts rejecting requests with OOM errors under load, but nvidia-smi shows several GB free per GPU. Most likely cause and fix?
AThe model weights are too large; switch to a smaller modelBKV-cache fragmentation — naive serving reserves contiguous max-length blocks per request, leaving unusable holes; switch to PagedAttention (vLLM/SGLang/TGI)CA memory leak in the tokenizer; restart the pods on a scheduleDInsufficient replicas; scale on GPU memory utilization
A RAG chat assistant reuses a 1,500-token system prompt and retrieved context on every call. How do you cut cost and latency most effectively?
ASwitch to a smaller model for everythingBLower the temperature to reduce token generationCEngineer the prompt so the stable prefix (system + retrieved context) comes first and the variable user turn last, then enable prefix caching for ~90% cost / ~85% latency on the prefixDIncrease max output tokens so fewer follow-up calls are needed
You’re configuring autoscaling for a vLLM deployment on GKE. Which scaling signal should you use?
AGPU memory utilization (DCGM_FI_DEV_FB_USED)BGPU duty cycle / utilizationCServer queue depth (or batch size) tied to a target latency, starting the queue target around 3–5 and tuning to hit the SLODRequests per second with a fixed replica count
An interviewer asks: “to halve per-token decode latency, would buying H200s instead of H100s help?” Strongest answer?
AYes — a faster GPU always lowers latency proportionallyBOnly partially — a bigger/faster GPU mainly helps prefill and concurrency; ITL is memory-bandwidth-bound over the KV cache, so the bigger lever is shrinking the KV footprint (quantization, smaller model, speculative decoding)CNo — GPU choice never affects latencyDYes — bigger memory means a smaller KV cache
Could you defend a per-stage latency budget, the LLM serving stack, GPU economics, and an autoscaling plan in a live round?
Not yetGetting thereConfident
Takeaways
Convert “how fresh?” into a per-stage millisecond budget before picking a model — and break the end-to-end SLO into named stages.
For LLMs, TTFT (prefill, compute-bound) and ITL (decode, bandwidth-bound over the KV cache) need different levers; a bigger GPU doesn’t fix ITL.
Continuous batching + PagedAttention is 8–23× over naive; KV-cache fragmentation causes OOMs with free memory — fix is paged attention.
Prefix caching is ~90% cost / ~85% latency on repeated prefixes, but exact-match — order the template stable-first, variable-last.
Autoscale GPU inference on queue depth / batch size tied to a latency SLO, never on GPU memory (preallocated) or utilization; keep a min-replica buffer for cold starts.
Model routing picks the cheapest model that clears the bar; a medium model + reranker often beats a giant-context call at ~20% of the cost.
Next: monitoring, drift & retraining — online vs offline metrics, drift detection, the retraining loop, rollback, and an incident scenario.