The model is trained — now serve it cheaply. Quantization quality is a function of method not bit-width (GPTQ/AWQ/fp8/GGUF with measured perplexity deltas), PagedAttention kills the 60–80% KV-cache waste, continuous batching is the 23–36x system win, speculative decoding buys 2–3x single-stream latency, and the vLLM/TGI/SGLang/TensorRT-LLM choice is workload- not vendor-driven — with the production numbers to budget against.
Inference is where the money goes
A fine-tune is a one-time cost; serving it is a bill you pay on every token, forever — which is why the interview research finds ~40% of an LLM-engineer loop is now inference, not modelling. The wins are large and they compose: PagedAttention (2–4x), continuous batching (up to 23–36x over naive), FP8 (~1.5x), speculative decoding (2–3x at batch 1). But each attacks a different term of the cost equation and each has a regime where it stops helping. This lesson is the five levers, the measured numbers, and the decision frame that says which lever a given SLO actually needs.
Start with the cost equation, because every optimisation maps to one term of it. A forward pass on a 70B-class model splits into two phases with different bottlenecks. Prefill (processing the prompt) is compute-bound for long prompts — it does a big parallel matmul over all input tokens. Decode (generating tokens one at a time) is memory-bandwidth-bound — each new token streams the entire weight set and the growing KV cache through the GPU, doing tiny matmuls. So latency-to-first-token (TTFT) is a prefill problem and time-per-output-token (TPOT) is a decode problem, and the five levers split along exactly that line: quantization shrinks the weight stream, KV paging shrinks the cache, batching fills idle decode cycles, speculative decoding amortises the per-token bandwidth, and stack choice composes them all.
The senior framing the research keeps hammering: quantization quality is a function of method, not bit-width. Two 4-bit quantizers can differ wildly, and the same bit-width can be production-safe or unusable depending on the algorithm. Know four families by mechanism. GPTQ (Frantar et al., ICLR 2023) is a one-shot, weight-only quantizer that uses second-order (Hessian) information to minimise per-layer reconstruction error, calibrating on ~128 sequences of 2048 tokens — it quantized OPT-175B in ~4 GPU-hours. AWQ (Lin et al., MLSys 2024) observes that protecting just the ~1% salient weight channels (identified from activation magnitudes, not weights) captures most of the quality, then scales those channels before uniform int4 — and it ships a fused dequant+GEMM kernel. FP8 (E4M3/E5M2) is the hardware-native path on H100/Blackwell. GGUF / llama.cpp uses mixed-precision "k-quant" blocks (e.g. Q4_K_M, Q5_K_M) that vary bit-width across a single matrix — the CPU/Apple-Silicon default.
code
1QUANTIZATION QUALITY (method drives the recoverable fraction, not bits)23 Method Bits Model Quality delta vs FP16 Speedup / note4 ------ ---- ---------- -------------------------- ------------------5 GPTQ int4 OPT-175B C4 ppl 10.13 -> 10.28 (+.15) 3.25-4.53x @3-bit6 GPTQ int3 OPT-175B 10.13 -> 10.67 (+0.54) single-GPU-only floor7 AWQ int4 Llama-2-70B protects ~1% salient chans >3x vs HF FP16 (kernel)8 FP8 fp8 Mistral-7B ~0 perplexity delta +33% tok/s, -24% $/Mtok9 GGUF Q5_K_M Llama-7B +0.0142 ppl 4.45 GB (vs 13.0 GB)10 GGUF Q4_K_M Llama-7B +0.0535 ppl 3.80 GB (recommended floor)1112 At 4-bit on 100B+ models BOTH GPTQ & AWQ hold <0.2 ppl delta; AWQ's edge is the13 fused kernel, NOT quality. int3 (+0.54) is reserved for must-fit-one-GPU cases.
The numbers decouple quality from speedup at modern bit-widths. GPTQ at int4 on OPT-175B moves C4 perplexity only 10.13 → 10.28 (+0.15); even int3 is just +0.54. AWQ at int4 on Llama-2-70B matches that quality and runs >3x faster than HF FP16 — the win is the kernel, not the bits. FP8 on a Mistral-7B/H100 stack (Baseten) measured +33% tokens/sec, +31% throughput, −8.5% TTFT, −24% cost per million tokens at near-zero perplexity delta; NVIDIA on Mixtral 8x7B with TensorRT-LLM saw ~50% more throughput within a 0.5s budget. Interview angle. The canonical question — "deploy a 70B on 8×A100 40GB, walk me through quantization" — wants you to reason from memory and hardware: FP16 ≈ 140GB (won’t fit), int8 ≈ 70GB (fits), int4 (AWQ/GPTQ) ≈ 35GB (comfortable with KV headroom), and FP8 is not an option on A100 (no FP8 tensor cores) — naming that hardware constraint is the senior tell.
So the choice rule is "pair the quantizer with the kernel stack that consumes it." AWQ + fused kernel (vLLM/TensorRT-LLM) for GPU when memory is the constraint; FP8 + TensorRT-LLM/vLLM as the default on Hopper/Blackwell because the GPU already runs FP8 GEMMs at ~2x FP16 throughput; GGUF Q4_K_M/Q5_K_M for CPU and Apple Silicon; SmoothQuant int8 on Ampere. Interview angle. "GPTQ vs AWQ in practice?" → similar int4 perplexity (~0.1–0.3 delta vs FP16), so the decision is almost always kernel ecosystem and throughput in your serving stack, not accuracy — and per-channel (per-row) scaling is standard for both because one scale-per-output-channel slashes error versus a single per-tensor scalar.
KV-cache reuse & PagedAttention — killing the 60–80% waste
The KV cache is the hidden memory hog of decode: every generated token must attend to all previous tokens, so the keys/values for the whole sequence are cached and grow with output length. A 13B model burns ~1 MB of KV state per token; on an A100 40GB, after ~26GB of weights you have room for only ~14k tokens of cache, and at 2048-token sequences that caps you at ~7 concurrent requests. The classic bug: allocating one contiguous max-length KV tensor per request wastes 60–80% of reserved memory to internal fragmentation and over-reservation. PagedAttention (Kwon et al., SOSP 2023) fixes it by borrowing OS virtual memory: the KV cache is split into fixed 4–16 token pages, a per-request block table maps logical→physical pages allocated on demand, and the result is <4% waste with copy-on-write sharing of identical prefixes.
code
1PAGEDATTENTION (KV cache as a page table, not one big tensor)23 NAIVE: one contiguous max-len KV tensor per request4 -> up to 60-80% wasted (internal fragmentation + over-reservation)56 PAGED: logical KV split into 4-16 token PAGES; block-table maps log->phys7 -> <4% waste (only the last partial page); pages allocated on demand8 -> shared prefixes (system prompt, few-shot, RAG doc) = ONE physical copy9 referenced by many block tables (copy-on-write) ["prefix caching"]1011 Measured: vLLM 2-4x throughput vs FasterTransformer/Orca at SAME latency.12 SGLang RadixAttention generalises sharing to a tree -> 74.1% hit rate (Vicuna prod).
The payoff scales with prefix overlap. The PagedAttention paper measured 2–4x throughput at the same latency versus FasterTransformer/Orca. The big multiplier is prefix caching: when many requests share a long system prompt or RAG document, the prefix’s KV blocks are computed once and reused — the scheduler allocates only the per-request delta. SGLang generalises this to a radix tree (RadixAttention) keyed on token prefixes and measured 52.4–99% cache hit rates, with 74.1% in Vicuna-33B production translating to a 1.7x first-token-latency reduction. LinkedIn’s production vLLM deployment reports the compound effect: ~10% TPS gain and 60+ GPUs saved at sub-600ms p95 for thousands of QPS. Interview angle. "What’s the 80%/4% number?" → naive KV allocation wastes up to ~80% to fragmentation; PagedAttention holds it under 4%. That pair is the single most-cited inference number in the question banks.
The tradeoff is honest: paging forces the attention kernel to gather scattered physical pages into a contiguous workspace (which FlashAttention’s tile model accommodates well), and prefix caching adds a small per-token hash-table cost paid in TTFT. The operational rule: enable prefix caching by default for any workload with >5% prefix overlap (RAG, agents, chatbots with long instructions); if every request has an unrelated context (single-shot open-user completion) the hit rate is low and the complexity isn’t justified. Per-user caches are not shareable across users for privacy reasons — only shared system prompts and public documents get prefix-cached.
Continuous batching — the largest system-level win
The single biggest serving win is a scheduling change, not a kernel. Static (request-level) batching waits for a fixed batch to fill, then runs every request to completion before swapping in the next — so the whole batch is padded to the longest sequence and short requests sit idle burning GPU. Under length variance, Anyscale measured static-batching throughput collapsing to ~81 tokens/s. Continuous (iteration-level) batching (Orca, Yu et al., OSDI 2022) reschedules after every forward pass: finished requests are evicted and new ones admitted in their slots, keeping the GPU full. Orca measured a 36.9x throughput improvement at the same latency versus FasterTransformer on GPT-3 175B; Anyscale’s vLLM reproduction reports 23x throughput while reducing p50 latency on Llama-13B, saturating near 1900 tokens/s at QPS≈8.
code
1BATCHING (the schedule is the bottleneck, not the kernel)23 STATIC (request-level): fill batch -> run ALL to completion -> swap4 pad to longest seq; short requests idle; ~81 tok/s under length variance56 CONTINUOUS (iteration-level, Orca/vLLM): reschedule EVERY forward pass7 evict finished, admit new into freed slots -> GPU stays full8 Orca: 36.9x vs FasterTransformer (GPT-3 175B) at SAME latency9 vLLM/Anyscale: 23x vs naive + lower p50; ~1900 tok/s @ QPS~8 (Llama-13B)1011 Cost: prefill (big) starves decode (small) in a shared batch -> TTFT spikes.12 Fix: CHUNKED PREFILL (split long prompt into ~512-tok chunks) or disaggregation.
Continuous batching is now table-stakes — static-batching implementations are obsolete for serving. But it shifts the bottleneck from "GPU idle" to "scheduler quality," and introduces a new failure: a long prefill (large, compute-heavy) shares the batch with many tiny decode steps and starves them, producing TTFT spikes correlated with prompt-length outliers. The first mitigation is chunked prefill — split a long prompt into ~512-token chunks so decode iterations interleave. The full fix is prefill-decode disaggregation: run prefill and decode on separate engine pools connected by a KV-transfer link, so prefill bursts can’t perturb decode latency and the two scale independently. Interview angle. "vLLM claims 23x — what bottleneck does it remove?" → name static-vs-continuous, the padding/idle waste, then layer PagedAttention (the 23x figure includes both) and quote a throughput number; "vLLM is fast" is the weak answer.
Disaggregation is the new SLO frontier, and the data shows it is not uniformly better — it’s a function of which SLO term binds. The TaiChi study (Llama-2-70B) measured: under relaxed TTFT (16s) + tight TPOT (60ms), disaggregation hits 98% SLO attainment vs aggregation’s 7%; but under tight TTFT (5s) + relaxed TPOT (250ms), aggregation wins 97% vs disaggregation’s 42%. Disaggregation wins when TPOT is the binding constraint (prefill can’t starve decoders); aggregation wins when TTFT binds (shared compute pool). Adopt disaggregation only when your SLO is asymmetric and prompt lengths are bimodal (long RAG + short chat) and you have NVLink/RDMA between pools — otherwise use chunked prefill.
Speculative decoding — 2–3x single-stream latency
Speculative decoding attacks the per-token bandwidth tax of decode. A small draft model cheaply proposes K tokens, and the large target verifies all K in a single forward pass via rejection sampling — and the output distribution is mathematically identical to vanilla decoding (this is the property to state; it’s lossless). The speedup is roughly K × α where α is the acceptance rate; with K=4–8 and α=0.5–0.8 you get 2–3x. Leviathan et al. (ICML 2023) measured on T5-XXL/TPUv4/batch=1: 2.6x@temp=1, 3.4x@temp=0 on translation; 2.3x/3.1x on summarisation. EAGLE removes the separate draft model by training a single draft head inside the target’s forward pass; Meta’s Llama-4 production deployment reports 1.4–2x at large batch and 10–30% over the vLLM baseline at single-stream, with a 2.6x compound when guided/JSON decoding is combined with speculation.
code
1SPECULATIVE DECODING (draft proposes K, target verifies in 1 pass)23 speedup ~= K * alpha (K = draft tokens, alpha = acceptance rate)4 K=4-8, alpha=0.5-0.8 -> 2-3x; output distribution IS IDENTICAL (lossless)56 Leviathan (T5-XXL, bs=1): 2.6x @T=1 / 3.4x @T=0 (translation)7 EAGLE (Meta, Llama-4 prod): 1.4-2x large batch; 10-30% over vLLM single-stream8 EAGLE 3.1 (vLLM, GB200): 2.03x @C=1 -> 1.71x @C=4 -> 1.66x @C=16910 WHEN IT EVAPORATES: at large batch the target is already throughput-bound, so11 draft compute becomes overhead. Best for bs<=8 templated traffic (chat/RAG/code).12 Acceptance rate is workload-dependent -> MEASURE it; stable prefixes accept more.
The crucial regime fact, and the thing that makes this the most under-deployed latency lever: speculative decoding helps when the target is bandwidth-bound (small batch) and stops helping — even hurts — when the target is already throughput-bound (large batch), because the draft compute becomes pure overhead. EAGLE 3.1 in vLLM shows the decay cleanly: 2.03x at concurrency 1, 1.71x at C=4, 1.66x at C=16. So enable it on latency-sensitive batch-1 traffic (chat, completions, code edit); at batch ≥8 on memory-bound decode, measure with your own target metric before turning it on. And because the benefit is K × α, the acceptance rate α must be measured per workload — templated prompts with long stable prefixes (chat, RAG) accept more draft tokens than open-ended generation. Interview angle. "What is speculative decoding and when does it stop helping?" → lossless draft-and-verify for 2–3x at bs=1; evaporates at high batch when the target is throughput-bound — stating the regime is the senior signal.
Serving stacks — workload-driven, not vendor-driven
The stack choice is dictated by load shape, not vendor prestige. vLLM (Berkeley) composes PagedAttention + continuous batching + chunked prefill + prefix caching + the broadest quantization support + first-class speculative decoding — the greenfield default. TGI (HuggingFace) wins TTFT at low concurrency (less per-batch scheduler overhead). SGLang (LMSYS) adds RadixAttention + a compressed finite-state-machine JSON decoder + cache-aware scheduling — built for multi-call structured-generation programs with heavy prefix reuse. TensorRT-LLM (NVIDIA) is the peak-throughput story on H100/H200 via FP8/SmoothQuant kernels and in-flight batching, at the cost of 30–60-minute per-(model,GPU,batch) engine builds. The November-2025 head-to-head: vLLM beats TGI 3.67x on Llama-2-7B at 100 concurrent (15,243 vs 4,156 tok/s), widening to 24x at 200; but TGI wins TTFT by 1.3–2x at low concurrency.
code
1STACK BY WORKLOAD (compose levers against the binding SLO)23 Workload profile -> Stack Why4 --------------------------------------------- -------- -----------------------5 High-throughput chat, 50+ concurrency -> vLLM 3-24x TGI; paged KV; quant6 Short-batch, low-conc, TTFT-bound -> TGI 1.3-2x lower TTFT at low C7 Multi-call structured (agentic/RAG/JSON) -> SGLang RadixAttention; FSM decode8 Peak throughput, hot model, ops can build -> TRT-LLM 10,000+ tok/s; FP8/MoE910 vLLM vs TGI (Llama-2-7B): 15,243 vs 4,156 tok/s @100 conc (3.67x), 24x @200.11 SGLang multi-call: up to 5x vs Guidance/vLLM (Llama-7B); 74.1% prod hit rate.12 Adoption order: (1) cont. batch + paged KV (2) FP8/int4 (3) prefix cache13 (4) spec-decode @bs=1 (5) chunked prefill (6) disaggregation.
Production numbers — what to actually budget against
Senior credibility is quoting numbers that triangulate — where the paper result and the production result agree. The pattern across documented deployments: lab speedups (2–4x PagedAttention, 36.9x Orca, 2–3x speculative) translate into single-digit-to-low-double-digit percent wins on real traffic, because real requests are dominated by scheduler, network, and kernel-launch overhead rather than the isolated kernel speedup. So you budget on the measured compound, not the corner-case ceiling. The numbers below are the ones the research surfaced from named operators — memorise a couple as evidence that your budget isn’t speculative.
code
1ENGINEERING-VALIDATED PRODUCTION NUMBERS (lab <-> prod triangulate)23 Operator Stack Workload Measured impact4 -------- ----------------------- --------------- --------------------------5 LinkedIn vLLM (paged KV + CB) thousands of QPS ~10% TPS, 60+ GPUs saved,6 sub-600ms p957 NVIDIA TRT-LLM + FP8 + FA Mixtral 8x7B/2xH100 38.4 req/s, 60 tok/s/user,8 ~50% more tput under 0.5s SLO9 Meta EAGLE spec-decode Llama-4 / 8xH100 4 ms/token, 10-30% over vLLM10 Baseten TRT-LLM + FP8 Mistral-7B bs=128 +33% tok/s, -24% $/Mtok11 LMSYS SGLang RadixAttention Vicuna-33B prod 74.1% hit rate, 1.7x TTFT1213 Plan against these (percent-level), not the 24-32x theoretical compound.
The deepest synthesis: the levers compose multiplicatively, not additively — PagedAttention (2–4x) × continuous batching (~2x over naive) × FP8 (~1.5x) × speculative decoding (2x at bs=1) is a ~24–32x theoretical ceiling, but real production measures land at single-digit-to-low-double-digit percent (LinkedIn’s ~10% TPS, 60+ GPUs) because the denominator on real traffic is request-scheduler overhead, network jitter, and kernel-launch latency — not the kernel speedup in isolation. So the lab numbers (2–4x, 23–36x) are real component wins, but you always measure the compound on your own traffic. Interview angle. "Design inference for 10M req/day at $0.001/req" → that’s $10K/day, so each request must cost ~1 GPU-second on H100: name a router (queue-depth aware), continuous batching, prefix caching, FP8 (or AWQ-int4 on A100), spec-decoding for the tail, autoscaling on QPS, and admission control for prompts that exceed the KV budget — the cost arithmetic plus the named levers is what scores.
You must deploy a 70B model on 8×A100 40GB (320GB total). Walk through the precision choice. What’s the strongest plan?
AFP8 — it’s the most cost-efficient precision and halves memoryBAWQ or GPTQ int4 (~35GB weights) — fits comfortably with KV-cache headroom; FP8 is off the table on A100, and int4 holds <0.2 ppl delta on a 70BCKeep FP16 and shard across all 8 GPUs
Your chatbot shares a long system prompt across most requests and TTFT is high under load. Which lever gives the biggest, most direct win?
APrefix caching on a paged KV store — the shared prompt’s KV is computed once and reused, so each request only pays for its deltaBSpeculative decoding to generate tokens fasterCSwitch to int3 quantization to free memory
You enabled speculative decoding and saw a nice latency win at low traffic, but at peak (high concurrency) it got slightly slower. Why?
AThe draft model’s acceptance rate drops to zero at high loadBSpeculative decoding silently changes the output distribution at scaleCAt high batch the target is already throughput-bound, so the extra draft compute becomes overhead rather than savings — spec-decoding helps when decode is bandwidth-bound (small batch)
An interviewer asks why vLLM reports ~23x over naive batching. What’s the strongest mechanism-level answer?
AIt batches more requests together onto the GPUBStatic batching pads to the longest sequence and runs the whole batch to completion (short requests idle); continuous batching reschedules every forward pass, evicting finished requests and admitting new ones — and PagedAttention removes the 60–80% KV waste, so the 23x figure is both togetherCIt uses int4 quantization to run faster
Your traffic is mostly multi-call agentic programs that reuse the same prompts in slightly varying forms and emit strict JSON. Which serving stack fits best?
ATensorRT-LLM, because it has the highest peak throughputBTGI, because it has the lowest TTFTCSGLang — RadixAttention exploits the heavy shared-prefix reuse (50–99% hit rates) and its compressed-FSM decoder handles JSON/grammar; it’s up to 5x faster than vLLM on multi-call Llama-7B
Serving is ~40% of an LLM-engineer loop now, and it rewards a systems mindset: name the binding constraint (TTFT vs TPOT vs throughput vs memory), pick the lever that attacks that term, quote a measured number, and state the regime where the lever stops helping. The recurring weak answer is "vLLM is fast" or listing every quantization format without a hardware-aware decision. Lead with the cost equation (prefill compute-bound, decode bandwidth-bound) and reason down from it.
01“Quantize a 70B on A100 40GB?” → reason from memory+hardware: FP16≈140GB no-fit, int8≈70GB, int4 (AWQ/GPTQ)≈35GB fits; FP8 needs H100/Blackwell, not A100.
02“GPTQ vs AWQ?” → both ~0.1–0.3 ppl at int4; AWQ protects ~1% salient channels + fused kernel; choose by your stack’s kernel support, not accuracy.
03“What’s the 80%/4% number?” → naive contiguous KV wastes 60–80% to fragmentation; PagedAttention pages it to <4% waste, enabling far higher concurrency.
04“Why does vLLM hit 23x?” → static→continuous (iteration-level) batching removes padding/idle; PagedAttention removes KV waste; the figure is both together.
05“What is speculative decoding / when does it stop?” → lossless draft-and-verify, 2–3x (K×α) at bs≤8; evaporates at high batch when the target is throughput-bound.
06“PagedAttention tradeoff?” → kernel must gather scattered pages; prefix caching costs a small per-token hash in TTFT; only pays off with >5% prefix overlap.
08“FP8 impact?” → ~+33% tok/s, −24% $/Mtok, near-zero ppl on H100; ~50% more throughput under a 0.5s budget (NVIDIA Mixtral) — the default on Hopper/Blackwell.
Going deeper, the follow-ups that separate offers: “KV cache size math?” (~1MB/token on a 13B; after weights, an A100 fits ~14k tokens → ~7 seqs at 2048 ctx — name the arithmetic); “per-channel vs per-tensor quantization?” (per-channel gives each output channel its own scale/zero-point, slashing error vs one per-tensor scalar; AWQ adds activation-aware scaling on top); “design a $0.001/req platform for 10M/day?” (= $10K/day → ~1 GPU-sec/req on H100: router, continuous batching, prefix cache, FP8/int4, spec-decode tail, autoscale, admission control); and “how do the levers compose?” (multiplicatively in theory ~24–32x, but real traffic lands at single-to-low-double-digit percent because scheduler/network overhead dominates — always measure the compound).
Could you choose a quantization method from hardware + memory, explain PagedAttention + continuous batching with their numbers, and say when speculative decoding and disaggregation stop helping?
New to itGetting thereConfident
Takeaways
Cost equation: prefill is compute-bound (TTFT), decode is memory-bandwidth-bound (TPOT); each lever attacks one term.
Quantization quality is a function of method, not bits: GPTQ-int4 +0.15 ppl on 175B, FP8 ~free; int3 (+0.54) is the must-fit-one-GPU floor. FP8 needs H100/Blackwell.
PagedAttention pages the KV cache (4–16 tok pages) → <4% waste vs 60–80%; prefix caching reuses shared prompts (2–4x throughput).
Continuous batching reschedules every forward pass → 23–36x over static; tradeoff is prefill starving decode (fix: chunked prefill / disaggregation).
Speculative decoding is lossless 2–3x at bs≤8 but evaporates at high batch (target throughput-bound); measure acceptance rate per workload.
Stack by load shape: vLLM (default/throughput), TGI (low-conc TTFT), SGLang (multi-call/JSON), TensorRT-LLM (peak). Levers compose multiplicatively; measure the compound.
Next: the capstone — LoRA-tune a small model and PROVE the gain with a proper eval comparison, not vibes.