Requirements: serve 100B+ models at <500ms TTFT, 80+ tok/s/user, with multi-tenant isolation and cost controls.
Architecture: model registry (S3/GCS); bootstrap pod pulls weights into NVMe; vLLM serving with continuous batching + paged KV + chunked prefill; multi-LoRA hot-swap; prefix cache; speculative decoding; HTTP/SSE gateway.
Placement: bin-packing on GPU SKUs (H100, A100, L40S) with a scheduler that treats KV occupancy as the load signal. Admission control based on projected KV memory (prompt length + max output); preemption policy under memory pressure (swap-to-CPU vs recompute — recompute usually wins).
Eval: golden prompts replayed nightly; quality parity vs upstream; a regression suite gates deploys.
Monitoring: TTFT, ITL, GPU SM occupancy, KV-cache hit ratio, queue depth, request admission rate — and cost per million tokens as the north-star metric.
Cost: spot when tolerable; the autoscaler scales on token throughput, not request count (target tokens/s-per-GPU).
Depth signals: continuous vs static batching math; prefix-cache amortization; speculative decoding heads; "why vLLM over TGI / TensorRT-LLM" — pick one and explain the trade.
Follow-up probes: How do you hot-swap a model without dropping traffic? How do you ensure fairness across tenants? How do you autoscale on bursty traffic without cold-start penalties? Prefill/decode disaggregation — why serve them on separate GPU pools?