Requirements: 100M DAU, prompt + 50-turn chat history, median first-token latency <500ms, p99 <2s, a token-throughput target, and a target cost per 1M tokens.
Architecture: clients → CDN + edge WAF → API gateway (auth, rate-limit, abuse) → request router → inference fleet of H100/A100 pods with continuous batching (vLLM-style), paged KV cache, and prefix sharing for system prompts.
State: stateless inference pods; persistent Redis session store (chat turn IDs) and Postgres for user/org metadata; object store for attachments; vector store for retrieval-augmented memory if applicable.
Data flow: gateway → LLM-serving fleet with a load balancer using "least loaded by KV occupancy", not "least loaded by request count"; stream SSE back.
Evaluation: weekly load test at 5× median; shadow new model routes on 1% traffic; canary on cost/latency deltas; rolling deploy with auto-rollback at +5% p99 regression.
Monitoring: golden-signal dashboards (latency, traffic, errors, saturation) plus AI-specific metrics — TTFT, ITL, GPU SM occupancy, KV-cache hit ratio, drift in token throughput.
Cost/latency: prefix caching moves ~30% of repeat prompts onto cheap shared prefixes; speculative decoding cuts p50 for short prompts.
Depth signals interviewers probe for: continuous batching and paged KV; KV occupancy as the load-balance signal; tenant-aware fairness when one customer saturates the fleet.
Follow-up probes: How do you contain a single tenant consuming >5% of GPU-hours? Model-weight rollout — blue/green vs in-place re-shard? How do you detect a silent regression in quality, not just latency? How do you handle a 10× burst from a launch event?