During autoregressive decoding, recomputing all key and value projections for the entire prefix at every new token is wasteful because past tokens never change. KV cache stores those K and V tensors from prior steps, so each new token only needs to compute its own Q, K, V; the attention output then concatenates to growing KV storage rather than recomputing history. The per-token compute drops from O(seq_len) to O(1), giving 10-100x decoding speedups. The cost is memory: ~2 * n_layers * n_heads * seq_len * head_dim * 2 bytes (fp16, K and V). For a 70B model with 80 layers, hidden=8192, seq=8k tokens, that's ~80 GB, often exceeding H100 memory. This is why techniques like GQA/MQA (sharing KV across heads), Multi-Head Latent Attention (low-rank KV compression), paged KV cache (vLLM style), and KV cache quantization (FP8 KV) are hot topics. Senior candidates should mention the compute-memory trade-off and FlashAttention-style recomputation as the alternative when memory is tight but FLOPs are cheap.