LLM decoding is iterative (token-by-token) and requests share structure: they all run the same model, just with different contexts. Static batching wastes GPU on padding. Dynamic (in-flight) batching: each decode step, walk the queue of active requests, batch together the sequences that have not finished, run a GPU forward pass on the union. Continuous (in-flight) batching further refines: completed sequences can exit the batch mid-step and new ones can join. Stall-free batching (e.g., vLLM's iteration-level scheduling) avoids head-of-line blocking by separating prefill (compute-bound) and decode (memory-bound): new requests' prefill chunks can interleave with ongoing decodes. Senior nuance: prefill is latency-sensitive (TTFT: time to first token), decode needs steady TPOT (time per output token); the right scheduler balances both. FairBatching-style work prevents large requests from monopolizing the GPU.