Speculative decoding uses a small "draft" model to generate K candidate tokens autoregressively, then the large "target" model scores them in a single forward pass. Tokens that match the target's argmax (or sampled with the adjusted distribution) are accepted; the first mismatch triggers a re-sample using the target's distribution. Effective when: (1) batch size is small (single or few concurrent requests): decoding is memory-bandwidth-bound and a draft model is cheaper to query; (2) draft acceptance rate is high (>60-70%): target and draft distributions must be close. Fails when: (1) batch size is large (~64+): the GPU is already compute-bound, and speculation adds overhead without throughput wins; (2) draft/target acceptance is low (e.g., domains where draft underperforms target): the wasted speculation costs more than it saves. Senior nuance: production systems using speculative decoding see small-batch latency wins of ~1.5-2.5x but large-batch throughput unchanged. This is the "Speculative Decoding Illusion" many candidates miss.