Distinguish prompt processing from output generation.
Mechanism and reasoning
An autoregressive request has two different workloads. Prefill processes the available prompt and prepares attention state. Decode repeatedly produces another token using prior state. Both matter to the user, but they create different pressure on the device. A large prompt can delay the first visible token even when later tokens arrive smoothly. A large output can hold capacity for a long time after a quick first token.
Time to first token includes waiting, prompt preparation, prefill and any service overhead before the first output reaches the client. Inter-token latency measures the gaps that follow. State the observation point and unit: a streamed output event can contain several tokens, so an event-gap histogram is not automatically a per-token timing distribution. The worksheet assumes one visible token per stated gap. End-to-end duration includes both. A single average latency hides these differences and can lead to the wrong fix.
Batching combines work from several requests. A static batch waits for the batch to complete before moving on. Continuous batching can admit and remove requests as generation progresses. This reduces idle capacity caused by different output lengths, but it does not remove scheduling trade-offs. A scheduler that admits every long prompt immediately can delay active decodes. Chunking prompt work creates opportunities to alternate work, with additional scheduling overhead.
Choose a benchmark with the traffic shape you expect. Fix input lengths, output lengths, arrival pattern and stopping rules. Record rejected and cancelled requests too. Otherwise, a run that silently drops hard work may look faster. In an interview, first decide which user experience must improve. A throughput increase is not sufficient evidence that an interactive service became better.
Read a mixed-workload trace
Consider two traffic classes on the same replica. Short chat uses 256 input tokens and 128 output tokens. Document requests use 8,192 input tokens and 64 output tokens. These invented classes have different input-to-output ratios. An average prompt length of 4,224 describes neither class well. Keep separate latency distributions so that a change which harms one group cannot hide behind the aggregate.
Trial
Short-chat p95 first token
Short-chat p95 token gap
Document p95 first token
Total output rate
Baseline
500 ms
35 ms
1,900 ms
900 tokens/s
Larger prompt batch
900 ms
70 ms
1,500 ms
1,080 tokens/s
Bounded prompt chunks
550 ms
40 ms
1,750 ms
1,010 tokens/s
Assume the short-chat limits are 700 ms first token and 50 ms token gap, while documents allow 2,000 ms first token. The larger batch improves aggregate throughput by 20 percent but fails both chat targets. Bounded chunks meet these stated percentile limits in this trial and improve output rate by about 12.2 percent. This is a candidate for a longer test, not proof that the same limits hold during a burst or after losing a replica.
The mechanism matters. Prefill can contain substantial parallel work across prompt tokens. Decode repeatedly advances active sequences, and the effective bottleneck depends on batch size, hardware, architecture and implementation. Avoid declaring that prefill is always compute-bound or decode is always memory-bound. Those descriptions are useful tendencies under some conditions, not evidence from this trace. Inspect utilization, memory traffic and execution timing if the question asks which resource limits the system.
code
1Illustrative scheduler trace, not a specific engine algorithm200 ms: decode step for active chat requests308 ms: process a bounded chunk of document prompt428 ms: decode step for active chat requests536 ms: process the next document chunk656 ms: decode step for active chat requests
The trace shows why chunking can reduce the period during which prompt work prevents decode progress. It does not prove that the chunk duration is stable. A longer chunk may improve kernel efficiency but increase the gap between decode opportunities. A shorter chunk can add scheduling overhead. A real measurement must include the work of moving between these tasks, cache allocation pressure and the number of active sequences. A token limit is a control input; observed latency is the result to verify.
Distinguish output token rate from request rate. A service generating long answers can show many tokens per second with few completed requests. A service producing short answers can complete many requests while generating fewer tokens. When comparing two settings, fix the distribution of requested lengths and stopping conditions. If one setting causes more early cancellations, its completed-request distribution can become shorter and create a false improvement.
Cancellation also changes the work accounting. Suppose a client disconnects after waiting for its first token. If the server continues generation, GPU time is spent on an output that no user receives. Count submitted, admitted, cancelled and completed work separately, and verify that cancellation reaches the scheduler. Cancelling unfinished work can increase measured useful throughput without improving the kernel. This is a system behavior change and should be explained as such.
For a practical tuning exercise, keep the model, precision, arrival trace and hardware fixed. Change only the prompt chunk limit between the first two trials. Record the short and long class distributions, maximum resident tokens, admission rejects and output length distribution. Then add a burst with the same total request mix to test whether the policy remains useful when the queue is nonempty. One controlled change gives a stronger causal explanation than simultaneously changing batching, model version and GPU type.
When presenting the result, write the decision in one sentence: “For the stated mix and targets, test bounded chunks next because they preserve both chat limits while raising useful output rate.” Then state the uncertainty: this small trace does not show p99 behavior, a cold model, or overload after a failure. This combination of a clear decision and a bounded claim is what makes a capacity answer credible.
Worked example
An invented request waits 120 ms, spends 30 ms in preprocessing and 250 ms in prefill, then returns its first token. Its first-token time is 400 ms. It generates 50 tokens in total, with 20 ms between subsequent tokens. Completion time is approximately 400 + 49 × 20 = 1,380 ms. A scheduler change reduces prefill to 180 ms but raises each decode gap to 28 ms. First-token time becomes 330 ms, while completion becomes 1,702 ms. The change helps responsiveness and hurts full-answer time.
Exercise
A second request has 100 ms of queueing, 40 ms preprocessing, 160 ms prefill and 31 output tokens at 25 ms gaps. Calculate first-token and completion time. Which metric would reveal a doubled queue?
Model solution and rubric
First-token time is 300 ms. There are thirty gaps after the first token, so completion is 1,050 ms. Doubling queueing adds 100 ms to both, while inter-token latency stays 25 ms. Measure queue duration separately to distinguish congestion from model execution. Use server timestamps and client observations to expose transport delay rather than attributing everything to the GPU.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
“Higher output rate means every request is faster.” A batch can improve device efficiency while adding wait between a request's decode steps. Check each traffic class.
“First-token time is the prefill kernel time.” Queueing, preparation and transport can contribute. Separate the stages before changing GPU scheduling.
Interview probe
Evidence class: recommended. Original practice.
Why can more batching increase tokens per second but make chat feel slower?
Strong answer: The scheduler can keep the GPU busy while making a request wait longer for its next turn. I would compare first-token and inter-token latency by prompt and output length, then tune batch token limits against the user target.
Follow-up: How would you test a mix of short chat messages and long document prompts?
Weak answer indicators: Reporting aggregate tokens per second alone; multiplying by fifty gaps for fifty tokens; assuming all GPU work has the same performance limit.
Sources
Technical references: vLLM metrics design; vLLM prefix caching. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
A larger batch passes document latency but fails chat gap limits. Which next experiment directly tests the scheduling tradeoff?
AHold workload fixed and bound prompt chunksBDouble both model size and hardware countCUse only document traffic in the next testDReport one combined latency average
A new trial has higher tokens/s but far longer requested answers. What conclusion is supported?
AIt has higher useful request capacityBIt has lower per-request latencyCIt needs a matched length distribution before a capacity comparisonDIt has eliminated scheduler overhead
Explain how you would distinguish prompt processing from output generation without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Separate waiting, prompt work and generation. Optimize the part of the timeline that causes the user problem.