Prioritize measurements and reversible mitigations during inference overload.
Mechanism and reasoning
An incident interview often supplies incomplete evidence. Resist the urge to name a root cause immediately. State the user impact, determine whether it is still growing, and choose a mitigation that limits harm while preserving evidence. Diagnosis and mitigation can proceed in parallel, but each action should have a predicted effect.
Start with the request timeline. If first-token latency rose while decode gaps stayed stable, waiting or prefill is a stronger initial hypothesis than slower generation. If both rose, device contention, longer outputs or scheduling interactions deserve attention. Compare traffic mix and recent changes before restarting everything.
A recent deployment is evidence of timing, not proof of causation. A simultaneous traffic shift can produce the same symptoms. Segment by serving version, prompt length and tenant. If only the new version fails on an unchanged slice, rollback gains support. If every version slows when one tenant submits long prompts, admission and isolation may matter more.
Choose a reversible action with a clear stop condition. Reduce bulk admission, route to a known warm version, cap output length where the product permits it, or add already available capacity. Do not apply all changes together and then claim to know which one worked. Some emergency combinations may be necessary, but record that they limit causal attribution.
Keep the recovery criterion user-facing. A falling queue depth can hide requests that expired or were dropped. Count useful completions, rejection reasons and oldest queue age. After service recovers, replay the event in a controlled environment to test the hypothesis and revise the admission or release policy.
Keep an incident notebook with separate evidence classes
Write the observations before the conclusion. A useful notebook names the metric, time window, population and source. “Latency is high” is too broad. “Client-observed first-token p95 for long prompts rose from 0.8 to 8 seconds between 10:00 and 10:02, with 2,000 requests in the slice” gives another engineer something to verify. State whether timed-out requests are included, excluded or censored.
Record type
Example
Observation
Long-prompt queue p95 rose to 6 seconds
Hypothesis
Bulk arrivals exceeded the prompt-processing budget
Action
Limit new bulk reservations for five minutes
Prediction
Queue age and chat first-token time should fall
Falsifier
Queue stays low but prefill execution remains slow
Recovery check
Useful completions, rejections and latency by class
The falsifier prevents the hypothesis from becoming a story that accepts every outcome. If the queue falls but prefill execution remains slow, admission may have relieved one problem while another persists. If short prompts recover and long prompts remain slow, the class boundary matters. Keep the result narrower than “the incident is fixed.”
code
1Illustrative request trace2arrival 10:02:00.0003reservation granted 10:02:05.9004prefill begins 10:02:06.0005first token emitted 10:02:06.7006first token received 10:02:06.7407subsequent gap p95 0.035 seconds
This individual request spends 5.9 seconds before reservation,0.1 seconds between reservation and prefill,0.7 seconds in the pre-first-token model interval, and 0.04 seconds in transport after emission. It supports investigation of admission and queueing for this request. Do not treat one trace as the population p95, and do not sum independently measured p95 components as if they belonged to the same request. Percentiles of component distributions generally do not add into an exact end-to-end percentile.
Choose a mitigation that the available evidence supports. Limiting bulk admission can protect interactive work, but it creates explicit rejections or delays for bulk users. Routing to a warm old version can address a release regression, but it can overload the old pool. Restarting every replica can clear some bad state while also discarding caches and triggering expensive cold starts. Every action has a predicted benefit and a cost; report both.
When two changes happen together, preserve the causal limit. If an emergency requires both rollback and admission reduction, the combined action can restore service without identifying the independent contribution of either change. Later replay the same trace with one factor changed at a time. Do not rewrite the incident report to credit the favored action solely because it was easy to describe.
Use version and traffic slices carefully. A new version may receive only a particular tenant or request type. Comparing its aggregate latency with the old version can confuse workload selection with performance. Match prompt length, output limit, cache condition and arrival pressure where possible. If exact matching is unavailable, say which confounders remain rather than declaring a precise regression percentage.
The recovery window should include more than a momentary drop. A queue can clear because requests expired, clients stopped retrying, or the router rejected new traffic. Track successful work that meets the user objective, the offered-request failure fraction and oldest useful queue age. If the service is stable only because half the traffic is rejected, state that the mitigation controls overload but does not restore the full service contract.
After the immediate event, convert the supported lesson into a bounded change proposal: a token-aware tenant limit, a warm-capacity requirement, a release slice check or a missing metric. Include a local replay that would have caught the failure. Avoid a long list of unrelated best practices. The best post-incident action closes the specific gap demonstrated by evidence and includes a way to verify that the gap is closed.
Worked example
Invented timeline: at 10:00, a tenant starts 16,000-token prompts. At 10:02, first-token p95 rises from 0.8 to 8 seconds. Decode gap p95 stays 35 ms. GPU activity remains high, and oldest queue age reaches six seconds. No release occurred. The first action is to bound the tenant's long-prompt admission and protect the interactive pool. Check whether first-token latency recovers while decode stays steady. This supports a queue/prefill explanation but does not identify the exact scheduler behavior without further measurement.
Exercise
A model release and traffic increase happen together. Old and new versions both slow for long prompts, but short prompts remain healthy. Name two useful comparisons and one bounded mitigation.
Model solution and rubric
Compare long-prompt latency on each version at the same arrival rate, and compare queue duration with prefill execution time. Temporarily limit long-prompt admission or move bulk work into a separate bounded pool. A blanket rollback might still be prudent if there is additional release evidence, but this trace alone does not isolate the release as the cause. Verify useful completions and queue age after the change.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
“The sum of stage p95 values is the request p95.” The slow samples can be different requests. Use joined request traces or the actual end-to-end distribution.
“A shrinking queue proves recovery.” Expiry and rejection can shrink it without serving users. Check useful completions and the offered-request outcome.
Interview probe
Evidence class: recommended. Original practice.
What is your first response to an eight-second first-token p95?
Strong answer: Confirm affected users and traffic slices, inspect queue versus execution time, then apply a bounded mitigation tied to that evidence. I would preserve request/version metadata and define recovery as restored useful latency, not simply a restarted process.
Follow-up: What would a stable queue but doubled prefill duration suggest?
Weak answer indicators: Restarting all replicas before checking warm-up cost; assuming a nearby deployment caused the incident; calling dropped work a performance improvement.
Sources
Technical references: vLLM metrics design; Google SRE alerting on SLOs. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
First-token latency rises; measured queue wait grows while matched prefill duration and streamed output gaps stay stable. Which first explanation best fits?
APrefill kernels became slower in the matched sliceBDecode execution became slower throughout the responseCMore time is spent waiting for generation to startDTransport buffering alone explains the measured server queue increase
The canary serves long documents and the baseline serves short chat. What would make the version latency comparison more informative?
AGather more baseline requests without changing either traffic mixBCompare only request counts, because each request uses one queue slotCNormalize by output token count while ignoring prompt and cache differencesDCompare matched prompt/output/cache/concurrency slices or model those confounds explicitly
Explain how you would prioritize measurements and reversible mitigations during inference overload without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Keep observations, hypotheses and actions distinct. A mitigation can work before the exact cause is known.