Lesson 3 of 4 · 25 min

Scale using queue behavior and warm capacity

Choose scaling signals that predict user-visible overload.

Mechanism and reasoning

GPU utilization reports device activity. It does not directly say whether requests meet their deadlines. A device can be busy with a healthy batch, or busy while a long queue grows. It can also be underused because tokenization, routing or input transfer is slow. Scaling solely from utilization treats these different situations as the same event.
Track arrival rate, completion rate, queue age, first-token latency and available token capacity. Queue age often expresses user harm more directly than queue length. Twenty long requests can represent more work than two hundred tiny requests. Label traffic classes so one class's failure does not disappear in a global average.
Startup delay is part of the scaling equation. A replica may need to download weights, allocate cache, compile kernels and pass a warm-up request. Counting it as capacity before those steps finish creates a false recovery signal. Keep a small measured warm reserve when the service cannot wait for cold capacity.
Use hysteresis to avoid oscillation. Scale-out and scale-in need different evidence and timing. Drain active generation before removing a replica where possible. A cooldown does not fix a poor signal, but it can stop noise from causing repeated expensive starts.
Set a maximum scale and a budget. Beyond that point the service must degrade or reject according to a defined rule. Unlimited scale can turn a traffic bug or retry storm into a large bill. Explain the stop condition as carefully as the growth condition. A practical interview answer links the scaling metric to a user deadline and then checks whether the control can act in time.

Calculate the recovery path, not only the target count

A scaling controller reacts after observing work. Observation delay, decision delay, provisioning, model loading and warm-up all contribute to the time before new capacity serves requests. Add those intervals explicitly. If observation takes thirty seconds and startup takes ninety seconds, a decision based on the latest sample can be two minutes behind the start of a burst. A five-second request deadline cannot be protected by that action alone.
code
1Illustrative control timeline200:00 arrival rate increases from 60 to 100 requests/minute300:30 queue-age window detects overload400:35 controller requests two more replicas501:35 processes start, but weights and kernels are not ready602:05 warm-up and readiness checks pass702:05 router begins using the new replicas
During the first 125 seconds, the original service rate remains sixty per minute. At a constant deficit of forty per minute, backlog grows by about 83 requests if nothing expires or is rejected. This arithmetic is approximate because real arrivals and service times vary, but it shows the scale of the problem. A controller that reports four desired replicas at 00:35 has not yet supplied four ready replicas.
When capacity arrives, calculate the surplus. If the new total is 120 per minute and arrivals remain 100, the ideal drain rate is 20 per minute. An 83-request backlog takes about 4.15 minutes to clear, even after warm-up. A readiness graph can look healthy long before the user-facing queue recovers. Deadline expiration can shorten the queue by removing useless work, but that is a failed service outcome and must remain visible in reporting.
StateReady service rateArrival rateQueue change
Before burst60/min60/minApproximately zero
Waiting for new replicas60/min100/min+40/min
Four replicas ready120/min100/min−20/min
One replica removed too early90/min100/min+10/min
The last row illustrates a scale-in failure. A utilization drop can occur while backlog is still being processed, depending on request mix and bottlenecks. Removing capacity before checking queue age and incoming work can restart the incident. Use a scale-in policy that considers sustained low demand, remaining queued work, drain time and the cost of starting the replica again. Hysteresis should be justified by measured delays and variability.
A warm reserve is one response to startup delay, but its cost is real. If expected bursts require two immediately ready replicas, compare the reserve cost with the product's allowed rejection or wait policy. Some bulk tasks can accept a longer queue and use cheaper delayed capacity. Interactive requests may need a protected warm pool. Separating these classes can be more effective than applying one global queue target to both.
Signals must identify the bottleneck. High queue age with low GPU activity can indicate slow tokenization, input fetching, a blocked router or unavailable execution slots. Adding GPUs may leave that bottleneck unchanged. Compare arrival and admission timestamps, CPU utilization, preprocessing duration and GPU execution metrics. A trace showing long time before the model receives work supports a different intervention from a trace showing a fully occupied decode scheduler.
Set an upper bound on scale and a response beyond it. A runaway client retry policy can generate more offered work as responses slow. Capacity growth alone may make that feedback expensive without restoring service. Rate limits, retry guidance and request identity can reduce amplification. The controller should expose why it cannot add more capacity, such as quota, budget or resource availability, so operators do not mistake a desired count for an achievable plan.
For verification, inject a known burst, measure every delay in the timeline and check the oldest useful request after capacity arrives. Then repeat with a failed replica and a cold cache. The purpose is to test the control loop against the service deadline. A single steady-state throughput number cannot answer whether the loop acts soon enough.

Worked example

Invented trace:
MinuteArrivals/minCompletions/minOldest queue ageReady replicas
060600 s2
11006020 s2
21006040 s2
310012030 s4
New replicas take two minutes to become ready. Scaling at minute one cannot save requests with a five-second deadline. Admit only useful work, maintain warm reserve for expected bursts, and use the trace to revise the policy. Minute three shows backlog draining, not instant recovery.

Exercise

Two replicas each serve thirty requests per minute. Demand rises to ninety. A third replica needs three minutes to start. With no rejection, calculate accumulated backlog before it is ready and explain whether it then drains.

Model solution and rubric

The deficit is thirty per minute for three minutes, giving ninety queued requests. Once ready, total service rate equals ninety per minute, so the existing backlog does not drain under constant demand. A fourth replica or reduced admission is needed to recover. Verify queue age after scale-out, not merely the ready count. The calculation assumes constant rates and no request expiry.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“Matching arrivals clears the backlog.” Equal rates stop further growth; existing queued work needs surplus capacity or explicit expiration.
“Low GPU utilization means excess serving capacity.” Another stage can starve the GPU while users wait. Find where time accumulates before scaling in.

Interview probe

Evidence class: recommended. Original practice.
Why did adding one replica stop queue growth but fail to restore latency?
Strong answer: The new capacity only matched incoming work. No spare capacity remained to drain the accumulated queue. I would estimate recovery time using service rate minus arrival rate and expire work that has missed its useful deadline.
Follow-up: How would you distinguish a tokenizer bottleneck from exhausted GPU capacity?
Weak answer indicators: Treating desired replicas as ready replicas; ignoring model warm-up; assuming equal arrival and service rates drain old work.

Sources

Technical references: AWS Builders’ Library on queue backlogs; vLLM metrics design; Kubernetes probes; Google SRE alerting on SLOs. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsAWS Builders’ Library on queue backlogsd1.awsstatic.comdocsvLLM metrics designdocs.vllm.aidocsKubernetes probeskubernetes.iodocsGoogle SRE alerting on SLOssre.google

Checkpoint

A 90-request backlog remains when capacity equals arrivals at 90/min. What happens under constant rates?

AIt drains in one minuteBIt doubles each minuteCIt remains approximately constantDIt disappears at readiness
Sign up free to answer and see why

Checkpoint

Queue age is high but GPUs receive little work; preprocessing duration rose. Which first action fits?

AAdd GPUs solely from queue lengthBScale in because GPUs are idleCIncrease decode batch size before tracing inputDInvestigate the preprocessing bottleneck
Sign up free to answer and see why

Checkpoint

Arrivals are 100/min, capacity 130/min and backlog 150. Ideal drain time?

A5 minutesB1.15 minutesC1.5 minutesD30 minutes
Sign up free to answer and see why

Checkpoint

A replica process starts but weights and warm-up are incomplete. How should the router treat it?

AImmediately ready because the process existsBUnavailable for normal workload until readiness passesCHalf capacity automaticallyDReady only for the largest prompts
Sign up free to answer and see why

Checkpoint

The service reaches its budget cap during a retry storm. What policy is required?

AUnbounded queued retries until resources returnBRaise the cap without checking the causeCBound admission and expose the overload outcome while controlling retry amplificationDRemove timeout metrics from the scaling signal
Sign up free to answer and see why

Explain how you would choose scaling signals that predict user-visible overload without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Scale from measured work and deadlines. Check surplus capacity and readiness before declaring recovery.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.