Lesson 1 of 4 · 25 min

Build a GPU memory budget

Calculate a defensible inference memory budget.

Mechanism and reasoning

A model can load successfully and still fail when users arrive. Weight memory is only the fixed part of the allocation. Each active sequence also needs attention state, often called the key-value cache. The runtime needs temporary buffers, compiled graph storage and allocator space. An estimate that uses parameter count alone answers whether weights might fit. It does not establish a safe serving limit.
For a conventional full-attention model, a useful first approximation for cache bytes is two times layer count times stored tokens times key-value head count times head dimension times bytes per value. The factor two represents keys and values. Use key-value heads, not query heads, for grouped-query attention. State whether stored tokens includes both input and output. Architectures with sliding windows, latent attention or recurrent state need a different calculation.
Treat this calculation as a bound for planning. Runtimes allocate cache in blocks, and different requests finish at different lengths. Fragmentation, prefix reuse and temporary peaks change actual use. Begin with an explicit reserve, then measure the largest supported workload. A reservation is not wasted capacity if it prevents an out-of-memory failure during a traffic burst.
An interview answer should expose units. Decimal gigabytes and binary gibibytes differ. A calculation that mixes them can appear precise while hiding a material error. Write the formula, substitute the model's actual configuration, and explain what you excluded. Do not invent a maximum concurrency from the advertised GPU capacity.

Capacity worksheet and admission boundary

Use a worksheet that separates the model architecture from the traffic promise. The architecture determines bytes for one stored token. The request contract determines how many tokens can be stored. The scheduler determines how much of that state is resident together. These are three different inputs. A configuration file that permits a long context does not prove that all admitted requests can reach that context together.
Budget itemStated assumptionPlanned allocation
Device memoryBinary units24 GiB
Weights7 billion values at 2 bytes13.04 GiB
Runtime reserveMeasured later, provisional now3 GiB
Available cacheDevice minus weights minus reserve7.96 GiB
KV per stored tokenStandard attention, 32 layers, 8 KV heads128 KiB
Cache at 4,096 tokensAll layers included0.5 GiB
Suppose the API permits 3,072 input tokens and up to 1,024 new tokens. Reserving 4,096 tokens per admitted request makes the fifteen-request bound conservative with respect to that declared maximum. If the API instead permits 4,096 input tokens plus 1,024 new tokens, the budget has changed. Each maximum request needs 5,120 × 128 KiB = 640 MiB, or 0.625 GiB. The simplified bound becomes floor(7.96 / 0.625) = twelve. The API field called context length must be interpreted correctly before admission uses it.
An alternative scheduler can reserve less than each request's maximum and allocate blocks as generation progresses. That can increase utilization when most requests finish early. It also creates a policy obligation when several requests grow together. Possible responses include preemption, recomputation, queuing additional work, or an explicit rejection before admission. None is free. Explain which latency objective you protect and what happens to already admitted requests when the free block pool becomes small.
code
1Illustrative admission ledger, 128 KiB per resident token2Request     Input tokens   Maximum new tokens   Reserved total3A                 2,048                1,024            3,0724B                 3,072                1,024            4,0965C                 1,024                2,048            3,0726TOTAL                                                   10,2407Cache reservation = 10,240 × 128 KiB = 1.25 GiB
This ledger also helps diagnose an out-of-memory event. Capture resident cache blocks, active requests, input and generated lengths, and the runtime allocation peak at the same instant. A dashboard with a one-minute average can miss a short allocation peak. A prompt-count graph cannot explain why one request used ten times the state of another. The investigation needs a timeline that connects admission to memory use.
Do not subtract a displayed weight footprint from nominal memory and treat the difference as usable cache automatically. Loading can use temporary copies; execution may initialize graph or workspace allocations later; allocator reservation differs from live tensor allocation. Conversely, the runtime may preallocate most of the cache pool, so a high displayed reservation does not mean every cache block is occupied. Read the engine's metrics and version-specific allocation behavior before declaring a leak.
The standard formula is intentionally limited. It does not directly describe every attention architecture, cache precision, offloading policy or sharing mechanism. The cited vLLM implementation article gives a per-layer, per-block formula for standard non-MLA attention. Multiplying by layers and resident token blocks produces this worksheet under its stated assumptions. Do not apply it unchanged to latent-attention models. In an interview, identifying that missing architecture detail is a better answer than inventing a precise total.
Finally, establish two separate test boundaries. A memory test should combine the longest permitted inputs with the longest permitted outputs and enough simultaneous work to exercise the planned reservation. A service test should replay a realistic arrival mix and check latency, rejection and cancellation. Passing the first establishes a narrower claim about allocation under that case. Passing the second establishes behavior for that workload and duration. Neither proves unlimited contexts or traffic.

Worked example

Teaching example: a 7-billion-parameter model uses two bytes per weight, so weights need 14 billion bytes, about 13.04 GiB. It has 32 layers, eight KV heads, head dimension 128 and two-byte cache values. Cache per token is 2 × 32 × 8 × 128 × 2 = 131,072 bytes, or 128 KiB. At 4,096 stored tokens, one sequence needs 0.5 GiB. On a 24 GiB device, reserve 3 GiB for runtime needs. The arithmetic leaves 7.96 GiB for cache, so fifteen such sequences fit this simplified budget. Fifteen is a starting test point, not a service promise.

Exercise

Use the same model and reserve on a 32 GiB device. Each request can reach 8,192 stored tokens. Calculate the simplified sequence limit and name one reason to test below it.

Model solution and rubric

Each sequence now needs one GiB. Remaining memory is 32 − 13.04 − 3 = 15.96 GiB, so the integer limit is fifteen sequences. Start below fifteen because temporary allocations and block overhead may exceed the reserve. Verify peak allocated and reserved memory using maximum input and output lengths, not only average prompts. A test that succeeds at short outputs cannot validate the stated limit.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“Weight quantization reduces every allocation by the same factor.” Cache precision and workspace behavior can remain unchanged. Recalculate each component separately and verify the runtime configuration.
“Fifteen sequences fit, so fifteen users are safe.” A user can have multiple active requests, and permitted token lengths can exceed the worksheet assumptions. Enforce the request contract and test peak state, not just concurrent user count.

Interview probe

Evidence class: recommended. Original practice.
Why did the server fail at twelve users when the model occupied only half of memory?
Strong answer: Concurrent sequence state and runtime allocations can consume the remainder. I would inspect stored token lengths, cache occupancy and temporary allocation peaks. User count is an incomplete load measure because one long conversation can retain much more state than several short requests.
Follow-up: What changes if the model uses four KV heads rather than eight?
Weak answer indicators: Equating weight size with total memory; using query heads without checking the architecture; describing the estimate as benchmark evidence.

Sources

Technical references: vLLM parallelism; vLLM metrics design; Inside vLLM, dated implementation anatomy. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsvLLM parallelismdocs.vllm.aidocsvLLM metrics designdocs.vllm.aidocsInside vLLM, dated implementation anatomyvllm.ai

Checkpoint

The API permits 4,096 input tokens and 1,024 new tokens. Which reservation matches a maximum request?

A4,096 tokens because only input is cachedB1,024 tokens because decode replaces input stateC5,120 tokens under the stated full-attention modelD4,096 tokens because generation shares the same positions
Sign up free to answer and see why

Checkpoint

Weight precision changes from two bytes to one, but cache precision stays two. What changes in this worksheet?

AWeight bytes halve; KV bytes per token remain fixedBBoth weights and cache halveCOnly runtime reserve halvesDConcurrency exactly doubles
Sign up free to answer and see why

Checkpoint

A server reserves its cache pool at startup. High reserved memory with few occupied blocks proves what?

AThe server has leaked occupied KV blocksBEvery admitted request is at maximum lengthCNew requests must fail immediatelyDReservation alone cannot establish live cache occupancy
Sign up free to answer and see why

Checkpoint

There are 6 GiB available for cache and each maximum request needs 0.75 GiB. What is the simplified integer limit?

A6 requestsB8 requestsC9 requestsD12 requests
Sign up free to answer and see why

Checkpoint

A latent-attention model replaces the standard attention model. What is the first defensible budget step?

AKeep the old per-token formula because layer count is unchangedBScale the old cache solely by parameter countCInspect the architecture and engine's actual cache state layoutDAssume cache memory is zero because fewer states are stored
Sign up free to answer and see why

Explain how you would calculate a defensible inference memory budget without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Keep the fixed and per-request allocations separate. Carry units through every step and validate the longest supported request.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.