Evaluate serving capacity with latency and cost evidence.
Mechanism and reasoning
A capacity experiment should answer a decision. For an interview, that decision might be whether two replicas can support a launch target. Define a successful request, the latency limit and the traffic distribution before running the test. If these definitions change after results arrive, the comparison no longer has a stable meaning.
An open-loop load generator sends arrivals according to a schedule. A closed-loop generator often waits for completion before sending the next request. During overload, the closed-loop generator can reduce its own arrival rate and hide the queue growth a real service would face. Neither approach is always wrong. They answer different questions and must be described correctly.
Averages are especially weak near saturation. Report a latency distribution and count timeouts, failures and admission rejections. Include warm and cold periods separately. A cache warmed by repeatedly sending the same prompt is not representative of mostly unique user input. Conversely, a cold-only test can miss useful prefix reuse in a stable workload.
Calculate cost using completed work that meets the objective. Price per GPU-hour divided by raw token count rewards outputs that arrived too late or failed quality constraints. A more useful unit can be cost per thousand successful requests at the specified latency target. State whether idle reserve, warm-up and failure redundancy are included.
A small test cannot establish all tail behavior. Keep the result bounded to its duration, request mix, hardware and software version. Good interview reasoning explains which uncertainty remains and proposes the cheapest next experiment that can resolve it.
Account for every arrival before comparing cost
A useful experiment produces a conservation table. Every submitted request must reach one clearly defined terminal outcome or remain explicitly pending. Without this table, a dashboard can count successful completions while silently dropping requests that timed out at the client. Use a request identifier that joins client and server observations, while excluding private prompt text from ordinary metrics.
Outcome at test cutoff
Count
Successful and within latency objective
7,000
Successful but late
1,500
Failed response
500
Rejected before execution
800
Still pending
200
Total submitted
10,000
These categories are disjoint for the example. If “failed” also includes late timeouts, do not subtract both again. The initial exercise counts responses, whereas this expanded table counts all submissions. Both can be correct if their denominators are named. With six cost units, the useful-work cost remains 6 / 7 = 0.857 per thousand timely successes. The offered-request success fraction is 7,000 / 10,000 = 70 percent. A cheap useful-work unit does not excuse an unacceptable success fraction.
Define how the cutoff handles pending work. One option stops arrivals and waits for all accepted work to finish, charging the drain interval to the run. Another reports pending requests and their age separately at a fixed deadline. Do not end the billing window while leaving expensive generation running and then report that partial cost as the full cost of the workload. Similarly, warm-up cost belongs somewhere, even if it is reported as a separate operational term.
code
1Experiment manifest, illustrative2Model revision: fixed digest M3Engine revision: fixed digest E4Arrivals: scheduled at 20 requests/s for 30 minutes5Mix: 80% short-chat, 20% document requests6Random seed: fixed for identical trace replay7Cache conditions: separate cold and representative-warm trials8Success: valid response, no transport error, meets class latency limit9Failure test: remove one replica after minute 1510Accounting: include warm-up and final drain as separate cost rows
The manifest makes reproduction possible without pretending that one seed covers all traffic. Replaying an identical trace isolates configuration changes. Repeating several seeds estimates sensitivity to request order and burst shape. Keep a record of the exact hardware, software and model configuration because kernel behavior and scheduling defaults can change across releases. When using a current documentation page, tie the actual experiment to the installed release rather than assuming today's default applied to an older run.
Open-loop arrivals are especially useful when the decision concerns an external traffic rate. If the generator sends twenty requests each second regardless of completion, queue growth exposes a service rate below twenty. A closed-loop test with twenty clients may send fewer requests per second as responses slow. That can be a realistic model of a bounded population of interactive users, but it does not demonstrate support for a fixed external arrival rate. Match the arrival model to the claimed use case.
Measure generator health too. A client process that cannot sustain its scheduled rate creates a different offered workload from the manifest. Record intended send time, actual send time and server arrival time when possible. Network saturation, connection limits or a blocked event loop in the generator can become the bottleneck. Do not interpret a flat server throughput graph as proof of GPU saturation until the input side is known to be delivering the intended load.
Tail statistics require care. A p99 from one hundred requests is dominated by very few observations. Repeated trials can reveal variability, but they do not automatically establish production reliability across days. State sample count and duration next to percentile results. If the experiment contains a long warm-up, report that interval separately rather than letting good warm-state results dilute it.
The final decision should combine useful throughput, latency by class, admitted-request failure rate, offered-request success fraction and cost. A configuration may be cheaper per success yet reject too much traffic. Another may pass latency in healthy operation but fail during replica loss. Write which contract the experiment passed, which it failed, and the next uncertainty to test. An honest negative result is useful engineering evidence because it prevents a capacity promise that the service cannot keep.
Worked example
Invented run: four GPUs cost a total of 12 currency units per hour. In one hour, 18,000 requests complete, but 3,000 exceed the agreed latency limit. Effective success count is 15,000, so cost is 12 ÷ 15 = 0.8 units per thousand successful requests. A second configuration costs 10 units and serves 12,000 requests, all within the target. Its effective cost is 10 ÷ 12 = 0.833. The first configuration is cheaper by this measure despite its higher bill, but its timeout policy still needs to match the product contract.
Exercise
A 30-minute run costs six units and returns 9,000 responses. Of these, 1,500 are too late and 500 fail. Assume those sets do not overlap. Calculate effective cost per thousand successes and identify one missing experimental detail.
Model solution and rubric
Seven thousand responses meet both conditions. Cost per thousand is 6 ÷ 7, about 0.857 units. Missing details include arrival pattern, prompt and output distribution, warm-up treatment or confidence from repeated runs. A raw response count would produce 0.667 and overstate value. Verify that failed and late responses were not counted twice before applying the arithmetic.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
“Rejected requests can be excluded from every success metric.” They may be outside execution latency, but they still matter to the offered-request service contract. Report both denominators.
“A fixed number of clients is a fixed arrival rate.” Waiting clients submit less work when the service slows. Record actual offered arrivals and choose a workload model that matches the claim.
Interview probe
Evidence class: recommended. Original practice.
Your benchmark says the new server is twice as fast. What would you ask before accepting it?
Strong answer: I would ask for arrival behavior, lengths, latency distribution, failure count, cache state and the exact successful-work definition. Then I would replay the same trace on both versions and compare resource cost and tail latency.
Follow-up: What can a five-minute run establish, and what does it leave uncertain?
Weak answer indicators: Treating throughput as user success; omitting failed requests; using repeated identical prompts without stating cache effects.
Sources
Technical references: vLLM metrics design; vLLM prefix caching; Google SRE alerting on SLOs. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
A client waits for each response before submitting another. During overload, which risk affects the capacity claim?
AIt necessarily generates duplicate requestsBIt can reduce offered arrivals as latency risesCIt always makes server latency largerDIt measures only cold starts
Two configurations use different output-length limits. What first repair makes a speed comparison more credible?
AReplay a matched request and stopping-rule distributionBCompare only maximum tokens/sCDiscard requests with long outputs from one runDUse the cheapest GPU price as the deciding metric
A load generator plans 100 arrivals/s but records only 60 actual sends/s. What can the run establish?
AThe server supports 100 arrivals/sBThe server rejects exactly 40 requests/sCGPU compute is saturatedDBehavior under the measured 60 arrivals/s, subject to other limits
Explain how you would evaluate serving capacity with latency and cost evidence without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
A capacity result is a conditional statement. Preserve the workload, objective and failure count with the number.