Design admission limits using token and time budgets.
Mechanism and reasoning
A request limit is useful only if it bounds the scarce work. Ten short prompts and ten long documents can have very different GPU costs. Admission should consider input tokens, allowed output, concurrent state and waiting time. Some costs are known before execution. Others need a conservative reservation followed by reconciliation when the request finishes.
Separate tenant quota from global capacity. A tenant may be within its monthly allowance while the shared pool is saturated. Conversely, an idle pool should not permit one tenant to exceed a contractual or financial limit. Keep both decisions visible in the response and metrics. A rejected request is an admission outcome, not a model execution failure.
A queue buys time, not capacity. Set a maximum wait that fits the user deadline. If a request has already waited most of its useful lifetime, launching expensive work may create output nobody can use. Cancellation should release the reservation and stop generation where the runtime supports it. Retries must not turn one request into several concurrent generations.
Fairness also needs a unit. Equal requests can be unfair when lengths differ. Token reservations, weighted queues or separate pools can provide better control, but estimated work can still be wrong. Reconcile actual usage and watch for tenants that repeatedly reserve small work then use the maximum. Do not claim a scheduler is fair without saying what resource and time interval it shares.
The practical test is a mixed trace. Send small interactive work while one tenant submits a burst of large requests. The interactive class should retain the agreed service level, and the bulk class should receive a clear bounded outcome.
Make the reservation a state machine
The admission decision must be atomic with respect to the counters it checks. Two concurrent requests can both read the same free capacity and then each reserve it. A process-local lock is insufficient when multiple gateways share the pool. Use one authoritative reservation operation or another mechanism that provides the required conditional update. The example below is an original pseudocode design for the contract, not a complete distributed implementation or a claim that vLLM provides this admission transaction.
code
1reserve(caller, request_id, tenant, parameters, tokens, deadline):2 authorize caller for this tenant and action3 key = (tenant, request_id)4 atomically against unique key and shared tenant/pool counters:5 if key exists:6 require stored parameters match this request; otherwise conflict7 return recorded outcome without creating another execution8 require deadline permits useful execution9 require tenant_active + tokens <= tenant_limit10 require pool_active + tokens <= pool_limit11 create reservation(key, parameters, tokens, state=RESERVED)12 increment tenant_active and pool_active by tokens1314finish(trusted_execution, key, actual_usage, reason):15 verify authentic runtime stop confirmation bound to this execution16 atomically:17 require execution identity and generation match the recorded owner18 if reservation is terminal: return recorded result19 decrement active counters by original reserved tokens20 record actual_usage and terminal reason
The operation must also define retry behavior. A client that retries after a network timeout does not know whether admission succeeded. A stable request identifier lets the server return the existing outcome instead of creating a second generation. Scope the identifier to a tenant and bind it to the request parameters. Reusing the same identifier for different input should produce an explicit conflict, not silently return an unrelated answer.
A crash between reservation and execution creates another problem: capacity can remain held without useful work. A lease with an owner and heartbeat can bound that leak, but expiration must be coordinated with execution. If a gateway expires a lease while the GPU still runs the request, admitting replacement work can exceed the physical budget. A lease is an ownership protocol, not permission to assume vanished bookkeeping means vanished computation. Define how the execution owner stops or renews. Fencing only a database write does not stop GPU work or release its memory. Reuse the reservation only after the runtime confirms that the old execution cannot consume the reserved capacity; otherwise keep that possible live work charged to the pool.
Event
Reservation state
Active tokens released?
Usage action
Request admitted
Reserved
No
None yet
Scheduler starts work
Running
No
Begin measured usage
Client asks to cancel
Cancelling
Not until stop is confirmed
Preserve usage so far
Runtime confirms stop
Cancelled
Full reservation, once
Record actual usage
Duplicate stop event
Cancelled
No further release
Return recorded result
This table distinguishes a requested cancellation from completed cancellation. A network disconnect alone does not prove that generation stopped. Releasing capacity too early can create oversubscription. Releasing only the used portion leaves the unused reservation stuck forever. Releasing the full amount twice drives counters below the true active amount. These bugs often appear only when retries, timeouts and worker failures overlap.
Fairness needs a stated objective as well as correct counters. Suppose tenant A sends forty small requests and tenant B sends four large ones with the same total reserved tokens. First-in, first-out admission can let either burst block the other. Weighted queues can limit how much work each tenant receives over a chosen interval. Separate interactive and bulk classes can protect a short deadline, but strict priority can starve bulk work indefinitely. Give the bulk class a minimum share or a bounded wait policy if its contract requires progress.
Now consider estimation error. Reserving maximum output tokens is simple and predictable but can reduce utilization when most outputs are short. Reserving an estimate permits more work but requires a policy when actual output grows. A safe design may request more reservation before crossing the initial bound, then stop or defer according to the user contract if capacity is unavailable. Silently allowing growth defeats the original budget.
Test the lifecycle with a deliberately small pool. Submit two concurrent requests that each require more than half its capacity. Only one should acquire the reservation. Cancel it, send duplicate completion events, retry its request identifier and restart the admission process. Verify the invariant: recorded active reservations equal the sum of nonterminal owned work under the documented recovery protocol. This is more useful than a happy-path test that merely checks one successful response.
Worked example
Invented pool: admission allows 40,000 reserved tokens. Request A reserves 12,000, B 16,000 and C 8,000. Total is 36,000. Request D asks for 7,000, so it cannot start immediately. If its maximum queue wait is two seconds and no reservation is released by then, reject it without model execution. When B completes using 10,000 tokens, release its full active reservation and record 10,000 as actual billable usage under the chosen policy. Active-memory reservation and billing counters serve different purposes.
Exercise
A tenant has a 15,000-token active limit and already holds 9,000. Global free capacity is 20,000. A new request reserves 8,000. Decide admission, then explain what changes when the earlier request ends.
Model solution and rubric
Reject or queue under the tenant rule because 9,000 + 8,000 exceeds 15,000, even though global capacity is available. After the earlier request ends and releases its reservation, the new request fits if its deadline remains useful. Record the reason as tenant concurrency or token capacity. Verify that completion, timeout and cancellation all release reservations exactly once, including duplicate completion notifications.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
“A cancellation request frees capacity immediately.” The runtime may still execute. Release only when the runtime confirms that the old work can no longer consume the reserved capacity.
“Idempotent billing is enough.” Duplicate generation can consume capacity even if the customer is charged once. Request identity must cover admission and execution as well as charging.
Interview probe
Evidence class: recommended. Original practice.
Why not let the runtime accept every request and rely on autoscaling?
Strong answer: Startup takes time, and an unbounded queue can spend the user's deadline before capacity arrives. Admission creates a predictable overload response and protects existing work. I would combine a measured warm pool with tenant and global token budgets.
Follow-up: How do retries reuse a request identity without charging or reserving twice?
Weak answer indicators: Using requests per second as the only limit; treating queue size as processing capacity; releasing only actual tokens from an active reservation.
Two gateways each see 6,000 free tokens and concurrently admit 4,000. What prevents oversubscription?
AAtomic conditional reservation against authoritative countersBReconcile billing at the end of the monthCUse a separate local counter in each gatewayDIncrease the client's timeout
The gateway lease expires while GPU execution may continue. What must happen before capacity is safely reused?
AAssume the process stoppedBDelete the usage recordCExtend the user's deadline onlyDConfirm that the runtime stopped the old execution and released its reserved capacity
Strict priority protects interactive traffic but bulk work never runs. What missing policy is exposed?
AA minimum bulk share or bounded wait contractBA larger maximum output for all requestsCA higher global concurrency limit without per-class schedulingDA shorter interactive timeout without a bulk progress rule
Explain how you would design admission limits using token and time budgets without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Reserve before execution, distinguish tenant policy from pool capacity, and close the reservation on every terminal path.