Lesson 4 of 4 · 25 min

Defend an inference platform design under failure

Build a capacity and recovery argument for a shared model service.

Mechanism and reasoning

A design interview should end with a checkable decision record. Begin with request volume, prompt and output distributions, latency targets, tenant boundaries and budget. Ask which conditions are contractual and which are estimates. The same model can need very different infrastructure for interactive chat and overnight batch processing.
Choose a replica shape from memory fit and measured performance. Then size the fleet for the failure condition. If the target must hold after losing one replica, normal utilization must leave enough spare capacity. That reserve has a cost. State it rather than hiding it in an unexplained multiplier.
Failure domains are larger than processes. Replicas on the same host, zone or shared storage path may fail together. A design that claims host tolerance must place surviving capacity outside that host. Include the router, model registry and authentication path in the dependency review. Replicating GPUs does not repair a shared dependency outage.
The platform needs a control plane for versions, policies and rollout, and a request path that continues to behave predictably when those controls are briefly unavailable. Cache configuration safely where appropriate, with explicit expiry and revocation behavior. Avoid making every generated token depend on a remote configuration lookup.
Use a small table to show normal capacity, failed capacity and maximum admission. Explain cold-start delay and recovery time. Close with the highest-risk assumption and a measurement plan. Interviewers can probe a concrete calculation. A diagram full of named products gives them less evidence that the candidate understands the service.

Make the failure requirement executable on paper

Write the requirement before choosing a topology. “Highly available” is too vague to size. A useful statement names the traffic mix, latency objective, tolerated failure unit and allowed recovery interval. For example: “Maintain 120 requests per minute at the measured mix after loss of one host, with no reliance on a cold replacement for the first three minutes.” This immediately excludes a design whose spare capacity exists only as an autoscaling target.
Hosts installedReplicas per hostHealthy capacityOne-host-loss capacity
22160/min80/min
32240/min160/min
42320/min240/min
Each replica in this example sustains forty requests per minute under the defined workload. The table shows why three hosts are the minimum for the 120-per-minute failed-state target under this placement. It does not show that all 240 healthy requests could be admitted while still promising the same failure tolerance. If the service admits 240 and then loses one host, surviving 160 capacity is insufficient. The admission promise must be aligned with the failure reserve.
A second failure domain can invalidate the table. Suppose all three hosts are in one zone and the contract changes to zone loss. The table no longer provides a surviving host. Adding another process in the same zone does not solve that requirement. Place enough independent capacity in other zones and check whether network, artifact storage and policy services remain reachable. The arithmetic must follow the actual dependency graph.
code
1Illustrative dependency review2Request path: gateway -> tenant policy cache -> model replica3Model startup: artifact registry -> weight download -> warm-up4Control path: release controller -> routing configuration5Runtime promise: existing ready replicas continue during brief registry outage6Limit: replacement replicas cannot start without artifact access7Revocation promise: cached tenant policy expires under a defined rule
The distinction between runtime and startup dependencies matters. A model registry outage may not stop already loaded replicas, but it can prevent recovery after a host failure. A design can therefore pass a single isolated host-loss test and fail during a combined artifact outage. State which combinations are within the contract and which remain risks. Do not claim tolerance for every simultaneous failure without paying for and testing it.
Policy caching presents a tradeoff. Serving from cached configuration can avoid making every request depend on a remote control service. However, stale authorization or quota policy can violate revocation expectations. Define maximum staleness, behavior at expiry and which policy changes require immediate enforcement. A general “cache everything” answer hides the difference between a harmless rollout label and a revoked tenant credential.
Use a capacity worksheet that includes startup and backlog. If one host fails and surviving capacity only equals arrivals, existing in-flight retries or queued work may not clear. Reserve extra capacity for the recovery burst or use bounded retry and admission policies. A client retry storm can make the effective arrival rate much larger than normal business traffic. Stable request identity and backoff are part of the capacity plan because they control repeated work.
Cost should be attached to the service promise. Three hosts cost more than the two that meet healthy demand. The extra expense buys immediate capacity after the named failure. If the budget cannot support it, propose a concrete alternative such as a lower admitted rate during failure, a longer bulk deadline, or a weaker failure target. Present the effect on users so the tradeoff can be decided explicitly.
End the interview with a verification plan. Remove one host during a matched workload, measure ready capacity, routing skew, retries, queue age and useful completions, and then observe replacement warm-up. Repeat with the relevant control dependency unavailable if that combination is in scope. The result should show whether the topology fulfills the written requirement. A diagram becomes useful when each dependency and reserve has a testable role.

Worked example

Teaching requirement: 90 requests per minute, each replica supports thirty within the latency target, and one-replica failure must be tolerated. Three replicas meet normal demand but leave only sixty after failure. Four leave ninety and satisfy this simplified condition. If replicas are placed two per host and a host failure is the requirement, four leave only sixty. Six replicas across three hosts, two per host, leave 120 after losing one host. The design depends on the stated failure unit and measured per-replica throughput.

Exercise

Demand is 120 requests per minute. One replica supports forty. Replicas are grouped two per host. How many hosts are required to meet demand after one host fails, assuming no other bottleneck?

Model solution and rubric

Each host provides eighty. After losing one of three hosts, two remain with 160, which meets demand. Two hosts would leave only eighty. Thus three hosts are the minimum under the fixed two-replica layout. Verify that the load balancer, model artifacts and tenant policy service remain available across the same failure. This arithmetic does not prove a shared dependency is resilient.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“Unused reserve is waste by definition.” If immediate failed-state service is required, reserve is the mechanism that supplies it. Its value depends on the agreed requirement.
“A healthy GPU fleet proves the platform is available.” Routing, authorization and artifact dependencies can still prevent service or recovery. Check the full path.

Interview probe

Evidence class: recommended. Original practice.
Why are you paying for six replicas when average demand needs three?
Strong answer: The requirement is continued service after a whole host fails, and replicas share hosts. I would show the failed-state capacity table and then ask whether a lower availability target or slower recovery is acceptable to reduce cost.
Follow-up: What if the failure unit changes from one host to one zone?
Weak answer indicators: Sizing only from average healthy load; counting co-located replicas as independent; adding redundancy without stating the failure requirement.

Sources

Technical references: vLLM parallelism; vLLM metrics design; Google SRE alerting on SLOs. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsvLLM parallelismdocs.vllm.aidocsvLLM metrics designdocs.vllm.aidocsGoogle SRE alerting on SLOssre.google

Checkpoint

Demand 120/min, each host 80/min, one-host loss required. Minimum hosts under these assumptions?

A3B2C4D5
Sign up free to answer and see why

Checkpoint

Three independent hosts pass a one-host-loss test, but all hosts share one zone. What follows about a zone outage?

AThe test does not establish zone survival because all tested capacity shares that dependencyBThe host-loss reserve proves zone survival whenever average utilization is below halfCThe zone survives because a scheduler can recreate three desired replicasDThe host test proves zone tolerance if all three process health probes pass
Sign up free to answer and see why

Checkpoint

Ready replicas continue during registry outage, but replacements cannot load. What is the right description?

AThe registry is not a dependencyBThe service tolerates all combined failuresCRuntime continues for existing capacity, while recovery capacity is constrainedDA registry outage instantly clears all GPU weights
Sign up free to answer and see why

Checkpoint

Failed-state capacity equals continuing arrivals, and 600 old requests remain queued with valid deadlines. What is needed to clear them?

AKeep admitting the full rate and increase only the queue limitBAdd spare service rate or apply a bounded admission/expiry policyCShorten retry backoff so each old request competes more oftenDRoute evenly across the same capacity and assume the backlog drains
Sign up free to answer and see why

Checkpoint

Budget cannot afford the agreed one-host-loss reserve. What response is defensible?

APresent an explicit lower failed-state service target or other bounded tradeoffBClaim autoscaling makes reserve freeCKeep the guarantee and omit cost from the designDCount cold desired replicas as immediately ready
Sign up free to answer and see why

Explain how you would build a capacity and recovery argument for a shared model service without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • State the failure unit before counting replicas. Defend reserve cost with a concrete service requirement.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.