Select a parallelism strategy from fit and communication constraints.
Mechanism and reasoning
Data-parallel serving places independent model replicas behind a router. Each replica handles separate requests. This is usually the simplest way to increase throughput when a model and its target cache fit on one device. The copies do not need to exchange intermediate activations for every generated token, but they duplicate weight memory.
Tensor parallelism divides work within model layers across devices. Pipeline parallelism divides layers into stages. These approaches can make a large model fit, yet they introduce communication and coordination costs. A deployment with more GPUs can therefore be slower for one request than a smaller deployment. Device interconnect and machine boundaries are part of the design, not minor configuration details.
Training uses similar names with different state and communication requirements. In distributed data parallel training, each process has a model copy and gradients are synchronized. Serving replicas do not normally synchronize gradients because they are not updating weights. Explaining this distinction prevents a common interview mistake.
Make two decisions separately. First, determine the smallest group of devices that can host one useful replica at the required precision and context size. Second, determine how many such replicas are needed for traffic and fault tolerance. If one replica spans four devices, losing one of those devices may remove the entire replica. Eight devices arranged as two groups of four do not provide eight independent failure units.
Compare feasible layouts with a fixed workload. Count cost per successful request and latency under a failed replica, not just peak output at full health.
Place replicas into real failure domains
Parallelism has both a performance topology and a failure topology. Draw both before choosing a device count. Suppose four GPUs are available on two hosts, with two GPUs per host. A two-device replica confined to one host uses its local interconnect. Placing one device from each host in each replica may expose both replicas to either host's failure. The same four devices and same nominal replica count can therefore give very different availability.
Layout
Replica 1
Replica 2
Loss of host A
Host-local groups
A1, A2
B1, B2
Replica 2 remains
Cross-host groups
A1, B1
A2, B2
Both groups lose a member
This table assumes that a replica cannot continue after losing one required tensor-parallel worker. Some systems support other recovery paths, but do not assume them without evidence. A process restart on the same host is not equivalent to continued service during a host outage. Restart requires healthy hardware, model loading and readiness checks; requests in flight may still fail.
A capacity plan also needs a routing policy. Suppose each healthy replica can sustain forty requests per minute at the target mix. Two replicas offer eighty nominally, but uneven routing can overload one at fifty while the other receives thirty. Simple request counts can be misleading if long prompts carry more work. A load-balancing decision can consider active tokens or queue estimates, while preserving tenant constraints and any cache-locality policy. The best policy depends on the workload and measurements, not just on the number of endpoints.
code
1Decision record for the teaching workload2Target: 55 requests/minute, including failure of one host3Measured capacity: 40 requests/minute per two-GPU replica4Placement: one complete replica per host5Minimum healthy replicas for target: ceil(55 / 40) = 26Installed replicas for one-host loss: 37Installed GPUs: 3 × 2 = 68Assumption: no shared router, storage or network bottleneck
The decision record makes an important point: four devices were sufficient for normal traffic but not the stated failed-state contract. If the available budget buys only four, choose an explicit degraded policy or revise the target. Do not hide the mismatch by quoting total healthy throughput. A queue can absorb a short burst, but sustained arrivals above surviving service capacity eventually exceed any finite deadline or queue bound.
Model fit is not the only reason to consider a larger parallel group. A latency requirement for one long request can make a larger group worth measuring, even when the model fits on one device. Conversely, distributing a small workload can spend more time communicating than it saves in computation. The claim must come from a matched benchmark. A successful fit test says nothing about whether tensor parallelism improved time to first token or generation cadence.
Pipeline parallelism introduces stage boundaries and scheduling effects. If stages have unequal work, a slow stage limits the pipeline. Do not divide layers equally and assume that memory, compute or communication will balance automatically. Embeddings and output heads can differ from transformer blocks; hardware and network paths can differ too. A useful answer asks what each stage stores and how long it runs before suggesting the split.
An interview follow-up may mix training with serving. In distributed data parallel training, ranks execute forward and backward work and coordinate gradient updates. In serving, independent replicas usually accept separate requests without gradient synchronization. Tensor parallelism may coordinate intermediate work in both contexts, but the stored optimizer state, communication schedule and fault recovery objectives differ. State which workload is under discussion before borrowing a familiar formula.
To validate this design, test three conditions: normal load with balanced routing, the same arrival trace after removing one complete failure domain, and recovery while a replica warms. Record the cold-start interval and whether the router sends work before the model is ready. Use the same token distributions across layouts. This method can disprove an attractive architecture quickly: the surviving replica count might be correct while warm-up or an overloaded shared gateway still breaks the target.
Worked example
A model needs 70 GiB for weights plus operating cache. Available GPUs provide 48 GiB each. A single device cannot host the target configuration. Four GPUs can form two two-device replicas, or one four-device replica. Suppose a measured teaching trace gives 40 requests per minute per two-device replica and 65 per minute for the four-device replica. Two smaller groups deliver 80 per minute and retain 40 after one group fails. The four-device group delivers 65 and has no independent replica. Choose two groups if those measured figures and the latency target hold.
Exercise
A model fits on one GPU. Two independent replicas each deliver 25 requests per minute. A two-GPU tensor-parallel replica delivers 35. Demand is 30. Compare normal capacity and capacity after a replica failure.
Model solution and rubric
Independent replicas provide fifty requests per minute normally and twenty-five after one fails. Tensor parallelism provides thirty-five normally and zero after that replica fails. Neither configuration meets thirty after a failure. A third independent replica gives seventy-five normally and fifty after one failure. Check router behavior and available memory before treating that as sufficient. The answer should distinguish device count from replica count.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
“Two replicas always survive one host failure.” Both can depend on the same host through cross-host parallel groups or shared infrastructure. Map the dependencies.
“A model fits on one GPU, so tensor parallelism cannot help.” Fit is only one constraint. A stricter latency target can justify measuring a larger group, though communication may erase the benefit.
Interview probe
Evidence class: recommended. Original practice.
Would you use all eight GPUs for tensor parallelism?
Strong answer: Only if fit or measured latency requires that group size. I would first find the smallest viable replica, then compare multiple independent replicas with a larger parallel group. More devices can add communication and enlarge the failure unit.
Follow-up: What changes when tensor communication crosses a slower inter-node link?
Weak answer indicators: Assuming linear speedup; treating distributed training and serving as identical; counting GPUs as independent replicas.
Sources
Technical references: vLLM parallelism; PyTorch distributed data parallel. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
Two replicas each span hosts A and B and require all their workers. What does loss of A remove?
AHalf the throughput with both replicas healthyBOnly the replica with the larger queueCNeither replica because weights are replicatedDBoth serving replicas
One GPU fits the model. Which evidence could justify tensor parallelism anyway?
AA matched benchmark shows it meets an otherwise missed request-latency targetBThere are idle GPUs in the accountCTraining uses gradient synchronizationDThe model has a larger vocabulary than last month
Each replica sustains 30 requests/minute. Demand is 50 after one replica loss. What is the minimum number of installed replicas under these assumptions?
Two equal replicas receive 50 and 30 requests/minute; each can serve 40. What first issue does the trace reveal?
AInsufficient combined healthy capacityBA requirement to increase every replica’s worker count before checking routingCUneven routing can overload one replica despite enough total capacityDA proven interconnect bottleneck
What distinguishes independent inference replicas from ordinary distributed data parallel training?
AInference replicas must synchronize gradients after each tokenBTraining replicas never exchange stateCInference replicas cannot share a routerDInference replicas normally serve separate work without gradient updates
Explain how you would select a parallelism strategy from fit and communication constraints without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Start with model fit, then measure the smallest viable replica. Keep failure units visible in the capacity plan.