Lesson 1 of 4 · 25 min

Diagnose a training slowdown across ranks

Locate a distributed training bottleneck from a timing trace.

Mechanism and reasoning

A distributed training step moves at the pace allowed by its dependencies. If ranks must meet at gradient synchronization, a slow rank can hold back the group. The visible symptom may be long communication time on otherwise healthy devices. That does not prove the network is the original cause. Those ranks may be waiting for a peer that has not reached the collective.
Split each rank's step into data wait, forward work, backward work and communication. Check that ranks execute matching collective operations in compatible order. Differences in control flow or failed workers can cause hangs that look like performance problems. Establish correctness before tuning bucket sizes or adding nodes.
Input skew matters. One rank may receive unusually long sequences, slow storage reads or expensive preprocessing. When the work is imbalanced, adding faster network hardware will not fix the slow input path. Compare distributions rather than one step, because occasional long batches can dominate tail step time.
Be careful with averages across ranks. A mean data-loading time can hide a rank that waits much longer than its peers. Collect aligned timestamps and the batch properties that explain them. In the interview, describe what evidence would separate input starvation, compute imbalance and communication bandwidth.
A useful diagnostic experiment changes one suspected cause while preserving the rest. Replacing the input with fixed synthetic batches can test the input path, though it does not represent real training throughput. Running a small collective benchmark can examine communication independently. Neither result alone proves the entire job is healthy. The answer should connect the local test back to the full trace.

Build a diagnosis that can survive a follow-up

Start by identifying what each timestamp measures. A CPU event marking entry into a collective is different from the GPU completing its previous kernels. Asynchronous execution can make a host-side interval look shorter or longer than the device interval of interest. Use aligned instrumentation and make synchronization introduced by the measurement explicit. A diagnostic trace can change performance if it forces every operation to wait.
code
1Illustrative observations from ten slow steps2Rank 2 input wait: 130–150 ms3Other ranks input wait: 15–25 ms4All ranks compute: 175–185 ms5Rank 2 sample source: remote shard S96Other ranks sample source: local cached shards7Collective completion after final arrival: 18–22 ms
This trace strengthens the input-path hypothesis because the collective's post-arrival duration is stable and short relative to the arrival skew. It still does not prove that the remote store is defective. Rank 2 might process a different batch shape, perform more decompression, or have a slower host CPU. The source label identifies a useful comparison, not a root cause.
A good experiment predicts more than “it gets faster.” If remote input is the cause, replaying the same decoded batches from local memory should reduce rank 2's input wait while leaving compute duration approximately unchanged. If compute also changes substantially, the experiment changed more than the input path or the first classification was incomplete. Preserve batch shapes and preprocessing outputs when possible.
HypothesisDiscriminating observationBounded experiment
Input storage delayLate input completion before computeReplay decoded batches locally
Unequal compute workSimilar input wait, longer device computeEqualize sequence lengths
Communication limitAligned entry, long collective completionIsolated matched-size collective test
Divergent collective orderRanks enter different operationsCompare operation sequence and failure logs
The final row is a correctness check. If one rank skips a collective because of data-dependent control flow, performance tuning cannot repair the mismatch. Check for worker exceptions, exhausted input iterators and inconsistent branches. A hang is not simply an extremely slow version of a healthy synchronization step. Establish that every participating rank follows the expected communication contract.
Now introduce sequence-length skew. Suppose rank 2 has twice as many padded tokens as the others. Equal sample count does not mean equal compute. Bucketing examples by length or using an appropriate distributed sampling strategy can reduce imbalance, but must preserve the training data contract. A shortcut that drops all long examples may improve step time while changing the learning problem. Report throughput in useful training tokens or examples under the same data policy.
A communication microbenchmark has a narrower role. It can reveal whether a specific message size and topology achieve expected behavior without the full training input path. It cannot show that the training graph issues the same communication pattern, overlaps it effectively, or avoids a late peer. Use the microbenchmark to challenge a hypothesis after reading the trace, then return to the real job to verify the effect.
The interview response should separate mitigation from investigation. If one data shard is temporarily slow and an approved replay path exists, switching to a healthy source can reduce harm. Record the data-order consequence and any repeated or skipped samples. If no safe mitigation exists, pause or checkpoint rather than silently change the training corpus. The correct action depends on whether the run values exact reproducibility, approximate continuation or simply valid optimization progress.
Close the diagnosis with a testable conclusion. For example: “The available trace shows rank 2 arriving late because its input stage is longer; the stable post-arrival collective duration makes network bandwidth a weaker first hypothesis. I will replay the same decoded batches locally and compare both input and compute intervals.” This is stronger than naming a library from a stack trace because it explains the evidence, its limits and the next result that could disprove it.

Worked example

Invented step trace in milliseconds:
RankInput waitComputeArrival at sync
020180200
120180200
2140180320
320180200
Ranks 0, 1 and 3 wait roughly 120 ms before rank 2 arrives. The immediate hypothesis is input delay on rank 2, not slower compute. Replay equal synthetic batches. If arrival times align, inspect that rank's storage and preprocessing. If delay remains, investigate runtime scheduling and measurement alignment.

Exercise

Four ranks arrive at synchronization at 250, 252, 251 and 410 ms. The last rank's compute is 160 ms longer. Estimate peer waiting and propose a discriminating experiment.

Model solution and rubric

The first three wait about 160, 158 and 159 ms for the slow rank, ignoring collective execution time. Equalize batch sequence lengths or replay identical synthetic work. If the difference disappears, workload skew is supported. If it persists, examine device health, contention and per-rank kernel timing. A faster collective cannot remove waiting caused before collective entry.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“Time inside all-reduce proves a network problem.” Early ranks can wait for a late peer. Compare arrival skew and post-arrival execution.
“Equal batch count means equal work.” Sequence length, preprocessing and padding can differ. Match useful work before comparing rank speed.

Interview probe

Evidence class: recommended. Original practice.
Every GPU shows time in all-reduce. Should we upgrade networking?
Strong answer: First align rank arrival times. Time attributed to a collective can include waiting for a late peer. I would compare input and compute timing and run a controlled communication test only after separating those causes.
Follow-up: What evidence would make you revise the input-skew hypothesis?
Weak answer indicators: Naming NCCL as the cause from a stack frame alone; comparing unaligned clocks; averaging away the slow rank.

Sources

Technical references: PyTorch distributed data parallel; PyTorch distributed checkpoints. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsPyTorch distributed data paralleldocs.pytorch.orgdocsPyTorch distributed checkpointsdocs.pytorch.org

Checkpoint

Rank 3 reaches synchronization 150 ms after its peers; collective work after its arrival is 20 ms. Which investigation best fits this trace?

ATune collective bucket size to remove the full 150 msBCompare rank 3 input completion and device compute against peer ranksCIncrease communication streams before checking collective arrival timesDAverage all rank input times to decide whether input is relevant
Sign up free to answer and see why

Checkpoint

Ranks enter different collective operations in the same step. What priority follows?

ATune bucket size firstBMeasure only average GPU utilizationCRepair the communication/control-flow mismatchDAdd input workers to every rank
Sign up free to answer and see why

Checkpoint

Arrival times are 200,210,205 and 360 ms. Rank 0's wait for the last arrival is?

A160 msB150 msC155 msD360 ms
Sign up free to answer and see why

Checkpoint

Local replay removes input skew but leaves unequal compute. What conclusion follows?

AThe entire incident was storage aloneBNetworking is proved faultyCThe input experiment failed because anything remained slowDInput delay was one factor; compute imbalance remains
Sign up free to answer and see why

Checkpoint

A synthetic collective benchmark is fast. What claim is supported?

AThe full training job cannot have communication problemsBThe tested message/topology case is fast in isolationCThe data loader is correctDEvery training batch has equal useful work
Sign up free to answer and see why

Explain how you would locate a distributed training bottleneck from a timing trace without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Find where ranks diverge before tuning communication. A waiting operation can reveal the victim rather than the cause.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.