Lesson 2 of 4 · 25 min

Design a checkpoint that can actually resume

Specify complete and verifiable training recovery state.

Mechanism and reasoning

A saved model is not always a resumable training checkpoint. Resuming can require optimizer state, scheduler state, random-number state, progress counters and information about the input position. Omitting these can change the training trajectory even if the model weights load successfully. Decide whether the requirement is approximate recovery or a stronger reproducibility target.
Distributed checkpointing adds coordination. Multiple ranks may write different shards. A directory appearing in storage does not prove all required shards are complete. Publish a manifest or completion marker only after the save has met its success criteria. A restore should reject an incomplete checkpoint rather than silently combine old and new pieces.
Checkpoint frequency trades write overhead against lost work. A short interval reduces recovery loss but consumes storage bandwidth and may pause training. An asynchronous save can reduce blocking, yet it still needs a consistent state and competes for memory and I/O. Copying live tensors without a consistency plan can create a checkpoint that never represented one valid step.
Version the checkpoint's schema and the code needed to read it. Record the model and optimizer configuration. Distributed checkpoint tools may reshard for changed world sizes, but that does not guarantee every custom state object or data loader can recover correctly. Test the required topology changes explicitly.
The practical acceptance test is a failure drill. Stop a disposable run after a known step, restore, and compare progress, optimizer state and expected behavior. A checksum proves bytes were transferred intact. It does not prove that the right state was saved or that a resumed job uses it correctly.

Define a checkpoint commit protocol

A recoverable checkpoint is a logical state at a defined training boundary. Decide whether the save represents the state before or after an optimizer update, and record that convention. A counter called step can mean a microbatch, a gradient accumulation boundary or an optimizer update. If progress and optimizer tensors refer to different boundaries, the job can repeat or skip an update after recovery.
code
1Illustrative manifest for checkpoint C10002schema_revision: 33optimizer_updates_completed: 10004model_revision: training-configuration digest5required_shards: rank0, rank1, rank2, rank36state: weights, optimizer, scheduler, RNG, data-progress7checksums: one for each required object8completion: published only after all required writes are durable
This is a conceptual application manifest, not an assertion that every framework writes these exact fields. The implementation must match the checkpoint library and storage semantics. A completion marker that appears before data is durable is not a reliable commit boundary. A marker written by one rank without confirming the required set is only a local success signal.
Use a unique checkpoint generation and keep incomplete writes separate from the last committed generation. A reader should select only a generation whose completion and required objects validate. Do not overwrite pieces of the last good checkpoint in place while producing the next one. If an interrupted save leaves partial new files, the previous complete generation remains a valid fallback.
Failure pointAvailable evidenceRestore decision
Before any new shard completesPrevious manifest validUse previous generation
Three of four new shards completeNew generation incompleteUse previous generation
All shards written, marker absentCompletion not establishedFollow explicit recovery policy; do not guess
Marker and objects validateComplete candidateRun semantic restore checks
The third row can have different implementations. A recovery tool may validate all objects and complete an interrupted publication under a documented protocol. This lesson's default restore path does not infer completeness from file count alone. The important requirement is that the decision is explicit and reproducible, not based on which directory has the newest timestamp.
Asynchronous saving changes the consistency problem. The job may continue updating tensors while another thread or process writes. A snapshot must capture a consistent state, often through staging or copying under the framework's supported mechanism. The extra memory and I/O can affect training even when the main thread no longer blocks for the whole write. Measure step-time tails during saves as well as average pause duration.
Recovery semantics for data deserve their own paragraph in an interview. Restoring weights and optimizer at update 1000 while resuming input from a later cursor can skip examples whose gradients were never committed. Resuming from an earlier cursor can repeat examples. Either may be acceptable for an approximate recovery contract, but neither should be called exact without evidence. Include sampler state, shuffle seed and the relevant data position or documented replay policy.
A changed world size adds another layer. A library can reshard model and optimizer tensors yet leave a custom sampler or external input cursor incompatible. Verify all state components, not just tensor loading. Check whether the new topology changes effective batch size and gradient accumulation. A run that resumes with twice the effective batch and the old learning-rate schedule may be a valid experiment, but it is not the same training continuation.
Finally, test recovery with a small deterministic run. Save at a known update, continue for several steps, interrupt a copy, restore and compare the next updates under controlled inputs. Exact numerical equality may depend on deterministic kernels and hardware; define tolerances or the weaker invariant being tested. Also corrupt or omit one shard deliberately and verify that the loader refuses the incomplete generation. The restore path is part of the training system and must be exercised before an expensive failure makes it necessary.
The checkpoint library does not discover every piece of application state. PyTorch DCP calls the state methods that a supplied Stateful object defines; the tutorial wrapper includes model and optimizer state. Custom sampler progress, external cursors and generator state need an explicit representation and restore rule. torch.get_rng_state() covers the default CPU generator; torch.cuda.get_rng_state_all() returns CUDA device generator states. Other generators used by the program need their own handling. A changed device topology requires an explicit state-mapping policy. Saving random state is necessary for some repeatability contracts, but it does not promise identical results across PyTorch versions, platforms or CPU/GPU execution. Test the stated continuation guarantee under the actual supported environment.

Worked example

Teaching run saves every twenty minutes; each synchronous save takes one minute. A failure occurs seventeen minutes after the last completed checkpoint. Restore takes four minutes. Lost compute is seventeen minutes, and recovery delay is four, excluding capacity replacement. If failures are roughly uniform within the interval, average lost work is approximately ten minutes, but this particular incident lost seventeen. Save overhead is roughly one minute per twenty-minute cycle under the stated convention. State whether the interval includes save time before comparing policies.

Exercise

The last complete checkpoint is step 1,000. A directory for step 1,100 contains only three of four required shards. The job failed at 1,130. Choose a restore point and list two checks after restore.

Model solution and rubric

Restore step 1,000 because 1,100 is incomplete under the stated four-shard requirement. Verify optimizer and scheduler progress and verify the data position or replay rule. Compare a known batch result or state digest where deterministic behavior is expected. The existence of three fresh shards is not permission to mix them with one old shard. Record that 130 steps need replay under this recovery choice.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“A checksum proves resumability.” It proves integrity of specific bytes, not that optimizer, progress and data state form a consistent training boundary.
“Resharding tensors guarantees recovery at any world size.” Custom data state and effective batch semantics may still change. Test the required topology explicitly.

Interview probe

Evidence class: recommended. Original practice.
The model loads after a crash, but training behaves differently. What might be missing?
Strong answer: Optimizer moments, scheduler position, random state or data progression may differ. I would compare the checkpoint manifest with the recovery contract and test a controlled interrupted run against an uninterrupted baseline.
Follow-up: Which parts of the promise change if the resumed world size differs?
Weak answer indicators: Treating weights as complete training state; declaring success from a nonempty directory; equating checksums with semantic correctness.

Sources

Technical references: PyTorch distributed checkpoints; PyTorch distributed data parallel; PyTorch CPU generator state, version 2.14; PyTorch CUDA generator states, version 2.14; PyTorch reproducibility limits, version 2.14. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsPyTorch distributed checkpointsdocs.pytorch.orgdocsPyTorch distributed data paralleldocs.pytorch.orgdocsPyTorch CPU generator state, version 2.14docs.pytorch.orgdocsPyTorch CUDA generator states, version 2.14docs.pytorch.orgdocsPyTorch reproducibility limits, version 2.14docs.pytorch.org

Checkpoint

Four shards are required; three new shards exist and the previous manifest validates. Restore choice?

AMix one old shard with three newBUse new weights but old optimizer without disclosureCUse the previous complete generationDChoose the newest directory timestamp
Sign up free to answer and see why

Checkpoint

A checkpoint records microbatch 20, with four microbatches per optimizer update. Why is a separate committed optimizer-update boundary needed?

AThe scheduler can always infer it from the current learning rateBThere must have been 20 optimizer updates because 20 batches ranCIt distinguishes accumulated but uncommitted gradients from completed updatesDRestoring the data cursor alone reconstructs all pending gradients
Sign up free to answer and see why

Checkpoint

All model weight files pass checksums. The recovery promise includes the same next optimizer update. What is still required?

AOnly a model-output check on one inference inputBA later checkpoint timestamp than the previous generationCOnly successful tensor allocation at the new world sizeDCompatible optimizer, progress, data and relevant random state at one boundary
Sign up free to answer and see why

Checkpoint

An asynchronous save reads tensors while training updates mutate them. What most directly establishes a consistent save boundary?

AUse the framework's supported staging or snapshot mechanism before later mutationBWrite the completion marker immediately and retry failed shards laterCCompute checksums after each shard without coordinating the state boundaryDRetain more generations while allowing each shard to capture a different update
Sign up free to answer and see why

Checkpoint

Tensor resharding succeeds when the world size changes. Which check addresses a remaining continuation risk?

ARequire the old number of shard files even if the loader supports reshardingBVerify custom data progression, effective batch, accumulation and random-state policyCVerify only that the same model architecture loadsDCompare only the checkpoint byte total before and after resharding
Sign up free to answer and see why

Explain how you would specify complete and verifiable training recovery state without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Define what resume means, save all required state, and prove the restore path with an interrupted run.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.