Lesson 2 of 4 · 35 min

A checkpoint is a continuation contract

Explain which state is required to resume an optimizer trajectory.

Loading model weights is sufficient for many inference tasks. Continuing training is different. The next update depends on optimizer state, scheduler position, data order, random-number state, and sometimes gradient-scaling state. If those are missing, training may continue, but it is not a faithful continuation of the interrupted run.
Momentum illustrates the issue. An optimizer carries information from earlier gradients. Restoring only the parameters resets that history and changes the next step. Adaptive optimizers also maintain running statistics. A learning-rate scheduler needs the correct progress count. Mixed-precision training may need a loss-scaler state. Record what the chosen stack actually requires instead of assuming one universal checkpoint layout.
Data state matters too. If a resumed run starts the epoch again, it can repeat examples and change the effective training distribution. If it skips too far, it can omit them. A resume protocol should define the sample cursor or sampler state and how distributed ranks reconstruct it. Exact bitwise continuation can still depend on the software and hardware environment.
Test interruption deliberately. Run a short reference for a fixed number of steps. In another run, stop at an intermediate checkpoint and resume. Compare the next batch identity, learning rate, optimizer state, and resulting parameters within the intended tolerance. A final metric alone may hide divergence.

Worked example

A teaching optimizer uses momentum velocity v = 0.9v + g and update w = w - 0.1v. At checkpoint, w = 1.0 and v = 2.0. The next gradient is 1.0. A correct resume gives v = 2.8 and w = 0.72. A weights-only resume resets v to zero, giving v = 1.0 and w = 0.90.
Both runs execute valid arithmetic, but they follow different trajectories. Calling the second one "resumed exactly" is false. If the intention is a new fine-tuning run from old weights, resetting state may be appropriate, but that intention must be explicit and compared under a new experiment identity.

Exercise and solution

A run saved at step 500 restores weights and optimizer, but the scheduler restarts at step zero and the sampler returns the first training batch. Name the two defects and a test.
The learning-rate schedule is misaligned, and the sample sequence repeats from the start. The test compares uninterrupted and resumed runs at step 501, checking learning rate, batch IDs, and parameter change. Award one point each for the scheduler defect, data-state defect, intermediate-state comparison, and distinction between resume and warm start. A matching model architecture does not establish continuation.

Lab artifact: a continuation trace

Use an uninterrupted reference and an interrupted copy with the same intended procedure. For this example, a step means one completed optimizer update; checkpoints are captured after that update and before the next batch is consumed.
ObservationUninterrupted next stepCorrect resumeDefective resume
Completed updates before step500500500
Next batch IDs[A91, A92][A91, A92][A01, A02]
Learning rate0.0010.0010.01
Momentum state hashV500V500V500
Next augmentation RNG stateR501R501R001
The defective run restores optimizer state but still changes both inputs and schedule. Comparing only a later accuracy score would not isolate either error. Compare state at the earliest boundary where the traces diverge. The same principle applies if a checkpoint occurs during gradient accumulation: partial gradients, microbatch position, and accumulation scaling may become part of the continuation state. The simpler safe teaching protocol checkpoints only at a completed update boundary.
An explicit conceptual envelope makes that boundary reviewable:
code
1checkpoint_schema: continuation-v22boundary: after completed optimizer update, before next batch3model_state: W5004optimizer_state: O5005scheduler_state: LR5006random_states: python, numpy-generator, framework-cpu, framework-device7sampler_state: epoch, global sample cursor, rank assignment rule8precision_state: scaler if used9completed_updates: 50010data_manifest: D3111environment_manifest: E9
Not every framework exposes all state in this exact form. This is a checklist for a specified continuation claim, not drop-in library code. Consult the stack's actual serialization and data-loader behavior. A saved data cursor may be insufficient when workers prefetch examples with random transformations. The resumed process needs a defined way to reconstruct which examples were committed to completed updates, rather than merely which examples workers had already fetched.

A second failure case: a torn checkpoint

Writing model state and optimizer state to separate mutable files can leave a mixed checkpoint after interruption: new weights with old optimizer state. Give the checkpoint bundle a generation identity and publish a completion manifest only after all required artifacts are successfully written and verified. On load, reject incomplete or mismatched generations. Depending on storage, this may use atomic rename or a versioned manifest; a rename's guarantees must match the actual filesystem or object store.
Keep the previous complete generation until the new one is confirmed. File existence alone is not a validity test. Check expected keys, shapes, manifest identity, and a known load-and-next-step fixture. These measures address serialization integrity separately from numerical reproducibility.

Exercise: specify an interruption test

A loop accumulates gradients over four microbatches, then updates once. It crashes after the third microbatch. The stored checkpoint is from the preceding completed update. What should a faithful restart do under this lesson's update-boundary protocol?
Restore that checkpoint and replay all four microbatches for the next update from their recorded data and random state. The three in-memory partial contributions were never committed in the checkpoint and must not be counted as completed training. If a design instead saves mid-accumulation, it must restore partial gradients and exact microbatch state; that is a different, more complex contract. Award two points for the boundary, two for replay/state, and one for distinguishing protocols.

Misconceptions to correct

“An exact next-step match guarantees all future steps match” exceeds the test; later nondeterminism or missing state can still appear. “Weights-only loading is always a bug” ignores valid inference and warm-start uses. Name the operation before judging its required state.
The momentum example assumes the stated recurrence, no Nesterov update, no dampening, and no weight decay. Other optimizer conventions can give different values. In an interview, write the recurrence first. This prevents an argument about library defaults from hiding the actual question: whether the state that determines the next update was preserved.

Interview probe

Original practice: Why does resumed training diverge from uninterrupted training? A strong answer inspects optimizer, scheduler, RNG, sampler, precision, and environment state. Follow up with intentional fine-tuning from a checkpoint. A weak answer assumes weight loading means every state was restored.

Sources

docsPyTorch: saving and loading modelsdocs.pytorch.orgdocsPyTorch: reproducibilitydocs.pytorch.orgdocsPyTorch: automatic mixed precision examplesdocs.pytorch.org

Checkpoint

For v_next=0.9v+g and w_next=w-0.1v_next, with v=2,g=1,w=1, what is next w?

A0.90B1.28C0.80D0.72
Sign up free to answer and see why

Checkpoint

A checkpoint is defined after an optimizer update and before the next batch. Which cursor is relevant?

AThe last item prefetched by any worker, regardless of consumption.BThe next sample sequence after the completed update, with necessary RNG state.CThe beginning of the current epoch by default.DThe final example in the dataset manifest.
Sign up free to answer and see why

Checkpoint

A crash occurs after three of four accumulation microbatches; the last checkpoint is at the previous completed update. What matches that protocol?

ARestore and replay all four microbatches from that boundary.BReplay only the fourth without restoring partial gradients.CAdvance the scheduler as if the unfinished update completed.DStart from the next optimizer step and omit the partial batch.
Sign up free to answer and see why

Checkpoint

Weights are generation 12 and optimizer state generation 11 after a write failure. What should a continuation loader do?

AAccept the mix if tensor shapes match.BUse the larger file as authoritative.CReject the incomplete generation and use a verified complete bundle.DReset optimizer silently and report exact continuation.
Sign up free to answer and see why

Checkpoint

Why can a correct next-step test still be insufficient for a universal exact-resume claim?

AA match in loss alone also proves the next batch and RNG state match.BOne matched update covers every later data-loader epoch boundary.CA matched update on one device proves the same distributed-rank mapping.DLater untested state or nondeterminism can still diverge.
Sign up free to answer and see why

Can you define a checkpoint boundary and reproduce the next data and optimizer update? Rate confidence from 1 to 5 and name a state not covered by weights alone.

Not yetGetting thereConfident

Wrap-up

  • Define whether a load is inference, warm start, or continuation. Test the next step, not just the existence of a checkpoint file.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.