← All questions
MediumCodingSystem design

Write a training loop with gradient accumulation and mixed precision

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below

Reference answer

Then expect these follow-ups

  • Is accumulation exactly equivalent to a bigger batch?

  • (Almost — BatchNorm statistics differ; LayerNorm models like transformers are fine.) Where does the memory actually go?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question