Lessons

1Optimize the critical path you actually measured35 min read

Use timing evidence and Amdahl's law to choose an optimization.

  • →Use a trace to identify the measured bottleneck
  • →Report throughput under a fixed quality protocol
  • →Separate overlap from additive timing
Read lesson
2Memory accounting before memory tricks35 min read

Build a first-order memory budget and choose a targeted mitigation.

  • →Estimate memory before selecting a mitigation
Read lesson
3Mixed precision changes arithmetic, so test the update35 min read

Explain loss scaling and the correct order for gradient clipping.

  • →Preserve optimization semantics with mixed precision
  • →Report throughput under a fixed quality protocol
Read lesson
4Distributed training must preserve the intended average35 min read

Calculate the difference between rank means and a sample-weighted mean.

  • →Check distributed gradient and sample-count semantics
Read lesson

Skills in this course

  1. 01Use a trace to identify the measured bottleneckUse a trace to identify the measured bottleneck.
  2. 02Estimate memory before selecting a mitigationEstimate memory before selecting a mitigation.
  3. 03Preserve optimization semantics with mixed precisionPreserve optimization semantics with mixed precision.
  4. 04Check distributed gradient and sample-count semanticsCheck distributed gradient and sample-count semantics.
  5. 05Report throughput under a fixed quality protocolReport throughput under a fixed quality protocol.
  6. 06Separate overlap from additive timingSeparate overlap from additive timing.