Lessons

1Build a GPU memory budget25 min read

Calculate a defensible inference memory budget.

  • →Calculate a defensible inference memory budget
Read lesson
2Explain prefill, decode and batching25 min read

Distinguish prompt processing from output generation.

  • →Distinguish prompt processing from output generation
  • →Choose batching limits that protect interactive latency
Read lesson
3Choose replication or model parallelism25 min read

Select a parallelism strategy from fit and communication constraints.

  • →Select a parallelism strategy from fit and communication constraints
Read lesson
4Design an honest capacity experiment25 min read

Evaluate serving capacity with latency and cost evidence.

  • →Evaluate serving capacity with latency and cost evidence
  • →Separate capacity estimates from measured service guarantees
Read lesson

Skills in this course

  1. 01Calculate a defensible inference memory budgetCalculate a defensible inference memory budget.
  2. 02Distinguish prompt processing from output generationDistinguish prompt processing from output generation.
  3. 03Select a parallelism strategy from fit and communication constraintsSelect a parallelism strategy from fit and communication constraints.
  4. 04Evaluate serving capacity with latency and cost evidenceEvaluate serving capacity with latency and cost evidence.
  5. 05Choose batching limits that protect interactive latencyChoose batching limits that protect interactive latency.
  6. 06Separate capacity estimates from measured service guaranteesSeparate capacity estimates from measured service guarantees.