Lessons
1Build a GPU memory budget25 min read
Calculate a defensible inference memory budget.
- →Calculate a defensible inference memory budget
2Explain prefill, decode and batching25 min read
Distinguish prompt processing from output generation.
- →Distinguish prompt processing from output generation
- →Choose batching limits that protect interactive latency
3Choose replication or model parallelism25 min read
Select a parallelism strategy from fit and communication constraints.
- →Select a parallelism strategy from fit and communication constraints
4Design an honest capacity experiment25 min read
Evaluate serving capacity with latency and cost evidence.
- →Evaluate serving capacity with latency and cost evidence
- →Separate capacity estimates from measured service guarantees
Skills in this course
- 01Calculate a defensible inference memory budgetCalculate a defensible inference memory budget.
- 02Distinguish prompt processing from output generationDistinguish prompt processing from output generation.
- 03Select a parallelism strategy from fit and communication constraintsSelect a parallelism strategy from fit and communication constraints.
- 04Evaluate serving capacity with latency and cost evidenceEvaluate serving capacity with latency and cost evidence.
- 05Choose batching limits that protect interactive latencyChoose batching limits that protect interactive latency.
- 06Separate capacity estimates from measured service guaranteesSeparate capacity estimates from measured service guarantees.