Lessons
1Put admission before the expensive work25 min read
Design admission limits using token and time budgets.
- →Design admission limits using token and time budgets
- →Design fair admission for unequal request sizes
2Make cache reuse measurable and safe25 min read
Evaluate prefix caching without assuming universal speedups.
- →Evaluate prefix caching without assuming universal speedups
3Scale using queue behavior and warm capacity25 min read
Choose scaling signals that predict user-visible overload.
- →Choose scaling signals that predict user-visible overload
4Release a model without losing the escape route25 min read
Specify a model rollout and rollback contract.
- →Specify a model rollout and rollback contract
- →Specify rollback checks for a model release
Skills in this course
- 01Design admission limits using token and time budgetsDesign admission limits using token and time budgets.
- 02Evaluate prefix caching without assuming universal speedupsEvaluate prefix caching without assuming universal speedups.
- 03Choose scaling signals that predict user-visible overloadChoose scaling signals that predict user-visible overload.
- 04Specify a model rollout and rollback contractSpecify a model rollout and rollback contract.
- 05Design fair admission for unequal request sizesDesign fair admission for unequal request sizes.
- 06Specify rollback checks for a model releaseSpecify rollback checks for a model release.