Lessons
1The ML system design framework46 min read
The spine every ML system design round rides — framing, success metrics, data, model, serving, monitoring — how the 60 minutes are actually budgeted, the six dimensions interviewers grade, and the three cardinal sins that silently sink strong engineers by minute 10.
- →Problem Framing and Metrics
2Data, features & embedding pipelines47 min read
Where the label actually comes from and why “click = positive” is the most common wrong default; feature stores and the offline/online split that creates training-serving skew; embeddings as a pipeline, not a model call; batch vs streaming features; and the leakage that fakes a great offline score.
- →Data and Feature Pipelines
3Serving, latency & cost at scale48 min read
The four serving modes and how to pick one; the LLM serving stack that actually wins — continuous batching, PagedAttention, prefix caching — with the real multipliers; GPU economics and the two-number latency SLO (TTFT and ITL); model routing; and the autoscaling lag that bounds it all.
- →Serving and Scale
4Monitoring, drift & retraining46 min read
The monitoring story is graded as heavily as the modeling — “watch accuracy on a hold-out” is the wrong answer. Online vs offline metrics, the drift taxonomy and which statistical test to use, shadow vs canary, the three layers of rollback, the retraining loop, and a worked incident scenario.
- →Monitoring and Recovery
- →Problem Framing and Metrics
5Worked design: recommendations & fraud49 min read
Two classic-ML designs end to end. Recommendations: the two-stage candidate-gen → ranking → rerank funnel, two-tower retrieval, watch-time labels, diversity. Fraud: the rules → ML → review cascade, asymmetric cost, the network moat — with real numbers and the tradeoffs interviewers grade.
- →Domain ML Design
- →Problem Framing and Metrics
6Worked design: enterprise LLM/RAG assistant49 min read
The fastest-growing prompt, worked end to end. Retrieval-conditioned generation over millions of internal docs: hybrid search, rerankers, chunking, grounding and citations, the eval problem without ground-truth labels, guardrails, multi-tenancy and permissioning, the scoring rubric, and the curveballs.
- →Knowledge System Design
- →Problem Framing and Metrics
Skills in this course
- 01Problem Framing and MetricsTurn a product request into measurable ML requirements, constraints, and tradeoffs.
- 02Data and Feature PipelinesDesign labels, features, splits, and serving paths without leakage or skew.
- 03Serving and ScaleChoose serving modes and capacity plans for latency, throughput, and cost.
- 04Monitoring and RecoveryDetect drift and quality loss with staged rollout, rollback, and retraining controls.
- 05Domain ML DesignApply ranking, recommendation, fraud, and review patterns to domain costs.
- 06Knowledge System DesignDesign a secure, evaluated enterprise retrieval assistant.