Lessons

1The ML system design framework46 min read

The spine every ML system design round rides — framing, success metrics, data, model, serving, monitoring — how the 60 minutes are actually budgeted, the six dimensions interviewers grade, and the three cardinal sins that silently sink strong engineers by minute 10.

  • →Problem Framing and Metrics
Read lesson
2Data, features & embedding pipelines47 min read

Where the label actually comes from and why “click = positive” is the most common wrong default; feature stores and the offline/online split that creates training-serving skew; embeddings as a pipeline, not a model call; batch vs streaming features; and the leakage that fakes a great offline score.

  • →Data and Feature Pipelines
Read lesson
3Serving, latency & cost at scale48 min read

The four serving modes and how to pick one; the LLM serving stack that actually wins — continuous batching, PagedAttention, prefix caching — with the real multipliers; GPU economics and the two-number latency SLO (TTFT and ITL); model routing; and the autoscaling lag that bounds it all.

  • →Serving and Scale
Read lesson
4Monitoring, drift & retraining46 min read

The monitoring story is graded as heavily as the modeling — “watch accuracy on a hold-out” is the wrong answer. Online vs offline metrics, the drift taxonomy and which statistical test to use, shadow vs canary, the three layers of rollback, the retraining loop, and a worked incident scenario.

  • →Monitoring and Recovery
  • →Problem Framing and Metrics
Read lesson
5Worked design: recommendations & fraud49 min read

Two classic-ML designs end to end. Recommendations: the two-stage candidate-gen → ranking → rerank funnel, two-tower retrieval, watch-time labels, diversity. Fraud: the rules → ML → review cascade, asymmetric cost, the network moat — with real numbers and the tradeoffs interviewers grade.

  • →Domain ML Design
  • →Problem Framing and Metrics
Read lesson
6Worked design: enterprise LLM/RAG assistant49 min read

The fastest-growing prompt, worked end to end. Retrieval-conditioned generation over millions of internal docs: hybrid search, rerankers, chunking, grounding and citations, the eval problem without ground-truth labels, guardrails, multi-tenancy and permissioning, the scoring rubric, and the curveballs.

  • →Knowledge System Design
  • →Problem Framing and Metrics
Read lesson

Skills in this course

  1. 01Problem Framing and MetricsTurn a product request into measurable ML requirements, constraints, and tradeoffs.
  2. 02Data and Feature PipelinesDesign labels, features, splits, and serving paths without leakage or skew.
  3. 03Serving and ScaleChoose serving modes and capacity plans for latency, throughput, and cost.
  4. 04Monitoring and RecoveryDetect drift and quality loss with staged rollout, rollback, and retraining controls.
  5. 05Domain ML DesignApply ranking, recommendation, fraud, and review patterns to domain costs.
  6. 06Knowledge System DesignDesign a secure, evaluated enterprise retrieval assistant.