Lessons

1Hypotheses & metrics47 min read

An experiment is only as good as the question it answers. Sharp hypotheses, the OEC, primary vs guardrail vs diagnostic metrics, MDE as the resolution of your test, and the metric-design failures (Goodhart, surrogation, missing counter-metrics) that ship harm — the way Microsoft, Airbnb, Netflix, and Uber actually set them.

  • →Hypotheses and Metrics
Read lesson
2Significance & power48 min read

p-values, statistical power, the sample-size equation read four ways, and confidence intervals — what each actually means (and what the ASA says it does not), why power is set by variance not effect, and the misinterpretations that end interview rounds and ship bad launches.

  • →Statistical Inference
Read lesson
3A/B test design & the peeking problem50 min read

The randomization unit is an architectural decision; SUTVA breaks under network effects; peeking inflates false positives 2-3x; and CUPED halves variance for free. Microsoft hash bucketing, DoorDash switchbacks, Meta cluster randomization, Spotify’s sequential-testing choice, and Bing’s CUPED result — the design decisions that make or break a trustworthy experiment.

  • →Experiment Architecture
  • →Statistical Inference
Read lesson
4Offline vs online metrics46 min read

Offline evaluation is a filter, not a verdict. Why an offline win (NDCG up) can be an online loss (CTR down), the ~97% agreement ceiling Amazon measured, the systematic reasons offline lies (position bias, presentation drift, counterfactuals), and the bridges — interleaving, off-policy estimation — that close the gap.

  • →Experiment Architecture
  • →Validity and Ship Decisions
Read lesson
5Data leakage & validation47 min read

The reason your model scored 0.95 offline and 0.70 in production. Target leakage, temporal leakage, group/cluster leakage, and the cross-validation splits that quietly cheat — plus the experiment-side cousins (SRM, clustered variance) that invalidate a test the same way. The single biggest source of overstated offline numbers.

  • →Validity and Ship Decisions
Read lesson
6Capstone: design an experiment49 min read

Put it together: take a real feature from hypothesis to ship/no-ship decision, defending every choice — OEC and guardrails, randomization unit, power and MDE, the monitoring plan, the analysis (CUPED, SRM, segmentation, novelty), and the curveballs interviewers throw at the end. The full experimentation round, walked end to end.

  • →Hypotheses and Metrics
  • →Experiment Architecture
  • →Validity and Ship Decisions
Read lesson

Skills in this course

  1. 01Hypotheses and MetricsDefine a testable hypothesis, primary metric, guardrails, and detectable effect.
  2. 02Statistical InferenceInterpret uncertainty, power, significance, and confidence intervals correctly.
  3. 03Experiment ArchitectureChoose randomization, variance reduction, and sequential methods that match interference and traffic.
  4. 04Validity and Ship DecisionsFind leakage and experiment invalidity, then make evidence-based ship decisions.