Hypothesize: one-sentence "if-then" prediction of the treatment effect's direction. Unit: who is randomized (user, session, geo)? Treatment: what does the new feature or model do? Primary metric tied to business (engagement, conversion, retention) - one primary metric or you're chasing noise. Power: MDE, baseline, alpha = 0.05, power = 0.80 -> sample size n per arm; account for intra-cluster correlation (e.g., ICC=0.05 multiplies the design effect). Duration: cover at least one weekly cycle; prevent peeking with sequential bounds or a pre-registered stopping rule. Guardrails: latency, error rate, paid-event rate. SRM check before reading the result. Run a CUPED variance reduction if traffic is plentiful but effect is small. Ship/no-ship decision tied to confidence interval, not point estimate.