Put it together: take a real feature from hypothesis to ship/no-ship decision, defending every choice — OEC and guardrails, randomization unit, power and MDE, the monitoring plan, the analysis (CUPED, SRM, segmentation, novelty), and the curveballs interviewers throw at the end. The full experimentation round, walked end to end.
The whiteboard question that is the whole job
The capstone of every experimentation interview — and the actual day job — is one prompt: “design and analyse an A/B test for <feature>, end to end.” A strong answer is a loop, walked out loud: hypothesis → metrics → unit → power/MDE → monitoring → analysis → decision, with the failure mode named at each step and a ship/iterate/no-ship at the end. This lesson runs that loop on a concrete case, then fields the curveballs (mid-test dip, p = 0.06, Simpson’s paradox, novelty, outliers) that separate someone who read a blog from someone who has shipped experiments.
The case: a B2C app wants to test a redesigned onboarding flow, hypothesising it lifts day-7 retention. We’ll carry it through every decision. The skeleton that scores across every prep source (Reddit r/datascience, Exponent, Datainterview, IGotAnOffer) is the same six beats — Hypothesis → Metrics → Design → Power → Analyse → Decide. Memorise the loop; name one concrete choice and one failure mode per beat. The weak version skips beats and ends at “it’s significant”; the strong version defends each choice and ends at a decision with quantified upside and risk.
Hypothesis. “If we replace the 4-step onboarding with a single guided screen, day-7 retention rises by ≥1.5% because we cut the step where ~12% of new users currently abandon.” Falsifiable, directional, mechanism-backed (L1). Metric hierarchy. Primary/OEC: day-7 retention of new users (a short-term measurable that predicts long-term LTV, validated against past onboarding experiments). Guardrails: onboarding-completion rate, day-1 crash-free rate, support-ticket rate, and time-to-first-action — each tied to a way the redesign could backfire. Diagnostics: per-step drop-off, time-on-screen. Trigger. Only new users hitting onboarding enter the experiment, analysed on that triggered population to avoid dilution (L1).
code
1THE DESIGN ONE-PAGER (lock before launch)23 hypothesis single-screen onboarding -> +>=1.5% D7 retention (cut 12% drop step)4 population NEW users who trigger onboarding (analyse the triggered set)5 OEC day-7 retention (new users) -- validated surrogate for LTV6 guardrails onboarding-completion, D1 crash-free, support-tickets, time-to-action7 unit user, hash(user_id + experiment_id) mod 1000 (SUTVA holds; no network)8 split 50/50 (max power for fixed N)9 alpha/power 0.05 two-sided / 0.8010 MDE 1.5% relative (the smallest business-relevant lift)11 variance CUPED on pre-period engagement covariate (expect ~30-50% reduction)12 duration >=14 days (cover weekly cycle + outlast novelty), fixed horizon13 monitoring group-sequential alpha-spending OR fixed-horizon; SRM chi-square daily14 analysis CUPED-adjusted t-test + CI on the difference; pre-registered segments
Beats 3–4: unit, power, and the runtime reality
Unit. User-level hashing with the experiment-id salt (L3); onboarding is per-user with no spillover, so SUTVA holds and we don’t need clusters or switchbacks. 50/50 split for maximum power. Power/MDE. Fix α = 0.05, power 0.80; the 1.5% relative MDE on the retention base rate gives a required n per arm via n = (z+z)²σ²/Δ² (L2). Convert to users via the new-user inflow and to days via daily signups — and if that runtime exceeds a feasible window, escalate (L1): apply CUPED to cut σ², or raise the MDE bar. We pre-commit to ≥14 days regardless of when significance appears, to cover a full weekly cycle and outlast the novelty window (L3).
Variance reduction. Each new user has little pre-period history, but we have some signal (acquisition channel, device, first-session depth) — CUPED or CUPAC on those covariates buys whatever (1−ρ²) reduction the correlation supports (L3); even a 30% cut meaningfully shortens runtime. Monitoring. Either a fixed horizon looked at once, or group-sequential alpha-spending if we want a licensed early-stop for a clear win/loss (Spotify’s default when max N is known, L3) — and a chi-square SRM check every day as the canary (L5). Interview angle. “How long do you run it?” → derive n from power, convert to days, take the max of (powered runtime, one weekly cycle, the novelty window) — and never stop early on a fixed-horizon test just because it crossed 0.05.
Beat 5: analysis, and the things that go wrong
Before reading the primary metric, run the trust checks: SRM chi-square (kill if p < 0.001 vs the planned split — L5), confirm the triggered population matches across arms, and verify no instrumentation gaps. Then analyse on the CUPED-adjusted metric with the CI on the difference (L2), variance clustered at the user level (L5). Report effect size with its CI, not a bare p-value. Run the pre-registered segmentation (device, channel, geo) to check for heterogeneous effects — and specifically to defend against Simpson’s paradox, Kohavi’s 4th puzzle, where the aggregate sign and every segment’s sign can disagree because the arms have a different mix on a confounder.
Three more analysis hazards the research flags. Novelty / primacy (L3): a redesign often shows a transient bump (curiosity) or dip (experienced users re-learning) in the first days; run a treatment×day interaction and don’t ship until the slope flattens — Kohavi’s “obviously good” feature that tested negative because users were still learning the old UI is the cautionary tale. Multiple comparisons: testing many metrics/segments inflates false positives toward 1 − 0.95^N (~40% at N=10); apply Benjamini-Hochberg (FDR) across the guardrail/secondary set. Outliers: Kohavi’s Amazon case — one user buying $2,500 skews a revenue metric over 100k users; cap contribution at the 99th percentile or use a trimmed mean before trusting the result.
code
1THE ANALYSIS CHECKLIST (run in this order; each maps to a lesson)23 1. SRM chi-square realized vs planned split; p<0.001 -> STOP, find bug (L5)4 2. CUPED adjustment Y - theta*(X - mu_X); theta pooled across arms (L3)5 3. cluster variance at the USER level, not session/observation (L5)6 4. CI on the difference report effect size + CI, not a bare p-value (L2)7 5. pre-registered segments device/channel/geo; check for Simpson reversal (L3)8 6. novelty/primacy treatment x day interaction; ship after slope flattens (L3)9 7. multiplicity Benjamini-Hochberg (FDR) across guardrails/secondaries (L1/L2)10 8. outliers cap at 99th pct / trimmed mean before trusting revenue (L2)
Beat 6: the decision, and the curveballs
The decision is ship / iterate / no-ship, combining the effect size and CI, the guardrails (any regression blocks), the novelty-corrected estimate, and the business value — the stat-sig result is one input among these, never the whole decision (L1). The curveballs the research catalogues, with the strong read: mid-test dip after a strong start → usually novelty decay; wait/extend, don’t kill on the early bump. p = 0.06 at α = 0.05 → you fail to reject; report the CI and decide between extending, re-running, or accepting a quantified false-positive risk — not “almost significant, ship.” Significant primary but a harmed guardrail → investigate the mechanism; consider shipping only to the segment where the guardrail held. Significant negative primary → assess cost of inaction vs partial rollback; iterate the hypothesis.
The end-to-end design question is the experimentation round. Interviewers grade whether you walk the full loop without dropping a beat, name the failure mode at each step, run trust checks before believing a result, and land a business decision. Lead with the loop, defend each choice, and finish with ship/iterate/no-ship and the upside.
01“Design an A/B test for <feature>, end to end.” → Hypothesis → Metrics (OEC + guardrails) → Unit → Power/MDE → Monitoring → Analyse (trust checks) → Decide; one choice + one failure mode per beat.
02“How long do you run it?” → derive n from power, convert to days, then take the max of (powered runtime, one weekly cycle, the novelty window); fixed horizon or group-sequential, no naive early stop.
03“Mid-test the metric dips after a strong start — kill it?” → usually novelty decay; run a treatment×day interaction and wait for the slope to flatten before deciding.
04“The result is p = 0.06 — what now?” → fail to reject at 0.05; report the CI, then extend, re-run, or accept a quantified false-positive risk — never “almost significant, ship.”
05“Aggregate says ship but every segment says iterate — what’s happening?” → Simpson’s paradox from a different arm mix on a confounder; trust the pre-registered segmented analysis.
06“Before you trust the win, what do you check?” → SRM chi-square, triggered-population match, user-level clustering, outlier capping, and a novelty check — symmetric skepticism on wins and losses.
07“You can’t randomise users (network effects / policy) — now what?” → switchback, geo holdout, or a quasi-experiment (diff-in-diff, synthetic control, RD) with pre-registered parallel-trends.
08“What makes this a ship vs a no-ship?” → significant + guardrails intact + novelty-corrected + business value clears the bar; the p-value alone never decides it.
Going deeper, the hardest follow-ups stress your judgement: “a GM wants to ship on day 3 because it’s already significant — defend your position” (interim looks on a fixed-horizon test inflate α and the lift; either commit to the horizon or run an always-valid design — and explain novelty in plain language); “you can’t run a clean A/B at all — what’s your plan?” (name why — network effects, ethics, one-shot launch — then the matching alternative with its identifying assumption); and “the experiment is a learning but not a launch — how do you frame that?” (a null or negative result that de-risks a roadmap bet is a win; report the MDE you ruled out). In every answer, end at a decision.
Checkpoint
Your onboarding test shows a strong day-1 retention bump that fades over the first week toward a small steady lift. Best interpretation for the ship decision?
ANovelty effect — the early bump is curiosity decaying to the true effect; run a treatment×day interaction and decide on the post-novelty steady stateBThe treatment is wearing off and will go negative — kill it nowCShip immediately on the day-1 bump before the effect disappears
Aggregate results say the redesign wins, but when you segment by device (pre-registered), it loses on both iOS and Android. What’s going on and what do you trust?
AThe aggregate is correct because it has the most data; ignore the segmentsBSimpson’s paradox — a different device mix across arms reverses the aggregate sign; trust the within-segment (consistent) resultCRandom noise — rerun until the aggregate and segments agree
The experiment ends at p = 0.06 against α = 0.05, with a CI on the lift of [−0.2%, +2.8%]. A VP says “basically significant, let’s ship.” Strongest response?
AAgree — 0.06 is close enough to 0.05 to call it a winBLower α to 0.10 retroactively so the result is significantCWe fail to reject at 0.05 and the CI includes 0; present the effect size and tradeoff, then decide to extend, re-run, or ship accepting a quantified false-positive risk
The primary metric (revenue per user) is up and significant, but you haven’t looked at anything else yet. What’s the right next step before declaring a win?
ADeclare the win — the primary metric is significant and positiveBRun the trust checks — SRM, user-level clustering, 99th-percentile outlier cap, guardrails, and segmentation — applying the same skepticism you would to a lossCImmediately roll out to 100% to capture the revenue before it disappears
Leadership wants to test a pricing-algorithm change in a two-sided marketplace where a clean user-level A/B isn’t valid. What do you propose?
AA switchback or geo holdout (or a pre-registered quasi-experiment), because network effects break SUTVA for a user-level splitBA standard 50/50 user-level A/B with a larger sample to overcome the noiseCSkip experimentation and ship to everyone, monitoring revenue after