Lesson 1 of 6 · 47 min

Hypotheses & metrics

An experiment is only as good as the question it answers. Sharp hypotheses, the OEC, primary vs guardrail vs diagnostic metrics, MDE as the resolution of your test, and the metric-design failures (Goodhart, surrogation, missing counter-metrics) that ship harm — the way Microsoft, Airbnb, Netflix, and Uber actually set them.

The decision is made before any data arrives

The most expensive A/B mistakes are not statistical — they are pre-statistical. A vague hypothesis, a single optimisation metric with no counter-metric, an effect size nobody agreed was worth detecting: each of these quietly invalidates a perfectly executed test. Microsoft, Airbnb, Netflix, and Uber all converged on the same discipline — lock the question, the metric hierarchy, and the minimum effect worth shipping before a single user is bucketed. This lesson is the part of the experimentation interview where strong candidates separate themselves: not by reciting a t-test, but by designing what to measure.
A real hypothesis is falsifiable, directional, and tied to a mechanism. Exponent’s rubric (used in PM/DS interviews) wants the exact form: “If we [change X], then [metric Y] will move by [magnitude Z] because [user-behaviour rationale].” Compare “we want to test the new checkout” (a wish) with “moving payment to a single page will lift checkout-completion by ≥2% because it removes a step where 8% of mobile users currently drop” (a hypothesis). The second commits you to a metric, a direction, an effect size, and a causal story you can be wrong about. Interview angle. The single most common low-score signal across every prep source is failing to state the treatment and the hypothesis before jumping to analysis.
The ultimate guide to A/B testingRonny Kohavi / Lenny’s Podcast

The OEC: one number the org agrees to optimise

Microsoft’s experimentation platform (ExP) gave the industry the Overall Evaluation Criterion (OEC): a single, agreed success metric that decides whether a change ships. The OEC exists for an organisational reason as much as a statistical one — teams need one number to align to, fixed at design time so no one can pick the favourable metric post-hoc. Airbnb states it plainly: “bookings and nights booked are a typical OEC for the company.” Netflix optimises a quality-of-experience engagement metric; LinkedIn, premium subscriptions; Uber, revenue per session. The senior move is treating OEC selection as upstream code, not analytics decoration.
The hard part of an OEC is that it must be a short-term measurable that predicts long-term value. Daily active sessions is measurable today; lifetime value is what you actually care about. A good OEC is a surrogate chosen because it correlates with the long-term goal and the treatment cannot easily game it. This is why Kohavi insists the OEC be debated and validated against historical experiments before it is trusted — a surrogate that diverges from the true goal under the treatments you run is worse than useless.

The metric taxonomy: primary, guardrail, diagnostic

Every mature platform describes the same three-tier hierarchy. Primary (OEC): the one metric that decides ship / no-ship. Guardrail (counter-)metrics: 2–4 metrics that block the ship if they regress, regardless of how good the primary looks. Diagnostic / secondary: an unlimited bag that explains the result but never overrules the OEC. Uber formalised this as Primary / Guardrail / DAS (Diagnostic, Auxiliary, Secondary) and runs it across 1,000+ concurrent experiments, with experimentation gating ~95% of revenue-impacting feature launches.
code
1THE METRIC HIERARCHY (lock all of this before bucketing a single user)23  tier          role                                example (checkout test)4  -----------   ---------------------------------   --------------------------5  PRIMARY/OEC   decides ship / no-ship (ONE)        checkout-completion rate6  GUARDRAIL     blocks ship if it regresses         page-load p95, refund rate,7                (2-4, tied to risk vectors)         support-ticket rate, crashes8  DIAGNOSTIC    explains movement, never overrules  CTR by step, time-to-pay,9  /SECONDARY    (unlimited)                         add-to-cart, scroll depth1011  Rule: a guardrail regression is a launch blocker even if the OEC is a big win.12  A metric that moves directionally but not in magnitude is a HYPOTHESIS, not a result.
Guardrails exist because any single optimisation metric is subject to Goodhart’s Law — “when a measure becomes a target, it ceases to be a good measure.” Kohavi’s canonical case: a Bing experiment lifted click-through rate while sessions-per-user dropped. More clicks looked like success, but users were clicking because results were worse and they had to work harder to find the answer. The team only avoided shipping harm because sessions-per-user was a pre-registered guardrail. Interview angle. “Your treatment lifts the primary metric but harms a guardrail — what do you do?” The strong answer names the risk vector, investigates the mechanism, and considers shipping only to the segment where the guardrail held — never “it’s a win, ship it.”
At scale this becomes infrastructure, not judgement. Airbnb’s Experiment Reporting Framework (ERF) computes ~2,500 distinct metrics per day against ~50,000 experiment×metric combinations — a typed catalog so every team reads the same metric definition, rather than each analyst hand-rolling “conversion.” The lesson for a senior candidate: ad-hoc metric creation by every team is what kills trust in an experimentation program; a governed metric catalog is what scales it.

Choosing the metric: sensitivity, surrogation, dilution

Not all valid metrics are good experiment metrics. Three properties decide usefulness. Sensitivity: can the metric actually move within your traffic and runtime? Revenue is the goal but is high-variance and slow; a leading indicator (add-to-cart, day-1 retention) often detects the same effect with far more power. Surrogation risk: the closer you optimise a surrogate, the more the treatment finds ways to move the surrogate without the goal — so surrogates need guardrails on the true outcome. Dilution: if only a slice of users are exposed to the change, a metric measured over all users dilutes the effect toward zero; measure on the triggered population (users who actually hit the changed surface) to recover power.
Triggering is a senior-level subtlety with real consequences. If a new banner shows to 3% of sessions but you analyse the metric over 100% of sessions, you have diluted a (say) 10% local lift to 0.3% globally — likely undetectable. The fix is trigger-based analysis: define the trigger condition, count only users who met it, and analyse the counterfactual triggered set in control. The trap, which we revisit under leakage, is that triggering can itself introduce bias if the trigger depends on post-treatment behaviour. Interview angle. “The feature only affects checkout but you measure site-wide revenue — what’s wrong?” → dilution; analyse the triggered population.
code
1DILUTION: why you analyse the TRIGGERED population, not everyone23  exposure to change ......... 3% of sessions hit the new banner4  true local lift ............ +10% on the triggered slice5  measured site-wide ......... 0.03 * 10% = +0.3%  <- swamped by variance, looks "flat"67  fix: count only users who met the trigger condition; compare to the8       counterfactual triggered set in control. Power is restored because9       the denominator is the population the treatment could actually affect.1011  caution: the trigger must NOT depend on post-randomization outcomes12           (that re-introduces selection bias -- see the leakage lesson).

MDE: the resolution of your experiment

The Minimum Detectable Effect (MDE) is the smallest true effect your test can reliably detect at your chosen power — think of it as the resolution of the experiment. It is set at design time and reported alongside any non-significant result. The formula (full derivation next lesson): MDE ∝ (z1−α/2 + z1−β) × √(σ²/n). Three levers move it — your significance level α, your power 1−β, and crucially the variance σ² of the metric relative to the sample you can gather. A small MDE achieved purely by huge n is not the same as one achieved by variance reduction; always report σ and the standardised effect MDE/σ.
The MDE must be set by the business, not the statistics. The question is “what is the smallest lift that would change our decision or pay for the engineering cost?” — not “what lift can we detect.” If the business needs to detect a 1% lift but your traffic only resolves 5% in a feasible runtime, the experiment is infeasible as designed, and the senior move is to escalate that before launch: raise the MDE bar, pick a more sensitive metric, or apply variance reduction (CUPED, L3). SplitMetrics’ worked example is worth memorising: a 20% baseline conversion rate, detecting MDE = 5%, needs ~11,141 total conversions — a concrete anchor for “how big.”

Significance vs importance: the decision layer

The final discipline: statistical significance is necessary but not sufficient for a launch decision. At 100M users a 0.01% lift is statistically detectable and operationally trivial. The decision layer must combine the effect size and its CI, the guardrails, novelty-corrected estimates (L3), the expected duration of the benefit, and plain business judgement. Kohavi’s framing: “getting numbers is easy; getting numbers you can trust is hard” — and even trusted numbers are an input to a decision, not the decision. The weak interview answer ends at “it’s stat-sig”; the strong one ends at “ship / iterate / no-ship, and here is the upside and the risk.”
Pick the OEC and at least two guardrails before any user is bucketed. — the recurring discipline across Microsoft ExP, Airbnb, Netflix, and Uber: the metric design is the experiment; the test is just arithmetic on top of it.
docsTrustworthy Online Controlled Experiments: A Practical Guide to A/B TestingKohavi, Tang & Xu (the canonical text)paperTrustworthy online controlled experiments: five puzzling outcomes explainedKohavi et al. (KDD 2012)articleScaling Airbnb’s Experimentation Platform (ERF)Airbnb EngineeringarticleGuardrail metrics: the complete guide to balanced product experimentationMixpanel

Interview prep

The metrics-and-hypothesis round tests whether you design what to measure with the same rigour you bring to the statistics. Interviewers at Uber, Airbnb, Meta, and Google probe a fixed set: can you write a falsifiable hypothesis, name an OEC and defend it as a long-term surrogate, build a guardrail set tied to risk vectors, and reason about MDE and dilution. Lead with the mechanism and the production consequence — never just a metric name.
  1. 01“How do you design an A/B test for feature X?” → state purpose → “if X then Y by Z because…” hypothesis → OEC + 2-4 guardrails + diagnostics → unit of randomization → power/MDE → decision rule.
  2. 02“What’s an OEC and how do you pick one?” → one agreed success metric, frozen at design time; a short-term measurable that predicts long-term value and the treatment can’t easily game.
  3. 03“Primary vs guardrail vs driver metrics?” → primary decides ship; guardrails block ship on regression (Goodhart insurance); drivers/diagnostics explain but never overrule.
  4. 04“Your OEC went up but a guardrail dropped — ship?” → no by default; investigate the mechanism (Bing CTR-up/sessions-down), consider segment-only ship, don’t Goodhart.
  5. 05“What is MDE and who sets it?” → smallest effect detectable at target power; set by the business (smallest lift worth shipping), then check feasibility vs traffic.
  6. 06“The feature only touches checkout but revenue looks flat — why?” → dilution; analyse the triggered population, not all users, to restore power.
  7. 07“Why not just optimise revenue directly?” → high variance + slow; use a sensitive leading indicator as OEC with revenue as a guardrail.
  8. 08“Is a statistically significant result enough to launch?” → no — stat-sig is the entry ticket, not the verdict; combine effect size, CI, guardrails, durability, and business value.
To go deeper, expect the follow-ups that separate “read a blog” from “ran experiments”: “what short-term metric would you trust as a surrogate for retention, and how would you validate it?” (correlate the surrogate with the long-term outcome across historical experiments; reject it if treatments move them in opposite directions); “how many guardrails is too many?” (each guardrail you test adds multiplicity — keep the set tied to genuine risk vectors and correct for it, L2); and “your metric moved but only in one segment — is that a win?” (heterogeneous effect; pre-register the segmentation, beware Simpson’s paradox, L3). In each, name the metric, the risk, and the decision.

Checkpoint

A PM says “let’s test the redesigned feed and see if engagement improves.” You’re the DS in the room. What’s the strongest first move?

AReframe it as a falsifiable hypothesis with an OEC, an expected direction/magnitude, and the guardrails that would block a shipBPick the largest available sample so the test is well poweredCRun it on all traffic and compare engagement week-over-week
Sign up free to answer and see why

Checkpoint

Your treatment lifts click-through rate by a significant +4%, but sessions-per-user drops by 2% (a pre-registered guardrail). Best read?

AShip — the primary metric (CTR) is significant and positiveBDrop CTR as the OEC and use sessions-per-user instead, then re-decideCTreat the guardrail regression as a blocker, investigate why CTR rose while sessions fell, and consider shipping only where the guardrail held
Sign up free to answer and see why

Checkpoint

A new help widget is shown to ~4% of sessions. Measured over all sessions, the conversion lift looks flat and non-significant. Most likely issue?

AThe widget genuinely has no effect; conclude no-shipBDilution — measuring over all sessions swamps a local effect; analyse the triggered population that actually saw the widgetCThe significance threshold is too strict; relax α to 0.10
Sign up free to answer and see why

Checkpoint

The business wants to detect a 1% lift in conversion, but at current traffic your MDE for a feasible 2-week run is ~4%. What’s the right call?

ARun the 2-week test anyway and report whatever p-value comes outBLower α to make the test more sensitiveCEscalate before launch: raise the MDE bar, switch to a more sensitive metric, or apply variance reduction — the test is infeasible as designed
Sign up free to answer and see why

Checkpoint

Leadership wants the OEC to be 12-month lifetime value so the metric “matches what we care about.” Your reaction?

AAgree — always optimise the metric closest to the true business goalBPick a sensitive short-term surrogate that predicts LTV, validate it against historical experiments, and keep LTV-linked guardrailsCUse revenue-per-session because it’s already in the metric catalog
Sign up free to answer and see why

Could you walk into a design review and define the hypothesis, OEC, guardrails, and MDE for a feature — and defend each choice?

New to itGetting thereConfident

Takeaways

  • The expensive mistakes are pre-statistical: lock the hypothesis, OEC, guardrails, and MDE before bucketing anyone.
  • A hypothesis is falsifiable and directional: “if X then Y by Z because…” — not “let’s see if it works.”
  • OEC = one short-term measurable that predicts long-term value; guardrails are Goodhart insurance and block the ship on regression.
  • MDE is the resolution of your test, set by the business; report it (and the CI) with every null result.
  • Measure the triggered population to beat dilution; a significant result is an input to the launch decision, not the decision itself.

Next: p-values, power, sample size, and confidence intervals — the one equation read four ways, done correctly.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.