Lesson 4 of 6 · 46 min

Offline vs online metrics

Offline evaluation is a filter, not a verdict. Why an offline win (NDCG up) can be an online loss (CTR down), the ~97% agreement ceiling Amazon measured, the systematic reasons offline lies (position bias, presentation drift, counterfactuals), and the bridges — interleaving, off-policy estimation — that close the gap.

Your best offline model can be your worst online model

Teams that ship ML rank candidates offline — on a logged dataset, with a metric like NDCG, AUC, or accuracy — and then assume the offline winner will win online. It frequently doesn’t. Kohavi’s most cited puzzle is an experiment where an offline ranking metric improved while online click-through dropped — the offline winner would have shipped harm. The senior mental model: offline evaluation is a winnowing filter, not a launch decision. This lesson is about why the gap exists, how big it is (Amazon measured it), and the bridges that narrow it — and it’s a favourite interview probe because the naive answer (“pick the model with the best offline score”) is wrong.
The structural reason offline and online disagree: offline metrics are computed on the logged policy’s exploration distribution; online metrics are computed on the policy you’re about to ship. The two distributions differ. Your logs record what users did when shown the old system’s slates at the old system’s positions. A new ranker reorders items, surfaces things the old system never showed, and changes the presentation context — so the click distribution re-weights in ways the offline metric, which assumes users would have seen the same slate at the same positions, cannot capture. This is counterfactual measurement, and it is irreducibly hard.
Amazon Science put a number on it: across product-ranking models, offline and online comparisons agree only ~97% of the time — meaning roughly 3% of offline winners silently underperform in production. That 3% is the cumulative effect of position bias (users click top results regardless of relevance), presentation drift (the UI context changed), and the irreducibility of counterfactual estimation. Three percent sounds small until you realise it’s a steady stream of confidently-shipped regressions that an offline-only process can’t catch. Interview angle. “Why not just ship the model with the best offline metric?” → because offline/online agreement caps around 97% and the disagreements are exactly the cases where presentation and position bias dominate — you need an online test for the ship decision.
code
1WHY OFFLINE DISAGREES WITH ONLINE (the ~3% that flips)23  source                what it does                       example4  -------------------   --------------------------------   ----------------------5  position bias         users click top slots regardless   offline NDCG rewards6                        of true relevance                  an order users won't act on7  presentation drift    the UI/context differs from logs   new card layout changes8                                                            click propensity9  counterfactual gap    logs only show the OLD policy's    new ranker surfaces items10                        slates; can't observe new ones     never logged -> unmeasurable11  distribution shift    train/eval data != live traffic    seasonality, new users1213  Amazon Science: offline vs online agree ~97% on product ranking14  -> ~3% of offline winners LOSE online. Offline = filter, online = decision.
Trustworthy Online Controlled Experiments at Large ScaleRonny Kohavi (lecture)

Proxy metrics: useful, and a trap

Because the true online metric is slow or expensive, teams optimise proxy metrics — offline surrogates believed to predict it. A proxy is legitimate only to the degree it’s validated against the online outcome it stands in for. The trap is twofold. First, a proxy that correlates with the goal in observational data may not correlate under the interventions you actually run — the relationship you fit on logs can break exactly where the treatment moves things. Second, the moment you optimise hard against a proxy, Goodhart bites: the model finds ways to move the proxy without the goal (offline NDCG up, real satisfaction flat). The discipline is to tune the offline metric to maximise its agreement with the online metric it predicts, and to keep the online metric as the final arbiter — never to rationalise a persistent gap as noise.

Interleaving: a faster online bridge for ranking

For ranking and recommendation specifically, there’s a powerful middle layer between offline and a full A/B: interleaving. Instead of showing user group A ranker X and group B ranker Y, you merge the outputs of both rankers into a single slate shown to one user in one request (team-draft interleaving), and attribute each click to the ranker that contributed the item. Because both rankers see the same user, same query, same context, the between-user variance collapses and you get a high-sensitivity signal. Netflix reported interleaving is 10–100× faster than a traditional A/B for comparing ranking algorithms; Etsy published the same architecture. The catch: interleaving measures which ranker users prefer, not the downstream business outcome — so you interleave to pick the candidate, then run a conventional A/B for the revenue/retention question.
code
1THE EVALUATION LADDER FOR RANKING (cheap -> trustworthy)23  stage          measures                 cost / speed         role4  ------------    ---------------------    ------------------   -----------------5  offline metric  NDCG/MRR/AUC on logs     cheapest, instant    winnow the pool6   (+ off-policy  IPS/doubly-robust est.   cheap, debiased-ish  rank candidates7    estimation)8  interleaving    which ranker users       online, 10-100x      pick the ranker9                  prefer (same request)    faster than A/B      (Netflix, Etsy)10  A/B test        business outcome         slowest, gold        the SHIP decision11                  (revenue, retention)     standard1213  Use each as a FILTER for the next. Never let an earlier rung make the14  decision the later rung is there to make.
There’s also off-policy / counterfactual estimation — inverse-propensity-scoring (IPS) and doubly-robust estimators that re-weight logged data to estimate how a new policy would have performed, correcting for the fact that the logs came from a different policy. It’s a more honest offline number than a raw metric because it explicitly models the counterfactual, but it has high variance when the new policy diverges far from the logging policy, and it depends on knowing (or estimating) the logging propensities. Treat it as a better filter, not a substitute for the online test. Interview angle. “How would you compare two rankers without a full A/B?” → interleaving for the preference signal (same-request, 10–100× more sensitive), off-policy estimation as a debiased offline filter — then an A/B for the business metric.

The Netflix split: same company, two evaluation strategies

Netflix is the clean case study because it deliberately uses different evaluation methods for different questions. For quality-of-experience (does this change make streaming better?), it runs user-randomised A/B tests with guardrails like stream-start time and rebuffer rate. For ranking-algorithm innovation, it runs interleaving first because the question is "which ranker is better," where same-request sensitivity wins, and only promotes survivors to an A/B for the engagement/retention impact. The lesson for a senior candidate: the evaluation method is chosen by the unit of inference and the question, the same way the randomization unit was in L3 — not by a one-size-fits-all "always A/B" or "trust the offline leaderboard."
paperHow well do offline metrics predict online performance of product ranking models?Amazon SciencearticleInnovating Faster on Personalization Algorithms at Netflix Using InterleavingNetflix Technology BlogarticleFaster ML Experimentation at Etsy with InterleavingEtsy — Code as Craft

Interview prep

The offline–online round tests whether you respect the gap. Interviewers want you to refuse to ship on an offline number alone, to explain why offline lies (position bias, presentation, counterfactuals), and to know the bridges (interleaving, off-policy estimation). Lead with the structural reason, then the number, then the funnel.
  1. 01“Why not ship the model with the best offline metric?” → offline/online agreement caps ~97% (Amazon); the disagreements are exactly where position bias and presentation dominate — online decides.
  2. 02“Why do offline and online disagree?” → logs come from the old policy’s exploration distribution; a new policy changes slates/positions/context, so the click distribution re-weights — a counterfactual problem.
  3. 03“Offline NDCG went up but online CTR went down — what happened?” → Kohavi’s canonical case: the offline metric rewarded an order users didn’t act on under real presentation/position bias.
  4. 04“What’s a proxy metric and its risk?” → an offline surrogate for a slow online goal; risk is the observational correlation breaking under intervention, plus Goodhart when you optimise it hard.
  5. 05“Compare two rankers without a full A/B?” → interleaving (same-request, 10–100× more sensitive, Netflix/Etsy) for the preference signal; off-policy estimation as a debiased offline filter.
  6. 06“What does interleaving NOT give you?” → the business outcome; it tells you which ranker users prefer, so you still need an A/B for revenue/retention.
  7. 07“How do you make an offline metric trustworthy?” → tune it to maximise agreement with the online metric it predicts; validate against historical A/Bs; never rationalise the gap as noise.
  8. 08“Where does each method sit?” → offline + off-policy = winnow; interleaving = pick the ranker; A/B = the ship decision — a funnel, not alternatives.
Going deeper, the follow-ups probe the mechanisms: “how does position bias corrupt an offline click metric?” (users click top slots regardless of relevance, so a metric scored on logged positions rewards orderings that won’t survive re-presentation); “when is off-policy estimation unreliable?” (high variance when the new policy diverges far from the logging policy, and it needs the logging propensities); and “your offline and online metrics keep disagreeing — what do you do?” (don’t dismiss it as noise — instrument the gap, add a "shipped offline-only" diagnostic, and re-tune the offline metric toward online agreement). Name the bias and the bridge in each.

Checkpoint

A candidate ranking model improves offline NDCG by 3% over production. Your manager wants to ship it on that basis. Best response?

AShip it — a 3% NDCG gain is a clear, measurable improvementBUse offline to confirm it’s a viable candidate, then run interleaving and/or an online A/B before shipping, because offline winners lose online ~3% of the timeCSwitch the offline metric to AUC and re-evaluate
Sign up free to answer and see why

Checkpoint

You need to compare two candidate rankers with high sensitivity but limited traffic. Which approach gives the strongest signal per user?

AA standard user-randomised A/B split between the two rankersBPick whichever has the higher offline NDCG and skip the online comparisonCInterleaving — merge both rankers’ results into one slate per request so the same user sees both, collapsing between-user variance (10–100× faster, per Netflix)
Sign up free to answer and see why

Checkpoint

An interviewer asks why offline and online metrics disagree even when the offline pipeline is bug-free. Strongest explanation?

AOffline metrics are computed on the logging policy’s distribution; a new policy changes slates, positions, and context, so the click distribution re-weights — a counterfactual the logs can’t captureBOnline metrics are simply noisier, so they drift away from the true offline valueCThe offline test set is too small; a larger one would make them agree
Sign up free to answer and see why

Checkpoint

Your team optimises hard against an offline proxy metric that historically correlated with revenue. Revenue is now flat in the A/B while the proxy soared. Most likely cause?

AThe A/B is underpowered; rerun with more trafficBGoodhart/surrogation — optimising hard against the proxy moved it without moving the goal; the observational correlation broke under the interventionCRevenue is the wrong online metric; switch to the proxy as the OEC
Sign up free to answer and see why

Checkpoint

You propose off-policy (IPS / doubly-robust) estimation to compare a new policy against logs. When is this least reliable?

AWhen the new policy is nearly identical to the logging policyBWhen the metric is binary rather than continuousCWhen the new policy diverges far from the logging policy, giving high-variance importance weights and poor counterfactual coverage
Sign up free to answer and see why

Could you explain why offline lies, quote the ~97% ceiling, and lay out the offline → interleaving → A/B funnel in an interview?

New to itGetting thereConfident

Takeaways

  • Offline evaluation is a filter, not a verdict — the online A/B makes the ship decision.
  • Offline and online agree only ~97% on ranking (Amazon); the disagreements are position bias, presentation drift, and counterfactuals.
  • Proxy metrics decouple from the goal under intervention and under hard optimisation (Goodhart) — validate and keep the true metric as arbiter.
  • Interleaving is 10–100× more sensitive for ranking (Netflix/Etsy) but measures preference, not business outcome.
  • Treat offline → off-policy → interleaving → A/B as a funnel; never let a cheaper rung make the decision a later rung exists for.

Next: data leakage & validation — the silent ways your offline number is inflated before the online test even runs.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.