Offline evaluation is a filter, not a verdict. Why an offline win (NDCG up) can be an online loss (CTR down), the ~97% agreement ceiling Amazon measured, the systematic reasons offline lies (position bias, presentation drift, counterfactuals), and the bridges — interleaving, off-policy estimation — that close the gap.
Your best offline model can be your worst online model
Teams that ship ML rank candidates offline — on a logged dataset, with a metric like NDCG, AUC, or accuracy — and then assume the offline winner will win online. It frequently doesn’t. Kohavi’s most cited puzzle is an experiment where an offline ranking metric improved while online click-through dropped — the offline winner would have shipped harm. The senior mental model: offline evaluation is a winnowing filter, not a launch decision. This lesson is about why the gap exists, how big it is (Amazon measured it), and the bridges that narrow it — and it’s a favourite interview probe because the naive answer (“pick the model with the best offline score”) is wrong.
The structural reason offline and online disagree: offline metrics are computed on the logged policy’s exploration distribution; online metrics are computed on the policy you’re about to ship. The two distributions differ. Your logs record what users did when shown the old system’s slates at the old system’s positions. A new ranker reorders items, surfaces things the old system never showed, and changes the presentation context — so the click distribution re-weights in ways the offline metric, which assumes users would have seen the same slate at the same positions, cannot capture. This is counterfactual measurement, and it is irreducibly hard.
Amazon Science put a number on it: across product-ranking models, offline and online comparisons agree only ~97% of the time — meaning roughly 3% of offline winners silently underperform in production. That 3% is the cumulative effect of position bias (users click top results regardless of relevance), presentation drift (the UI context changed), and the irreducibility of counterfactual estimation. Three percent sounds small until you realise it’s a steady stream of confidently-shipped regressions that an offline-only process can’t catch. Interview angle. “Why not just ship the model with the best offline metric?” → because offline/online agreement caps around 97% and the disagreements are exactly the cases where presentation and position bias dominate — you need an online test for the ship decision.
code
1WHY OFFLINE DISAGREES WITH ONLINE (the ~3% that flips)23 source what it does example4 ------------------- -------------------------------- ----------------------5 position bias users click top slots regardless offline NDCG rewards6 of true relevance an order users won't act on7 presentation drift the UI/context differs from logs new card layout changes8 click propensity9 counterfactual gap logs only show the OLD policy's new ranker surfaces items10 slates; can't observe new ones never logged -> unmeasurable11 distribution shift train/eval data != live traffic seasonality, new users1213 Amazon Science: offline vs online agree ~97% on product ranking14 -> ~3% of offline winners LOSE online. Offline = filter, online = decision.
Because the true online metric is slow or expensive, teams optimise proxy metrics — offline surrogates believed to predict it. A proxy is legitimate only to the degree it’s validated against the online outcome it stands in for. The trap is twofold. First, a proxy that correlates with the goal in observational data may not correlate under the interventions you actually run — the relationship you fit on logs can break exactly where the treatment moves things. Second, the moment you optimise hard against a proxy, Goodhart bites: the model finds ways to move the proxy without the goal (offline NDCG up, real satisfaction flat). The discipline is to tune the offline metric to maximise its agreement with the online metric it predicts, and to keep the online metric as the final arbiter — never to rationalise a persistent gap as noise.
Interleaving: a faster online bridge for ranking
For ranking and recommendation specifically, there’s a powerful middle layer between offline and a full A/B: interleaving. Instead of showing user group A ranker X and group B ranker Y, you merge the outputs of both rankers into a single slate shown to one user in one request (team-draft interleaving), and attribute each click to the ranker that contributed the item. Because both rankers see the same user, same query, same context, the between-user variance collapses and you get a high-sensitivity signal. Netflix reported interleaving is 10–100× faster than a traditional A/B for comparing ranking algorithms; Etsy published the same architecture. The catch: interleaving measures which ranker users prefer, not the downstream business outcome — so you interleave to pick the candidate, then run a conventional A/B for the revenue/retention question.
code
1THE EVALUATION LADDER FOR RANKING (cheap -> trustworthy)23 stage measures cost / speed role4 ------------ --------------------- ------------------ -----------------5 offline metric NDCG/MRR/AUC on logs cheapest, instant winnow the pool6 (+ off-policy IPS/doubly-robust est. cheap, debiased-ish rank candidates7 estimation)8 interleaving which ranker users online, 10-100x pick the ranker9 prefer (same request) faster than A/B (Netflix, Etsy)10 A/B test business outcome slowest, gold the SHIP decision11 (revenue, retention) standard1213 Use each as a FILTER for the next. Never let an earlier rung make the14 decision the later rung is there to make.
There’s also off-policy / counterfactual estimation — inverse-propensity-scoring (IPS) and doubly-robust estimators that re-weight logged data to estimate how a new policy would have performed, correcting for the fact that the logs came from a different policy. It’s a more honest offline number than a raw metric because it explicitly models the counterfactual, but it has high variance when the new policy diverges far from the logging policy, and it depends on knowing (or estimating) the logging propensities. Treat it as a better filter, not a substitute for the online test. Interview angle. “How would you compare two rankers without a full A/B?” → interleaving for the preference signal (same-request, 10–100× more sensitive), off-policy estimation as a debiased offline filter — then an A/B for the business metric.
The Netflix split: same company, two evaluation strategies
Netflix is the clean case study because it deliberately uses different evaluation methods for different questions. For quality-of-experience (does this change make streaming better?), it runs user-randomised A/B tests with guardrails like stream-start time and rebuffer rate. For ranking-algorithm innovation, it runs interleaving first because the question is "which ranker is better," where same-request sensitivity wins, and only promotes survivors to an A/B for the engagement/retention impact. The lesson for a senior candidate: the evaluation method is chosen by the unit of inference and the question, the same way the randomization unit was in L3 — not by a one-size-fits-all "always A/B" or "trust the offline leaderboard."
The offline–online round tests whether you respect the gap. Interviewers want you to refuse to ship on an offline number alone, to explain why offline lies (position bias, presentation, counterfactuals), and to know the bridges (interleaving, off-policy estimation). Lead with the structural reason, then the number, then the funnel.
01“Why not ship the model with the best offline metric?” → offline/online agreement caps ~97% (Amazon); the disagreements are exactly where position bias and presentation dominate — online decides.
02“Why do offline and online disagree?” → logs come from the old policy’s exploration distribution; a new policy changes slates/positions/context, so the click distribution re-weights — a counterfactual problem.
03“Offline NDCG went up but online CTR went down — what happened?” → Kohavi’s canonical case: the offline metric rewarded an order users didn’t act on under real presentation/position bias.
04“What’s a proxy metric and its risk?” → an offline surrogate for a slow online goal; risk is the observational correlation breaking under intervention, plus Goodhart when you optimise it hard.
05“Compare two rankers without a full A/B?” → interleaving (same-request, 10–100× more sensitive, Netflix/Etsy) for the preference signal; off-policy estimation as a debiased offline filter.
06“What does interleaving NOT give you?” → the business outcome; it tells you which ranker users prefer, so you still need an A/B for revenue/retention.
07“How do you make an offline metric trustworthy?” → tune it to maximise agreement with the online metric it predicts; validate against historical A/Bs; never rationalise the gap as noise.
08“Where does each method sit?” → offline + off-policy = winnow; interleaving = pick the ranker; A/B = the ship decision — a funnel, not alternatives.
Going deeper, the follow-ups probe the mechanisms: “how does position bias corrupt an offline click metric?” (users click top slots regardless of relevance, so a metric scored on logged positions rewards orderings that won’t survive re-presentation); “when is off-policy estimation unreliable?” (high variance when the new policy diverges far from the logging policy, and it needs the logging propensities); and “your offline and online metrics keep disagreeing — what do you do?” (don’t dismiss it as noise — instrument the gap, add a "shipped offline-only" diagnostic, and re-tune the offline metric toward online agreement). Name the bias and the bridge in each.
Checkpoint
A candidate ranking model improves offline NDCG by 3% over production. Your manager wants to ship it on that basis. Best response?
AShip it — a 3% NDCG gain is a clear, measurable improvementBUse offline to confirm it’s a viable candidate, then run interleaving and/or an online A/B before shipping, because offline winners lose online ~3% of the timeCSwitch the offline metric to AUC and re-evaluate
You need to compare two candidate rankers with high sensitivity but limited traffic. Which approach gives the strongest signal per user?
AA standard user-randomised A/B split between the two rankersBPick whichever has the higher offline NDCG and skip the online comparisonCInterleaving — merge both rankers’ results into one slate per request so the same user sees both, collapsing between-user variance (10–100× faster, per Netflix)
An interviewer asks why offline and online metrics disagree even when the offline pipeline is bug-free. Strongest explanation?
AOffline metrics are computed on the logging policy’s distribution; a new policy changes slates, positions, and context, so the click distribution re-weights — a counterfactual the logs can’t captureBOnline metrics are simply noisier, so they drift away from the true offline valueCThe offline test set is too small; a larger one would make them agree
Your team optimises hard against an offline proxy metric that historically correlated with revenue. Revenue is now flat in the A/B while the proxy soared. Most likely cause?
AThe A/B is underpowered; rerun with more trafficBGoodhart/surrogation — optimising hard against the proxy moved it without moving the goal; the observational correlation broke under the interventionCRevenue is the wrong online metric; switch to the proxy as the OEC
You propose off-policy (IPS / doubly-robust) estimation to compare a new policy against logs. When is this least reliable?
AWhen the new policy is nearly identical to the logging policyBWhen the metric is binary rather than continuousCWhen the new policy diverges far from the logging policy, giving high-variance importance weights and poor counterfactual coverage