Offline improvement can miss the product objective
Separate an offline prediction gain from a causal user-outcome claim.
An offline metric measures a function on recorded examples. A product experiment measures the effect of changing what users experience under a specified assignment. These are related, but they answer different questions. A ranking model can predict historical clicks better without causing more useful clicks when deployed.
Historical logs reflect the old policy. Items shown near the top received more exposure than items never shown. A model trained to reproduce these clicks may learn position or popularity effects. Offline evaluation on the same logging pattern can reward that behavior. It does not reveal outcomes for unseen alternatives without additional assumptions or data.
Begin diagnosis by checking the experiment itself. Verify assignment, exposure logging, feature parity, latency, and the outcome definition. A slower model may improve relevance while reducing engagement because users abandon the page. A metric may rise through low-quality repeated actions while the meaningful downstream outcome falls. These possibilities require evidence, not a story selected after the result.
Use a causal design appropriate to the decision. Randomized assignment can estimate the effect of the deployed policy under its conditions, but interference, carryover, missing logging, and unequal exposure can still complicate interpretation. Preserve the assignment unit and analysis plan. Do not repeatedly change the success metric until one turns positive.
Worked example
In an invented recommendation experiment, the candidate improves offline ranking score by 4%. Live assignment is by user. Click-through rate rises from 10% to 11%, but completed purchases stay at 2%. Candidate p95 latency rises from 100 ms to 500 ms.
The evidence supports a click increase under the tested policy, assuming the experiment is valid. It does not support a purchase increase. The latency change is a plausible contributor to downstream behavior, but it is not proven causal merely because it moved. The next test can isolate latency or compare a faster candidate while keeping the objective fixed. The team should also inspect whether the extra clicks lead to low-intent items.
Exercise and solution
A search model increases clicks per session but reduces successful task completion. The interviewer asks if you would ship it. Give the next decision steps.
A strong answer checks the validity and uncertainty of both metrics, identifies task completion as the stated objective if that is the product goal, and blocks a broad release until the tradeoff is understood. It examines repeated searches, latency, and result quality as possible mechanisms. Award one point each for objective alignment, experiment validity, mechanism hypotheses labelled as such, and a targeted follow-up. Do not equate engagement with usefulness without a product definition.
Lab artifact: diagnose numerator and denominator
A rate can rise because its numerator rises, its denominator falls, or both. Consider this original seven-day experiment summary. Assignment is by user, with one thousand assigned users in each arm and complete outcome logging.
Arm
Assigned users
Sessions
Clicks
Completed tasks
Control
1,000
2,000
400
300
Candidate
1,000
1,000
300
280
Clicks per session rise from 0.20 to 0.30, while clicks per assigned user fall from 0.40 to 0.30. Completed tasks per assigned user fall from 0.30 to 0.28. The table contains descriptive differences, not a significance test. It does not tell you whether the differences exceed expected sampling variation or generalize beyond this period. It does show why choosing a favorable denominator after seeing results can reverse the apparent story.
Microsoft's experiment guidance provides direct context for monitoring denominators and telemetry quality. The numerical table and interpretation here are original. To estimate uncertainty, respect the user assignment unit and repeated sessions within users. Treating all sessions as independent observations can give a misleadingly narrow estimate when the same person generates many sessions.
A second failure case: better logging looks like better behavior
Suppose the candidate also changes the event collector. Control records ninety percent of actual clicks, while candidate records all clicks. If true click behavior were equal, recorded candidate clicks would appear about 11.1% higher because 1 / 0.9 is approximately 1.111. This calculation assumes the stated capture rates and otherwise equal behavior. It is a diagnostic counterexample, not an estimate of the earlier experiment's cause.
Before interpreting a click lift, inspect event collection versions, missing-event checks, join rates, and whether errors differ by assignment. A single balanced user count does not establish equally complete outcome measurement. Where an independent server-side event can validate a client event, compare their relationship within each arm.
code
1Hypothesis H1: recommendation quality increased useful engagement2Discriminating evidence: successful downstream tasks per assigned user3Hypothesis H2: candidate event collector records more existing clicks4Discriminating evidence: server/client click reconciliation by arm5Hypothesis H3: repeated clicks reflect confusion6Discriminating evidence: repeated same-item clicks and failed-task sequences
Each hypothesis names an observation that could weaken it. A list of plausible explanations without discriminating checks is not a diagnosis.
Exercise: avoid selecting on a treatment consequence
A candidate makes users spend longer in the app. An analyst compares only users who spent at least thirty minutes after assignment and reports a lift. Explain the problem and a better primary analysis.
Time spent can be changed by the treatment, so membership in the selected subgroup can differ because of assignment. The comparison no longer preserves the original randomized grouping in a simple way. Use the prespecified assigned population for the primary analysis. If a segment is needed, prefer a relevant pre-assignment characteristic such as prior-week activity. A post-assignment analysis may still be descriptive or require a more careful causal method, but should not be presented as the original randomized effect. Award one point for identifying the post-treatment condition, one for its selection mechanism, and two for the primary and segment analyses.
Misconceptions to correct
“Random assignment makes every subsequent filter unbiased” fails when the filter depends on behavior that treatment changes. “An offline metric and a live metric disagree, so one is broken” fails when they measure different objectives or populations. Either can be correctly measured and still be insufficient for the product decision.
A good interview conclusion states the supported effect, the validity checks not yet complete, and the next experiment. If latency is suspected, a controlled change that isolates latency can distinguish that mechanism from recommendation quality. Merely correlating slow requests with failed tasks does not establish the latency effect, because complex requests may both run slowly and fail more often.
Interview probe
Original practice: Why might offline NDCG improve while revenue stays flat? A strong answer considers logging policy, objective mismatch, latency, and experiment validity. Follow up by asking which observation would distinguish two hypotheses. A weak answer invents a single cause from the aggregate.
Assuming valid measurement, more clicks with unchanged purchases directly supports which claim?
AA measured click increase, without a demonstrated purchase increase.BA purchase increase hidden by the metric definition.CA proven latency mechanism.DAn invalid experiment because outcomes disagree.
Control has 400 clicks/2,000 sessions and candidate 300/1,000; both have 1,000 assigned users. What happened?
ABoth clicks per session and per user increased.BBoth rates decreased.CClicks per session rose while clicks per user fell.DNeither changed because users are balanced.
Why is filtering to users with at least 30 post-assignment minutes risky?
AThe treatment can change membership in that selected group.BEqual total assignment counts guarantee this filtered subgroup remains balanced.CPost-assignment activity is equivalent to activity measured before assignment.DFiltering on a treatment consequence preserves randomization whenever both arms use the same cutoff.
Latency and failed tasks are correlated. Which next check better tests latency as a cause?
ADeclare latency causal from the correlation.BSelect only successful tasks before comparing latency.CRun a valid controlled change that isolates latency under the same objective.DSwitch the objective to clicks after inspecting results.
Can you diagnose an apparent lift using denominators, logging coverage, and assignment units? Rate confidence from 1 to 5 and propose a check that could refute your preferred cause.
Not yetGetting thereConfident
Wrap-up
Connect offline evidence to a valid live decision. Keep causal claims within the experiment's assignment and measured outcomes.