Lesson 5 of 6 · 47 min

Data leakage & validation

The reason your model scored 0.95 offline and 0.70 in production. Target leakage, temporal leakage, group/cluster leakage, and the cross-validation splits that quietly cheat — plus the experiment-side cousins (SRM, clustered variance) that invalidate a test the same way. The single biggest source of overstated offline numbers.

When the evaluation itself is cheating

Leakage is when information that would not be available at prediction time sneaks into training or evaluation — so your offline metric is measuring a problem easier than the real one. It is the single biggest reason a model that scored beautifully offline collapses online, and it is insidious because nothing errors: the pipeline runs, the number is high, the chart looks great. Every technique in the previous four lessons assumes the evaluation is honest; leakage breaks that assumption silently. This lesson catalogues the forms — target, temporal, group — the cross-validation splits that cheat, and the experiment-side cousins (SRM, clustered variance) that invalidate a live test the same way.
The canonical reference is Kaufman, Rosset & Perlich (“Leakage in Data Mining,” KDD 2011), who define leakage as the introduction of information about the target that should not legitimately be available. Their formal test for whether a feature leaks: would this feature’s value be known, in this form, at the moment you must make the prediction? If the answer is no — because it’s computed using the future, the label, or other rows — it leaks. The discipline is to reconstruct the exact information state at prediction time and forbid anything outside it. Interview angle. “How do you check a feature for leakage?” → ask whether it would be available, unchanged, at prediction time; suspiciously high single-feature importance is the smoke that signals it.

Target leakage: the label hiding in a feature

Target leakage is when a feature is a proxy for, or is computed from, the label. The classic example: predicting whether a patient has a disease, with “was prescribed disease-X medication” as a feature — it’s available in the historical data but only because the diagnosis already happened; at true prediction time you don’t have it. Another: a churn model with “account_closed_date” or “number of retention-team calls,” both of which exist only for users who already churned. The tell is a feature with implausibly high importance and a model that’s “too good” — near-perfect AUC on a genuinely hard problem is almost always leakage, not genius.
code
1THE LEAKAGE TAXONOMY (info unavailable at prediction time, by form)23  type          mechanism                         classic tell4  -----------   -------------------------------   --------------------------5  TARGET        feature is a proxy for / derived  "prescribed drug X" predicts6                from the label                    "has disease X"; AUC ~0.997  TEMPORAL      training uses data from AFTER     random CV on time series;8   (look-ahead) the prediction timestamp         features computed with future9  GROUP         same entity in train AND test;    same user/patient/store rows10                leaks identity, not just label    split across the CV fold11  PREPROCESSING fit a scaler/imputer/encoder on   StandardScaler.fit(ALL data)12                the FULL dataset before split     before train/test split1314  Rule: reconstruct the EXACT information state at prediction time.15        If a feature wouldn't exist then, in that form, it leaks.
A subtler target-leakage source is preprocessing fit on the full dataset. If you fit a StandardScaler, an imputer, a target encoder, or a feature-selection step on all the data and then split into train/test, the test set has leaked into the transform — the scaler’s mean, the encoder’s category statistics, the selected features all “saw” the test rows. The fix is to fit every data-dependent transform inside the cross-validation loop, on the training fold only (scikit-learn’s Pipeline exists precisely to make this hard to get wrong). Target encoding is the worst offender: encoding a categorical by its mean target value leaks the label directly unless done with out-of-fold encoding.

Temporal leakage: training on the future

Temporal (look-ahead) leakage is using information from after the prediction timestamp to predict the past. The most common form is doing random k-fold cross-validation on time-series data: random folds put future rows in the training set and past rows in test, so the model learns from the future it would never have at prediction time. Any time the data has a temporal order and the deployed model predicts forward, you must validate forward — a time-based split (train on [t₀, t₁], test on (t₁, t₂]) or rolling/expanding-window CV. Feature engineering leaks the same way: a “30-day rolling average” must use only the 30 days before each prediction point, never a window centred on it.
Temporal leakage is especially dangerous because it interacts with non-stationarity: a randomly-folded model looks great offline (it’s seen the regime) and then degrades the moment the world moves, because it never had to generalise across time. This is the offline–online gap (L4) wearing a different costume — the offline number was inflated by a validation scheme that doesn’t match how the model is used. Interview angle. “How do you cross-validate a forecasting / time-series model?” → never random k-fold; use a forward-chaining (rolling/expanding window) split so every test period is strictly after its training period, mirroring deployment.

Group leakage: the same entity in train and test

Group (or cluster) leakage is when rows from the same entity appear in both training and test folds, so the model can memorise the entity rather than learn the pattern. If a patient has 10 visits and a random split puts 7 in train and 3 in test, the model can recognise “this patient” and the test score is optimistic. The same happens with multiple sessions per user, multiple photos per person, multiple transactions per merchant. The fix is grouped cross-validation (scikit-learn’s GroupKFold / StratifiedGroupKFold): split on the group key so every entity is entirely in train or entirely in test, never both. This is the validation-time twin of the randomization-unit decision from L3 — both insist the unit of independence is the entity, not the row.
code
1CROSS-VALIDATION: MATCH THE SPLIT TO THE INDEPENDENCE STRUCTURE23  data structure          WRONG (leaks)          RIGHT4  ---------------------   -------------------    --------------------------5  i.i.d. rows             --                     k-fold / stratified k-fold6  time-ordered            random k-fold          time-split / rolling window7   (forecast forward)     (future in train)      (forward chaining)8  repeated per-entity     random k-fold          GroupKFold on the entity key9   (user/patient/store)   (same entity both)     (entity wholly in one fold)10  imbalanced classes      plain k-fold           StratifiedKFold (keep ratio)1112  And ALWAYS: fit scalers/encoders/selection INSIDE the fold (use a Pipeline),13  never on the full dataset before splitting.

The experiment-side cousins: SRM and clustered variance

Leakage has direct analogues in online experiments that invalidate a live test just as silently. Sample Ratio Mismatch (SRM): the realised traffic split deviates from the intended ratio, detected by a chi-square goodness-of-fit test on the bucket counts. It is your free canary — an SRM almost always means a bug (a broken trigger, a bot, a feature-flag leak, dogfooding), and it nullifies randomisation. Kohavi’s production example: a 50/50 Bing split came in at 821,588 vs 815,482 (50.2%), an SRM with p = 1.8×10⁻⁶ — a 1.6-percentage-point deviation is wildly significant at that scale. The MSN 16-vs-12-slide test read 49.8% vs 50.0% because heavy users were filtered as bots, breaking the split. The rule: run a chi-square SRM check before you trust any p-value; if it fails (p < 0.001 vs the planned split), kill the analysis and find the bug.
code
1SRM: THE FREE CANARY (chi-square on realized bucket counts vs planned split)23  Bing 50/50:   821,588 (treatment)  vs  815,482 (control)4                realized 50.2/49.8  ->  chi-square p = 1.8e-6  -> SRM!5                (a 1.6-pt deviation is hugely significant at this N)67  MSN 16-vs-12 slides:  49.8% vs 50.0%  ->  heavy users filtered as bots8                        -> the "weird ratio" was a data-quality bug, not a result910  common SRM causes: broken trigger, bot filtering, feature-flag leak,11                     dogfooding, redirect/latency asymmetry, concurrent-exp clash1213  rule: chi-square SRM check FIRST. p < 0.001 vs plan -> kill analysis, find bug.
The second cousin is clustered variance — the experiment-side version of group leakage. When a user contributes many observations (sessions, purchases) and you compute the metric’s variance at the observation level, you treat correlated rows as independent, the standard error comes out too small, the CI too narrow, and false positives jump 2–10× in sticky products. The fix is the same unit-of-independence principle: cluster the variance at the user level (the delta method on user-aggregated metrics, or a hierarchical model). It’s the exact mirror of GroupKFold — both refuse to count the same entity’s correlated observations as independent evidence. Interview angle. “Users have many sessions each — how does that affect your analysis?” → cluster the variance at the user level; observation-level variance understates SE and inflates significance.
paperLeakage in Data Mining: Formulation, Detection, and AvoidanceKaufman, Rosset & Perlich (KDD 2011)docsCross-validation: the right and wrong ways (GroupKFold, TimeSeriesSplit, Pipelines)scikit-learn User GuidearticleAddressing the Challenges of Sample Ratio Mismatch (SRM) in A/B TestingDoorDash Engineering

Interview prep

The leakage round tests whether you’re suspicious of your own good numbers. Interviewers want you to name the three forms, match the CV split to the data structure, and connect it to the experiment-side cousins (SRM, clustering). Lead with the “available at prediction time?” test and a concrete example.
  1. 01“What is data leakage?” → information unavailable at prediction time sneaking into train/eval, making the offline number measure an easier problem than reality.
  2. 02“How do you detect target leakage?” → ask if each feature would exist, unchanged, at prediction time; a single feature with implausible importance / “too-good” AUC is the smoke.
  3. 03“How do you cross-validate a time series?” → never random k-fold; forward-chaining (rolling/expanding window) so test is strictly after train, mirroring deployment.
  4. 04“Patients have many visits each — how do you split?” → GroupKFold on the patient key so each entity is wholly in train or test; random folds leak identity.
  5. 05“Where does preprocessing leak?” → fitting scalers/encoders/selection on the full dataset before the split; fit inside the fold via a Pipeline; target-encode out-of-fold.
  6. 06“What is SRM and why care?” → realized split ≠ planned (chi-square); it’s a free canary for bugs (broken trigger, bots, flag leak) and nullifies randomisation — check it before any p-value.
  7. 07“Users contribute many sessions — what breaks?” → observation-level variance understates SE; cluster at the user level (delta method / hierarchical) or false positives jump 2–10×.
  8. 08“Your model scores 0.97 offline, 0.72 online — first hypothesis?” → leakage in the offline pipeline (target/temporal/group) or a CV split that doesn’t match deployment.
Going deeper, the follow-ups probe the subtle cases: “is using future data ever legitimate?” (only if it’s genuinely available at prediction time for the deployed use case — e.g. backfilled features that arrive before scoring; otherwise no); “how do you catch leakage you didn’t anticipate?” (audit top feature importances, compare offline vs a strict time-split holdout, and treat any large offline–online gap as a leakage hypothesis first); and “how is GroupKFold related to your randomization unit?” (both enforce that the entity, not the row, is the unit of independence — the same principle on the validation side and the experiment side). Name the unit of independence in every answer.

Checkpoint

A churn model hits 0.99 AUC offline. Among its top features is “number_of_retention_team_calls.” What should you suspect?

AA breakthrough model — 0.99 AUC means the features are highly predictiveBTarget leakage — retention-team calls happen because the user was already flagged/churning, so the feature won’t exist in that form at prediction timeCNothing — drop the feature only if it lowers AUC
Sign up free to answer and see why

Checkpoint

You evaluate a daily demand-forecasting model with standard random 5-fold cross-validation and get great scores. What’s wrong?

ANothing — 5-fold is the standard validation methodBYou should use 10 folds instead of 5 for a more stable estimateCTemporal leakage — random folds put future data in training; use forward-chaining (rolling/expanding window) so each test period is strictly after its training period
Sign up free to answer and see why

Checkpoint

A medical model is trained on X-rays where each patient contributed several images. You used random k-fold. Why might the test score be optimistic?

AGroup leakage — the same patient appears in train and test, so the model recognises the patient rather than the pathology; split with GroupKFold on patient idBClass imbalance — use StratifiedKFold to fix itCToo few images per patient — collect more images
Sign up free to answer and see why

Checkpoint

Your A/B was planned 50/50 but the realised split is 50.4% / 49.6% on 1.2M users, chi-square p = 3×10⁻⁵. What do you do before reading the primary metric?

AProceed — a 0.8-point deviation is tiny and within normal noiseBTreat it as an SRM, stop and investigate the cause (trigger bug, bot filtering, flag leak), and don’t trust the primary metric until the split is explained or fixedCRe-weight the arms to 50/50 and analyse as planned
Sign up free to answer and see why

Checkpoint

In a sticky app where users have many sessions, you compute the conversion-rate variance at the session level and get a significant lift. What’s the senior concern?

ANone — more sessions means more data and tighter estimatesBSession-level variance treats correlated within-user observations as independent, understating the SE; cluster the variance at the user level (delta method / hierarchical)CSwitch the metric from conversion rate to revenue
Sign up free to answer and see why

Could you spot target/temporal/group leakage in a pipeline, pick the right CV split, and run an SRM check before trusting a result?

New to itGetting thereConfident

Takeaways

  • Leakage = information unavailable at prediction time inflating your offline number; nothing errors, the chart just lies.
  • Target leakage = label-derived features; the “available at prediction time?” test and suspiciously high importance catch it.
  • Temporal leakage = random folds on time data; use forward-chaining. Group leakage = same entity in both folds; use GroupKFold.
  • Fit every scaler/encoder/selection inside the fold (Pipeline); target-encode out-of-fold.
  • Experiment cousins: run a chi-square SRM check before any p-value, and cluster variance at the user level to avoid 2–10× false positives.

Next: the capstone — design and analyse a complete A/B test end to end, defending every choice.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.