Lesson 2 of 4 · 35 min

Split by the unit you want to generalize to

Choose grouped and temporal validation from deployment requirements.

A random row split answers a specific question: how well does the model predict new rows drawn under roughly the same sampling process? It may not answer how well the model handles new users, future months, new hospitals, or new devices. Define the deployment unit before choosing a splitter.
Repeated observations create dependence. If one user's near-identical sessions appear in training and validation, the model may recognize the user instead of learning behavior that transfers. Grouping all rows from the same entity protects against that overlap when the goal is generalization to unseen entities. It is not always the right goal. A service that predicts future behavior for known users also needs a time-aware evaluation.
Time and group constraints can both matter. A new-customer model may require future records from customers absent in training. A forecasting model may need rolling cutoffs and a gap that accounts for feature windows or label overlap. A splitter's name is not a guarantee that it matches the task. Inspect actual entity sets and date intervals after splitting.
Use validation for model selection and reserve a final test for the final protocol. Repeatedly checking the test set and changing features based on its score turns it into another validation set. Report the selection process and keep a record of which data informed each choice. A slightly lower but valid score is more useful than an impressive score on a contaminated comparison.

Worked example

A fictional dataset has 100 patients and ten visits per patient. A random 80/20 row split places visits from most patients in both sets. A patient-ID feature then helps identify familiar cases. If deployment concerns unseen patients, use a group split by patient ID: 80 patients for training and 20 for validation, with no shared IDs.
Now change deployment to predict next month's visits for existing patients. The grouped split no longer matches that question by itself. Train on visits before the cutoff and evaluate later visits, using only historical features. Report the known-patient and unseen-patient results separately if both populations matter. These two evaluations estimate different behavior and should not be collapsed without a stated population mix.

Exercise and solution

You have monthly purchase records from January through June for 300 stores. The model will launch in July at existing stores and 20 newly opened stores. Propose evaluation slices.
A strong solution uses a temporal holdout for future months at known stores and a separate held-out-store slice for new-store transfer. It checks that features obey each cutoff and reports denominators. It does not use July labels that do not yet exist or claim the new-store slice perfectly predicts July conditions. Award one point each for temporal realism, group separation, separate metrics, and an explicit limitation.

Audit the split with identities and time ranges

A split function's name is not evidence that the resulting partition matches the target. Save a small audit table for each fold.
CheckTrainingValidationRequired relationship
Customer IDs for unseen-customer targetSet TSet VIntersection empty
Prediction times for future targetBefore cutoffAfter cutoffOrdered as declared
Label windowsMature under train cutoffMature under evaluation dateNo immature labels
Learned preprocessingFitted hereTransform onlyNo validation fit
Duplicate problem groupsGroups TGroups VSeparate when target requires new groups
The relationships depend on the question. For known-customer future prediction, customer overlap may be intentional, while future information remains forbidden. For unseen-customer transfer, empty identity overlap is necessary but may still leave near-duplicate household or organization information. Choose the group level that matches the actual dependence.
This pseudocode is a review aid. It does not replace a full splitter.
code
1if target == "unseen_customers":2    assert intersection(train.customer_ids, valid.customer_ids) is empty3assert max(train.prediction_time) < validation_cutoff4assert min(valid.prediction_time) >= validation_cutoff5assert every label used satisfies its maturity rule6assert preprocessing.fit_ids are a subset of training_ids7save split IDs and audit results with the experiment
For a rolling evaluation, repeat the audit at each cutoff. Do not fit one scaler on all months and reuse it inside every historical fold, because that scaler can contain information from later periods. A pipeline must be fitted inside the relevant training portion.

A second worked case: household dependence

A lending-style teaching dataset has applications from two people in each household. The target is performance for entirely new households. A person-level group split puts one household member in training and the other in validation. Addresses and shared account features make the records strongly related.
The correct grouping unit for the stated target is household, not person. This does not mean household grouping is universally required for every application. It follows from the target population and available shared information. If household IDs are incomplete, report the limitation and audit likely duplicates rather than claiming perfect separation.
A similar issue appears in document datasets where multiple paragraphs come from one source, or coding tasks generated from one template. Row count can greatly exceed the number of independent groups. A large validation set with weak separation can still give an optimistic estimate.

Choose the final evaluation role deliberately

If you tune a model using performance on a held-out hospital, that hospital has become development evidence. It can remain a useful regression slice, but it is no longer an untouched test of new-hospital transfer. Reserve another independent site or use a suitable outer evaluation procedure if the data supports it. With very few sites, state the resulting uncertainty rather than pretending many visits create many independent sites.
The same principle applies to repeated daily inspection. You do not need gradients for a dataset to influence model selection. Feature choices, thresholds, and checkpoint selection all use information.

Misconceptions to reject

"Stratification prevents leakage" confuses class proportions with independence. Stratified rows can still share an entity or a future-derived feature across partitions.
"A group split guarantees future performance" protects one boundary only. It may not reproduce temporal drift, availability delays, or new-site conditions.

Transfer exercise

A dataset contains 1,000 support messages from 100 conversations, with ten messages per conversation. The model will classify the first message of new conversations. A random message split gives high accuracy. Design a more relevant evaluation.
Use conversation groups, include only information available at the first-message decision, and evaluate on new conversations under a temporal holdout if future deployment matters. Later messages cannot be input features for that first-message target. Score one point for grouping, one for the decision cutoff, one for time realism, and one for reporting the actual number of conversations rather than only messages.

Interview probe

Original practice: Should every dataset use GroupKFold? A strong answer starts with the deployment target and explains when entity separation is necessary. Follow up with known-user forecasting. A weak answer treats one splitter as universally correct.

Sources

docsscikit-learn: cross-validationscikit-learn.orgdocsscikit-learn: common pitfalls and data leakagescikit-learn.org

Checkpoint

Repeated visits share patient identity, and deployment targets unseen patients. Which split addresses that overlap?

AGroups by patient.BStratified visits without grouping.CAlternating visits from each patient.DRandom visits.
Sign up free to answer and see why

Checkpoint

The target is new households, but people within each household are split separately. What is the gap?

AStratifying people by outcome removes all household dependence.BDropping the household-ID feature guarantees the model cannot exploit shared information.CThere is no gap because person IDs differ.DShared household information can cross the evaluation boundary.
Sign up free to answer and see why

Checkpoint

A temporal fold uses a scaler fitted on all months. What should change?

ARemove all temporal ordering.BKeep it because the model itself trains only on earlier months.CFit the scaler inside each fold's training portion.DFit the scaler on validation only.
Sign up free to answer and see why

Checkpoint

A held-out hospital repeatedly guides feature changes. What role does it now serve?

ADevelopment evidence for model selection.BAn untouched test set if no optimizer gradients use it.CA valid final estimate after reporting only the final feature set.DAn independent test set whenever preprocessing excludes its rows.
Sign up free to answer and see why

Checkpoint

What should a split audit report beyond row counts?

AOnly the splitter's function name.BOnly class balance.CGroup intersections, time ranges, label maturity, and preprocessing fit IDs.DOnly the best fold score.
Sign up free to answer and see why

Explain how you would audit group and time separation for the actual deployment population. Use one supplied fixture and identify a condition that would invalidate your conclusion. Rate confidence from 1 to 5.

Not yetGetting thereConfident

Wrap-up

  • Choose splits to represent the future decision population. Inspect group and time overlap directly.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.