Exclude features that were unavailable when a decision was made.
A training row represents a decision at a particular time. Its features must reflect information available then. The latest database state often contains later corrections, resolved outcomes, or backfilled fields. Joining that state to an old label can produce a model that appears excellent offline and cannot work online.
Distinguish event time from availability time. An event may occur at 09:00 but arrive in the feature pipeline at 09:07. A prediction at 09:03 could not use it, even though its event timestamp is earlier. For each feature, define the cutoff and the delay policy. If production uses data that is at least ten minutes old, offline extraction should reproduce that condition.
Labels belong after the prediction cutoff because they describe the outcome to learn. Features derived from those outcomes do not. A support ticket's final resolution can be a label for triage research, but including the resolution timestamp as an input leaks future information. The same risk appears in refund flags, account closure reasons, and post-purchase review fields.
Use point-in-time joins that select the latest eligible feature record for each entity and prediction timestamp. Preserve source versions so the extraction can be rerun. A SQL query that uses only the maximum timestamp per entity is insufficient when it ignores the prediction cutoff. A pipeline library prevents some transformation leakage, but it cannot repair an incorrectly defined historical join.
Worked example
An invented churn model scores account A at 10:00.
Record
Event time
Available time
Value
Usage snapshot U1
09:30
09:35
12 actions
Usage snapshot U2
09:55
10:05
3 actions
Cancellation C1
13:00
13:00
cancelled
At prediction time, U1 is eligible and U2 is not. C1 can contribute to a later churn label, but cannot be an input. The correct feature is 12 actions. A latest-state join would select 3 actions or even the cancellation flag and make the task artificially easy.
The dataset contract records prediction time, source availability time, feature version, and label horizon. A test creates exactly these three records and asserts that the row uses U1. This is a small test with a direct causal purpose. Randomly splitting a large leaked dataset does not expose the problem because the leak exists in every split.
Exercise and solution
A fraud decision occurs at 12:00. A device event happened at 11:58 and arrived at 12:02. A chargeback arrived seven days later. Which fields are eligible inputs and which may define labels?
The device event is unavailable at 12:00 and must be excluded unless production also delays the decision. The chargeback can define a future label after a stated maturity window. Neither is an eligible immediate feature. Award one point each for the availability distinction, the prediction cutoff, future-label separation, and a reproducible extraction test.
Inspect an as-of join
The following SQL-like pseudocode shows the intended eligibility conditions. It is not tied to a particular warehouse. The created timestamp must mean when the value became available to the decision system; a database insertion timestamp is not sufficient if another pipeline delay occurs afterward.
sql
1SELECT prediction_id, entity_id, feature_value2FROM (3 SELECT p.prediction_id, p.entity_id, f.feature_value,4 ROW_NUMBER() OVER (5 PARTITION BY p.prediction_id6 ORDER BY f.event_time DESC, f.available_time DESC7 ) AS position8 FROM predictions p9 LEFT JOIN feature_history f10 ON f.entity_id = p.entity_id11 AND f.event_time <= p.prediction_time12 AND f.available_time <= p.prediction_time13 AND f.event_time >= p.prediction_time - INTERVAL '24 hours'14) eligible15WHERE position = 1;
The event-time condition limits what happened before the decision. The availability condition limits what the system could know. The freshness condition limits how old a value may be under this feature contract. They answer separate questions. A value can satisfy two and fail the third.
The pseudocode needs production details before use: a deterministic tie-breaker for identical timestamps, explicit handling of no match, time-zone normalization, and a definition of corrections. A left join should preserve prediction rows with missing features rather than silently dropping them. Otherwise, evaluation may exclude exactly the cases where serving has missing data.
A second worked case: a late correction
Suppose account B has two feature records for the same event time. The first reports balance 100 and was available at 08:00. A correction reports balance 20 and was available at 11:00. The prediction occurred at 09:00.
Record
Event time
Available time
Balance
Eligible at 09:00?
Original
07:00
08:00
100
Yes
Correction
07:00
11:00
20
No
Later event
10:00
10:01
10
No
A latest-correction query may choose 20 because it is now the best representation of the historical account. But the model at 09:00 saw 100. If the evaluation asks whether the live system would have made a good decision then, the original available value is the correct input. If the research asks a different retrospective question using corrected facts, state that target explicitly.
Feast's documentation makes a related distinction. Its default point-in-time retrieval constrains event time, while created-time filtering is a separate option with support and timestamp assumptions. Do not assume that a feature store's "point-in-time" label automatically reproduces online availability under every backend and schema.
Test the extraction, not only the model
Create a fixture containing a current eligible record, a late-arriving older event, a future event, a stale record outside the freshness window, and an entity with no record. The expected extraction includes one eligible value or a documented missing state for every prediction. Check row count, selected record identity, and cutoff conditions.
A high model score cannot validate this query. In fact, leakage can make the score higher. The extraction test needs known timestamps and expected records, so it can fail even when downstream accuracy looks excellent.
Misconceptions to reject
"Historical event time guarantees historical availability" ignores delayed ingestion and corrections. The information may have existed in the world before it reached the system.
"Dropping rows with missing feature matches cleans the evaluation" can select an easier population than production. Missing-input cases need a documented policy and their own coverage count.
Transfer exercise
A prediction at 15:00 has a 12-hour freshness limit. Candidate records are: event 02:00 available 02:05 value 9; event 14:00 available 15:10 value 3; event 13:00 available 13:05 value 7. Select the record and explain both exclusions.
The chosen value is 7. The 02:00 record is thirteen hours old and outside the window; the 14:00 record arrived after prediction. Award one point for the selected value, one for each distinct exclusion, and one for preserving the prediction as missing if no candidate remained.
Interview probe
Original practice: Your offline score is excellent, but live performance collapses. How do you inspect temporal leakage? A strong answer reconstructs what was available at each prediction time, including ingestion delay and backfills. Follow up with a record whose event time is early but arrival is late. A weak answer only changes the random seed.
An event occurred before prediction but arrived afterward. Which input treatment matches immediate serving?
AExclude it until it is available to the decision system.BInclude it only in validation.CReplace its arrival time with event time.DInclude it based only on event time.
A corrected value keeps the old event time but arrives after prediction. Which online-replay value is valid?
ALatest correction, because event time is before prediction.BLatest correction if label collection happened after it arrived.CLatest value actually available at the prediction cutoff.DOriginal and corrected values together, to represent uncertainty available now.
Why preserve prediction rows with no eligible feature match?
AAn inner join is equivalent if the missing rate is low.BDropping them can exclude real missing-input cases from evaluation.CImputing the latest future value preserves the intended population without leakage.DRemoving them is safe whenever train and test apply the same removal.
A feature store promises event-time joins. What should you verify for online availability replay?
AThat its historical join sorts only by event time.BWhether availability filtering is supported and its timestamp reflects online availability.CThat correction deduplication always chooses the latest exported record.DThat the query returns one row per prediction, regardless of record timing.
Explain how you would reconstruct a prediction row under late arrival, correction, and freshness limits. Use one supplied fixture and identify a condition that would invalidate your conclusion. Rate confidence from 1 to 5.
Not yetGetting thereConfident
Wrap-up
A valid row reconstructs available knowledge at a decision time. Audit availability, not just event timestamps.