Lesson 4 of 4 · 35 min

Drift is a symptom to investigate

Choose the next diagnostic check from a drift report.

A distribution change tells you that something moved. It does not tell you why it moved or whether model quality fell. Feature drift, label prevalence change, calibration shift, and service bugs can produce similar dashboards. An ML engineer should connect each alert to a diagnostic question.
First check the measurement path. A renamed category, changed unit, missing-value default, or logging bug can create a large feature shift without a change in user behavior. Compare raw inputs and transformed features by version. Next inspect traffic composition: a new region, device class, or acquisition channel may explain the aggregate. Then use mature labels to assess quality within relevant slices.
Do not retrain automatically on every alert. New data may contain broken labels or reflect a temporary incident. Retraining can encode the problem into the model. A retraining decision needs a valid dataset, a plausible reason that updated learning will help, and an evaluation under the intended policy. Sometimes the correct fix is restoring a feature service or adjusting an operating threshold after a justified cost or prevalence change.
Separate leading indicators from outcomes. Missing-feature rates and score shifts are fast but indirect. Fraud loss or conversion may arrive slowly. Keep the label maturity rules from training in monitoring. Comparing yesterday's incomplete labels with last month's mature labels produces a false quality alarm.

Worked example

A teaching classifier's positive prediction rate rises from 4% to 12%. Feature missingness also rises from 0.5% to 18% immediately after a deployment. The affected feature uses zero as an imputation value. The first investigation compares feature-service versions and missingness by request type. It does not conclude that fraud prevalence tripled.
The team discovers that a category field changed spelling and bypassed a lookup. Restoring the mapping returns missingness and score rates near their prior range. Mature outcome labels still need review to quantify impact. The incident record distinguishes the observed serving bug from any unproven claim about lost revenue.

Exercise and solution

A model's precision falls after traffic expands into a new region. Missingness and latency are stable. Labels are mature. What checks distinguish population change from a general model regression?
Calculate precision and other relevant metrics by region, compare prevalence and score distributions, verify label definitions, and compare the old region before and after expansion. If the old region is stable and the new region differs, the issue is localized rather than a universal regression. Award one point each for slice analysis, prevalence check, label validity, and a limited conclusion. Avoid claiming that any single drift statistic proves causation.

Lab artifact: separate composition from within-group change

An aggregate precision change can come from a different mix of otherwise stable groups. The following invented audit uses mature labels and the same decision threshold in both periods.
PeriodRegionFlagged casesTrue positivesPrecision
BeforeEstablished90054060%
BeforeNew1002020%
AfterEstablished50030060%
AfterNew50010020%
Before expansion, aggregate precision is 560/1,000 = 56%. After expansion it is 400/1,000 = 40%. Neither regional precision changed. The aggregate drop is real for the actual reviewed population, but describing it as a universal within-region regression would be wrong. At a fixed fifty-fifty region mixture, both periods have forty percent precision. Standardization helps explain composition; it does not erase the operational fact that reviewers now encounter a lower-yield queue.
The next question is why the new region has lower yield. Check class prevalence, score calibration, error patterns, and whether the review action has the same cost. A common threshold is not necessarily a common operating tradeoff. Conversely, separate thresholds require a justified policy and evaluation. The table alone cannot tell you which policy is appropriate.

A second failure case: comparing different label ages

Suppose yesterday's flagged transactions include ten confirmed positive outcomes out of one hundred, while last month's flagged transactions include thirty out of one hundred. If confirmations can arrive within seven days, yesterday's apparent ten percent precision is an interim lower count, not directly comparable with last month's mature thirty percent. Keep cohort age visible in the monitoring dataset.
code
1monitor_record:2  cohort_date: 2026-09-203  policy_version: review-v74  label_window: 7 days5  extraction_date: 2026-09-226  status: immature7  observed_positive_count: 108  eligible_for_final_precision: false
The final metric also requires a functioning outcome collection path. Seven elapsed days do not make labels complete if the event pipeline failed for three days. A label-maturity flag should reflect both the outcome window and observed collection coverage. This links monitoring to the training-data contract rather than inventing a different definition for dashboards.

Exercise: choose a remedy from a diagnostic table

A deployment changes temperature values from Celsius to Fahrenheit without updating the model's preprocessing. Missingness stays at zero, request latency stays at fifty milliseconds, and mean raw temperature rises from twenty to sixty-eight. Scores shift. Mature performance labels will arrive next week. The model contract explicitly requires Celsius.
The immediate remedy is to restore or correctly convert units and replay a small parity fixture. Retraining on the shifted inputs would entrench a known contract error. A stable missingness chart does not rule out semantic failure because the values are present and numerically valid. Use an example: sixty-eight Fahrenheit equals twenty Celsius, so the normalized feature should match the earlier twenty-Celsius value. Award one point for identifying unit mismatch, one for rejecting missingness as sufficient, one for the conversion check, and two for the limited impact claim and verified repair.

Misconceptions to correct

“No feature drift means no model problem” fails when labels or the relationship between features and outcomes change while feature marginals remain stable. A loan applicant mix can look similar while a changed economic environment changes repayment. “A large drift statistic means retraining will help” fails when the root cause is logging or units. Diagnostics need a mechanism connecting the proposed repair to the observed failure.
The release packet should end with a short incident note: observed facts, competing explanations, discriminating checks, chosen action, and unresolved impact. For the category-mapping incident, the known facts are the deployment boundary, missingness increase, and reproduced lookup failure. Revenue impact remains unresolved until valid outcomes and exposure records can support it. Avoid filling that gap with a precise-looking estimate from an unrelated population.
A useful follow-up check can disprove your favored explanation. If reverting the category mapping repairs missingness but score rates remain shifted within the same traffic slices, investigate another cause. Restoration of one dashboard is not proof that every effect is resolved. Preserve the pre-repair examples, the corrected transformation, and the scope of the replay so a reviewer can assess what was actually verified.

Interview probe

Original practice: Does feature drift mean you should retrain? A strong answer checks serving correctness, population, mature outcomes, and the proposed remedy's mechanism. Follow up with a unit conversion bug. A weak answer schedules retraining whenever a histogram changes.

Sources

docsGoogle: rules of machine learningdevelopers.google.comdocsscikit-learn: probability calibrationscikit-learn.org

Checkpoint

Scores shift and missingness rises after a feature release. What is the most direct first investigation?

ARetrain immediately on the newest labelled cohort before checking its feature version.BCompare feature retrieval and transformations across the release.CLower the threshold until the old flag rate returns.DWait for next month's labels before inspecting the input path.
Sign up free to answer and see why

Checkpoint

Regional precision is unchanged at 60% and 20%; the mixture of reviewed cases shifts from 90/10 to 50/50. What explains aggregate precision falling 56% to 40%?

AA proven within-region regression.BA change in label definition is necessary.CThe aggregate metric was calculated incorrectly.DA composition change, despite stable regional precision.
Sign up free to answer and see why

Checkpoint

A seven-day outcome window has elapsed but collection failed during days 4–6. Are negative labels complete?

ANot established; observation coverage must also be checked.BYes, elapsed time alone makes them mature.CYes, unless a positive event is already present.DYes, if the number of prediction records matches the expected traffic volume.
Sign up free to answer and see why

Checkpoint

Celsius inputs changed to Fahrenheit; missingness is zero. What should happen before retraining?

ANormalize the new values with the old scaler and monitor recall.BRestore the feature unit contract and verify parity examples.CTreat the higher values as population drift and accept them.DChange only the score threshold to recover the old flag rate.
Sign up free to answer and see why

Checkpoint

A mapping repair restores missingness but within-slice score distributions remain shifted. Which conclusion is justified?

AThe full incident is resolved because missingness recovered.BRetraining is automatically the remaining fix.CAt least one remaining cause needs investigation.DThe original mapping failure was impossible.
Sign up free to answer and see why

Can you separate composition change, input-contract failure, and immature outcomes before selecting a remedy? Rate confidence from 1 to 5 and state your next discriminating check.

Not yetGetting thereConfident

Wrap-up

  • Use drift alerts to select investigations. Confirm the measurement path and mature outcomes before selecting a remedy.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.