Lesson 3 of 4 · 35 min

Labels can be missing for a reason

Explain delayed and selectively observed labels without treating them as negatives.

A label is a measurement process, not an unquestionable fact. Operational systems often observe outcomes only after a delay or only for items selected by an earlier policy. An ML engineer must describe that process before treating a column as ground truth.
A missing label may mean the outcome has not happened, has not arrived, or cannot be observed. These cases differ. For a purchase-conversion model with a seven-day horizon, a one-day-old session has not matured. Labelling it negative favors recent traffic and changes the class distribution. For a fraud model, only investigated transactions may receive a confirmed fraud label. Uninvestigated transactions are not automatically known clean.
Selection can create a feedback loop. A model recommends items, users see only those items, and the resulting clicks train the next model. The logs describe behavior under the old exposure policy. They do not reveal what users would have done with every unseen item. Log the action policy and relevant probabilities if the system intends to use methods that need them. Do not invent counterfactual labels from absence.
Start with a label contract: outcome definition, horizon, maturity rule, source, revision policy, and unknown state. Audit disagreements with domain experts on a small sample. The goal is not to remove every ambiguous case until accuracy rises. It is to understand what the model is being asked to predict and which examples can support that claim.

Worked example

A teaching conversion dataset is collected on September 22. The label is "purchase within seven days of signup". A user who signed up September 20 with no purchase is immature. A user who signed up September 10 with no purchase is a mature negative, assuming event ingestion is complete. A user who signed up September 19 and purchased September 21 is a known positive within the horizon, but the training protocol should still handle its observation window consistently.
One valid extraction includes only signups whose full horizon plus ingestion allowance has elapsed. This sacrifices recent rows for cleaner labels. Another approach models censoring explicitly, but that is a different statistical task. Simply filling unknown values with zero does not implement it.

Exercise and solution

A hiring platform labels only applicants who received an interview. It wants to predict success for every applicant. Explain the evidence gap and propose a first audit.
The observed labels are selected by the prior interview policy. They cannot directly establish outcomes for rejected applicants. Audit selection rates and feature distributions, inspect the decision process, and define a target supported by available data or a justified data-collection design. Award one point each for selection recognition, unknown-versus-negative distinction, policy logging, and a bounded claim. Do not promise that reweighting automatically removes bias without its assumptions and required data.

Make label state explicit

A binary target does not require every raw record to have a binary label immediately. Keep the observation state alongside the eventual label.
SignupEvaluation datePurchase observedSeven-day windowLabel state
Sep 10Sep 22NoneCompleteMature negative if ingestion complete
Sep 20Sep 22NoneIncompleteImmature
Sep 19Sep 22Sep 21IncompletePositive observed, protocol still controls inclusion
Sep 10Sep 22Unknown due to outageComplete in time onlyUnreliable observation
The outage row shows why time passing is not sufficient. A label window can be complete while the observation pipeline is incomplete. Label maturity should include the expected event delay and any known data-quality incident. Otherwise, a logging outage can create a large artificial negative class.
For a simple fixed-horizon classifier, one defensible protocol includes only cohorts whose full horizon and ingestion allowance have elapsed. It may exclude recent known positives along with immature negatives to keep cohort inclusion consistent. Other designs can use partial observation under explicit statistical assumptions, but filling missing values with zero does not implement those designs.

A second worked case: selected verification

A teaching review system investigates 100 of 1,000 transactions. Of the investigated set, 20 are confirmed positive and 80 confirmed negative. The other 900 are not investigated.
code
1observed_positive: 202observed_negative: 803unverified: 9004unsupported shortcut:5    replace unverified with negative6result of shortcut:7    apparent prevalence = 20 / 1000 = 2%8what the labels directly establish:9    positive rate among reviewed cases = 20 / 100 = 20%
The 20% is also selected and should not be advertised as population prevalence. The data directly describes the reviewed group. To estimate population prevalence or train a broader target, inspect how cases were chosen and what additional collection or assumptions are available. The correct response is neither to call all unverified cases clean nor to generalize the reviewed rate automatically.
If review selection depends on an existing model score, the observed labels can reinforce its blind spots. Cases the model never selects may remain unverified. Log the selection policy and changes over time. Any correction method must be justified by the information it needs, such as known selection probabilities or an appropriately sampled audit set.

Review disagreement is another measurement signal

Two reviewers can assign different labels because the rubric is ambiguous or because the case lacks evidence. Store disagreement and adjudication rather than forcing every case through majority vote without inspection. A disagreement concentrated in one category may indicate that the target definition needs refinement.
A label audit should sample more than obvious positives. Include mature negatives, unverified cases where audit is permitted, recently revised labels, and cases near the operational boundary. Record the sampling scheme so the audit's error rate is not misrepresented as a population estimate.

Misconceptions to reject

"A database null is a negative class" converts a storage state into a semantic conclusion. Missingness can represent delay, selection, outage, or uncertainty.
"The reviewed positive rate is the population positive rate" ignores the selection process. A review queue often deliberately over-samples suspicious cases.

Transfer exercise

A support model predicts escalation within 24 hours. At extraction, 60 tickets are older than 48 hours with complete logs: 15 escalated and 45 did not. Forty tickets are only six hours old with no escalation yet. Under a complete-window protocol, what is the training label set?
Use the 60 mature tickets, with 15 positives and 45 negatives. Mark the 40 recent tickets immature and exclude them from this binary extraction until the window and allowance complete. Award one point for each count, one for the observation-state distinction, and one for checking logging completeness. Do not report 15% prevalence from all 100 as if every outcome were observed.

Interview probe

Original practice: Can you train on all rows by replacing missing labels with zero? A strong answer distinguishes mature negatives, delayed outcomes, and unobserved counterfactuals. Follow up with labels collected only from investigated cases. A weak answer treats a database null as a semantic label.

Sources

docsGoogle: rules of machine learningdevelopers.google.comdocsscikit-learn: model evaluation metricsscikit-learn.org

Checkpoint

A seven-day outcome is absent for yesterday's signup. Under a complete-window binary-label protocol, what state is justified?

ANegative if all currently collected events are complete.BNegative for training but unknown for validation.CNegative with a lower sample weight.DImmature or unknown until the required window completes.
Sign up free to answer and see why

Checkpoint

A label window elapsed, but collection was unavailable during part of it. What follows?

AAbsent events are negative because the observation clock completed.BTime maturity alone does not establish complete observation.CA larger cohort makes the missing interval irrelevant to label validity.DUsing the same outage-affected labels in train and test removes the defect.
Sign up free to answer and see why

Checkpoint

20 of 100 reviewed cases are positive; 900 cases are unreviewed. Which rate is directly observed?

A98% confirmed population negatives.B2% population prevalence.C20% among reviewed cases.D20% population prevalence.
Sign up free to answer and see why

Checkpoint

Why record the prior review policy?

AIt helps explain which outcomes were selected for observation.BIt removes the need for label maturity.CIt makes all unreviewed cases usable as negatives.DIt guarantees unbiased labels.
Sign up free to answer and see why

Checkpoint

60 tickets have complete mature logs: 15 escalated, 45 not. Another 40 are six hours old for a 24-hour target. Which extraction matches the complete-window protocol?

A15 positives and 45 negatives, but use the 40 immature tickets to choose the best threshold.B15 positives and 85 negatives with lower weights on the recent tickets.C15 positives and 85 negatives at equal weight.D15 positives, 45 negatives, and 40 immature exclusions.
Sign up free to answer and see why

Explain how you would distinguish mature negatives from delayed, selected, and unreliable observations. Use one supplied fixture and identify a condition that would invalidate your conclusion. Rate confidence from 1 to 5.

Not yetGetting thereConfident

Wrap-up

  • Document how outcomes become labels. Missing, immature, and negative are different states.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.

Labels can be missing for a reason · Training data…