Lesson 2 of 4 · 35 min

Why precision falls when the world changes

Calculate precision from prevalence and conditional error rates.

Precision depends on the population being scored. Even if a classifier's true-positive rate and false-positive rate remain unchanged, a lower positive prevalence can reduce precision sharply. This is a useful interview calculation because it separates model discrimination from the consequences of deployment context.
Let prevalence be the fraction of truly positive examples. Multiply it by the true-positive rate to obtain the fraction of all examples that are true positives. Multiply the negative prevalence by the false-positive rate to obtain the fraction that are false positives. Precision is the first quantity divided by their sum. Write the counts for an imagined population if the formula feels abstract.
The calculation assumes the conditional rates transfer to the new population. That may be false when the new traffic differs in other ways. Treat it as a diagnostic baseline, then inspect actual labels and score distributions. A prevalence shift can also affect calibration, so old probability values may no longer support the same expected-cost calculation.
Do not respond by changing the threshold blindly. A higher threshold may improve precision while reducing recall. If the application has a review-capacity constraint, the threshold and workload need joint evaluation. If the cost of a missed positive changes, the preferred decision may move in another direction. State the policy objective before choosing a remedy.

Worked example

A teaching classifier has true-positive rate 0.8 and false-positive rate 0.1. In 1,000 cases with 10% prevalence, there are 100 positives and 900 negatives. Expected true positives are 80; false positives are 90. Precision is 80/170, about 47.1%.
Now use 1% prevalence in another 1,000 cases while assuming the same conditional rates. There are ten positives and 990 negatives. Expected true positives are eight; false positives are 99. Precision becomes 8/107, about 7.5%. The model's conditional rates did not change in this scenario, yet most alerts are now false. A report that calls this proof of a coding regression is too strong.

Exercise and solution

A classifier has true-positive rate 0.9 and false-positive rate 0.02. At 5% prevalence in 10,000 examples, calculate expected true positives, false positives, and precision.
There are 500 positives and 9,500 negatives. Expected true positives are 450 and false positives are 190. Precision is 450/640, about 70.3%. Award one point each for population counts, conditional counts, precision, and the transfer assumption. The numbers are expected counts, not a guarantee for a particular sample. Explain that a deployment comparison needs mature labels from the new population.

Lab artifact: derive the denominator

Write the confusion table in fractions of the whole population before substituting a sample size:
TruthPredicted positivePredicted negative
Positive, fraction pp × TPRp × (1 − TPR)
Negative, fraction 1 − p(1 − p) × FPR(1 − p) × (1 − FPR)
Precision is p × TPR divided by [p × TPR + (1 − p) × FPR]. The denominator is the alert fraction, so it also gives expected workload after multiplication by traffic volume. Under prevalence one percent, TPR 0.8 and FPR 0.1, the alert fraction is 0.107. At ten thousand daily cases, the expected workload is 1,070. A team with capacity two hundred cannot use that operating point unchanged.
This is original arithmetic under fixed conditional-rate assumptions. scikit-learn's metric definitions support the count relationships; they do not guarantee that rates transfer across deployments. If a new region differs in language, device mix, or labeling practice, the old TPR and FPR may be wrong as well as the prevalence.

A second failure case: a balanced test set hides workload

A team creates a test set with five hundred positive and five hundred negative examples to get enough error examples for analysis. At TPR 0.8 and FPR 0.1, it records four hundred true positives and fifty false positives, giving precision 88.9%. This precision applies to the fifty-percent-positive sample. It does not directly describe a deployment population with one-percent prevalence.
A class-stratified evaluation can estimate conditional rates if sampling within each class is representative and label quality is sound. It then needs the target prevalence for the workload calculation. The balanced table is useful; the mistake is treating its artificial class mixture as production. If selection within classes favors easier cases, even conditional rates may be biased.
python
1def expected_alerts(n, prevalence, tpr, fpr):2    tp = n * prevalence * tpr3    fp = n * (1 - prevalence) * fpr4    return {"tp": tp, "fp": fp, "alerts": tp + fp,5            "precision": tp / (tp + fp) if tp + fp else None}6# Inputs are assumptions or measured estimates, not universal model constants.

Exercise: compare two thresholds under changed prevalence

The high threshold has TPR 0.6 and FPR 0.005. The lower threshold has TPR 0.85 and FPR 0.03. There are ten thousand cases, prevalence two percent, and capacity two hundred reviews. Assume both sets of conditional rates transfer.
High yields 120 true positives and 49 false positives, or 169 reviews with precision about 71.0%. Lower yields 170 true positives and 294 false positives, or 464 reviews with precision about 36.6%. High fits the stated average daily capacity; lower does not. If false negatives are costly, that still does not make an infeasible queue feasible. Explore a defined intermediate policy or another action under valid evaluation. Award one point for each threshold's counts, one for precision, and two for capacity and assumption limits.
Now suppose prevalence doubles while traffic and conditional rates stay fixed. High gives 240 true positives and 48 false positives: 288 reviews. A threshold that fit previously can overload the queue even though its false-positive rate improved nothing and its code did not change. This is why monitoring workload and prevalence matters alongside model versions.

Misconceptions to correct

“Stable ROC behavior means stable alert quality” ignores the prevalence-dependent precision denominator. “A balanced evaluation is always invalid” discards a useful design for estimating class-conditional errors. The correct question is which population each statistic describes and which weighting assumptions permit transfer.
End an interview answer by asking for mature labels in the new population and the intended action. A score meant to inform a human may tolerate different errors from an automatic block. Your calculation should make this decision explicit, with uncertainty where estimated rates are based on limited samples.

Interview probe

Original practice: Precision collapsed, but the ROC curve looks similar. Is that possible? A strong answer computes the effect of prevalence and then checks whether conditional-rate stability is plausible. Follow up with a fixed review budget. A weak answer insists one metric must be wrong.

Sources

docsscikit-learn: model evaluation metricsscikit-learn.orgdocsscikit-learn: probability calibrationscikit-learn.org

Checkpoint

At fixed positive TPR and nonzero FPR, prevalence falls. What happens to precision under those assumptions?

AIt rises because fewer positives exist.BIt is determined by TPR alone.CIt stays fixed with ROC ordering.DIt falls as false alerts occupy more of the alert population.
Sign up free to answer and see why

Checkpoint

A balanced test set has TPR 0.8 and FPR 0.1. Which quantity is needed to project deployment precision?

AThe precision from the balanced sample, with no population adjustment.BDeployment prevalence plus a justified transfer of conditional rates.CDeployment traffic volume alone, while retaining balanced-sample precision.DThe training class ratio, assumed equal to deployment.
Sign up free to answer and see why

Checkpoint

At prevalence 2%, N=10,000, TPR 0.6 and FPR 0.005, expected alerts equal?

A169B120C49D220
Sign up free to answer and see why

Checkpoint

The preceding policy has capacity 200. Prevalence doubles with fixed rates. Expected alerts become 288. What follows?

AThe classifier's code must have changed.BCapacity is still satisfied because FPR is unchanged.CThe old feasible operating point now exceeds capacity.DThe alert count must remain 169 because N is fixed.
Sign up free to answer and see why

Checkpoint

When can a class-balanced sample estimate useful TPR and FPR?

AAlways, even if positives were selected for easy examples.BWhen within-class sampling and labels support the target conditional distributions.COnly when deployment prevalence is exactly 50%.DNever, because class balancing invalidates every metric.
Sign up free to answer and see why

Can you turn conditional rates into counts, precision, and workload under a new prevalence? Rate confidence from 1 to 5 and name the assumption you must verify.

Not yetGetting thereConfident

Wrap-up

  • Translate rates into counts under the deployment prevalence. Then check whether the conditional-rate assumption holds.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.