Lesson 2 of 4 · 35 min

Ranking quality does not make scores trustworthy probabilities

Evaluate calibration with a concrete prediction group.

A model can rank risky examples well while giving poor probability estimates. Ranking asks whether positives tend to receive higher scores than negatives. Calibration asks whether examples assigned a probability near p have an observed positive frequency near p. These properties support different decisions.
If an application uses scores only to select the top 100 items, ranking may dominate. If it computes expected loss or gives users probability statements, calibration matters directly. A monotonic transformation can preserve ranking while changing every probability. For example, replacing each score p with p squared preserves order on the interval from zero to one but changes its numerical meaning.
Inspect calibration on held-out data. Group predictions into bins and compare mean predicted probability with observed frequency. Bins are a diagnostic approximation: results depend on binning and sample size. A small high-risk bin can look badly calibrated through random variation. Use proper scoring rules such as Brier score or log loss as additional evidence, while remembering that these summarize more than one aspect of prediction quality.
A calibrator also learns from data. Fitting it on the same examples used to train the base model can overfit. Use a valid held-out or cross-validation procedure and keep a final evaluation set separate. Calibrating probabilities does not repair incorrect labels, temporal leakage, or distribution shift. It adjusts score meaning under the data and method assumptions.

Worked example

A teaching model assigns 100 examples scores around 0.8, with mean 0.80. Only 50 are positive. The observed rate is 0.50. For a simplified case where all predictions equal 0.8, the mean Brier loss is half of 0.2 squared plus half of 0.8 squared, or 0.34. Predicting 0.5 for this group gives 0.25.
This group result suggests overconfidence. It does not prove that replacing every prediction with 0.5 is a good global model, because other groups may contain useful ranking information. A calibration curve across the score range and an independent evaluation of a fitted calibrator are needed. If the deployment prevalence differs, recheck calibration rather than assuming the old mapping remains valid.

Exercise and solution

A group of 200 predictions has mean probability 0.25 and 50 positives. A second group of 20 predictions has mean probability 0.90 and 14 positives. Describe what can be concluded.
The first group matches its observed rate of 25% at this resolution. The second shows 70% observed positives against 90% predicted, but its smaller sample makes the estimate less precise. Neither group alone establishes calibration for the whole model. Award one point each for the rates, group-specific interpretation, sample-size caution, and rejecting a global claim from two bins.

Lab artifact: probability audit

Inspect this independent evaluation table. Within each group every example has the displayed probability, so its Brier loss can be calculated exactly. This simplifying assumption is essential: a bin mean alone is not enough to calculate the loss of variable individual probabilities.
GroupCountProbability per examplePositive labelsBrier loss
Lower1000.20200.16
Higher1000.80500.34
Combined200mixed700.25
The lower group's loss is 0.2 × 0.8² + 0.8 × 0.2² = 0.16. The higher group's loss is 0.34 as above. Equal group sizes give a combined mean of 0.25. A constant prediction of the overall prevalence, 0.35, gives 0.35 × 0.65² + 0.65 × 0.35² = 0.2275 on these labels. The more detailed model has worse Brier loss in this example despite assigning the higher-risk group a higher score. This shows why a ranking claim does not settle probability quality.
That comparison uses observed labels to explain a metric. It is not permission to choose a constant on the final test set and report the resulting score as a clean estimate. A deployable constant must be estimated from allowed training or calibration data. If this table drove the choice, obtain new evaluation data or use a planned validation design.
A compact audit function keeps the denominator visible:
python
1def brier(probabilities, labels):2    assert len(probabilities) == len(labels) and len(labels) > 03    assert all(0 <= p <= 1 for p in probabilities)4    assert all(y in (0, 1) for y in labels)5    return sum((p - y) ** 2 for p, y in zip(probabilities, labels)) / len(labels)6# Calculate from individual predictions, not only bin means.

A second failure case: good aggregate calibration, bad subgroup meaning

Consider two groups, each with one hundred predictions of 0.5. Group North has eighty positives; Group South has twenty. Across both groups the observed frequency is fifty percent, so the combined bin looks perfectly calibrated. Within groups the same statement is misleading. A risk communication system serving people in either group may need these conditional diagnostics.
This does not mean adding a group-specific calibrator is always justified. First verify group definitions, label quality, sample size, intended use, and whether the subgroup is available and appropriate at decision time. Small groups can overfit a flexible calibrator. Report uncertainty and avoid claiming all conditional calibration is established merely because a few chosen groups look good.

Exercise: numerical scores and unchanged ordering

Four examples have scores 0.1, 0.4, 0.6, 0.9. Squaring gives 0.01, 0.16, 0.36, 0.81. Under strict order and no numeric ties, the ranking is unchanged. At a numerical threshold of 0.5 the original model selects the last two; the squared scores select only the last. To preserve the old actions, transform the threshold too: 0.5 squared is 0.25. This action equivalence does not make squared scores calibrated probabilities.
Award one point for the transformed scores, one for unchanged order, one for changed fixed-threshold actions, and one for the equivalent transformed threshold. A fifth point requires explaining that ranking invariance and probability validity are separate.

Misconceptions to correct

“Lower Brier loss proves calibration improved” is too broad because the score reflects both calibration and the useful separation of outcomes. Use reliability diagnostics as well. “A monotonic transformation leaves all metrics unchanged” confuses order-based measures with fixed-threshold actions and proper probability scores. State the score domain: squaring preserves order on nonnegative probabilities, but it does not preserve order for arbitrary signed margins.
For the release packet, retain the base-model version, calibrator version, calibration population, fit window, evaluation window, and subgroup sample counts. The release must deploy the exact calibrated pipeline that was evaluated. Attaching a new calibrator while retaining the old threshold can change workload substantially, so rerun the capacity checks from lesson one.

Interview probe

Original practice: Can two models have the same ROC AUC and different calibration? A strong answer gives a monotonic score transformation and explains how ranking stays fixed while probability values change. Follow up with expected-cost decisions. A weak answer treats AUC as a probability-quality measure.

Sources

docsscikit-learn: probability calibrationscikit-learn.orgdocsscikit-learn: model evaluation metricsscikit-learn.org

Checkpoint

Scores in [0,1] are squared without ties. Which property can change?

ATheir strict ordering.BThe numerical probability calibration.CThe positive label count.DTheir rank-based ROC ordering.
Sign up free to answer and see why

Checkpoint

A bin has mean probability 0.5 and half positive labels. Can its Brier score be calculated from these facts alone?

AYes, it must be 0.25.BYes, it must be zero.CNo; individual scores and their association with labels matter.DNo; Brier requires a decision threshold.
Sign up free to answer and see why

Checkpoint

Two equal groups predict 0.5 throughout; their positive rates are 80% and 20%. What does the aggregate 50% rate establish?

AEvery group is calibrated.BThe model has perfect ranking.CA separate calibrator is certain to improve both groups.DThe pooled bin agrees, while subgroup calibration differs.
Sign up free to answer and see why

Checkpoint

Scores 0.1, 0.4, 0.6, 0.9 are squared. What transformed threshold preserves decisions from the original threshold 0.5?

A0.25B0.5C0.75D0.05
Sign up free to answer and see why

Checkpoint

A calibrator was selected after examining final-test reliability plots. How should that set now be treated?

AStill untouched because base-model weights did not change.BAs a calibration-training set whose fitted score is an unbiased test estimate.CAs development evidence; obtain valid independent evaluation.DAs an evaluation set if only the threshold changes next.
Sign up free to answer and see why

Can you distinguish a ranking claim, a bin-calibration claim, and a proper-score comparison using the supplied numbers? Rate confidence from 1 to 5 and name one missing assumption.

Not yetGetting thereConfident

Wrap-up

  • Check score order and probability meaning separately. Fit calibration without reusing final evaluation data.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.