Lesson 1 of 4 · 35 min

When the average tells the wrong comparison

Calculate subgroup and aggregate rates to expose a composition effect.

An aggregate metric combines performance with the mixture of cases. If two methods are evaluated on different mixtures, their overall rates may reverse the within-group comparison. This is a form of aggregation bias often discussed through Simpson's paradox. The interview skill is to inspect counts and denominators before attributing the difference to the method.
First verify that the groups and outcomes mean the same thing for both methods. Then compute each subgroup rate and the aggregate from counts. Do not average percentages without considering their denominators. If the intended target population has a known mixture, standardize both methods to that mixture for a descriptive comparison. A causal interpretation requires additional assumptions or a valid assignment design.
A reversal is not always a paradox to "fix". The actual production mixture may be the decision target. In that case, the aggregate can be useful, but the comparison still needs comparable exposure. If a new policy changes who enters each group, conditioning on those groups may introduce another bias. Define the causal or descriptive question before deciding which rate is authoritative.
Small cells need caution. A subgroup with one success out of one case has a 100% observed rate but little precision. Report counts and uncertainty rather than ranking groups by raw percentages. The same discipline applies to model evaluation across languages, task difficulty, and user cohorts.

Worked example

The following invented data evaluates two methods on different task mixtures.
MethodEasy correct/totalHard correct/total
A90/1001/10
B19/2020/100
A's easy rate is 90%, and B's is 95%. A's hard rate is 10%, and B's is 20%. B is better within both groups. Yet A's aggregate is 91/110, about 82.7%, while B's is 39/120, or 32.5%, because A received mostly easy tasks.
Under a common 50/50 easy-hard mixture, A's standardized rate is 50%, and B's is 57.5%. This descriptive standardization supports the within-group pattern. It does not prove that B causes better outcomes if assignment or task labels have other confounds.

Exercise and solution

A system reports 80% on 100 short tasks and 40% on 100 long tasks. Another reports 85% on 20 short tasks and 45% on 180 long tasks. Compute both aggregates and a common 50/50 standardized comparison.
The first aggregate is 60%. The second has 17 plus 81 correct out of 200, or 49%. Under 50/50 weighting, the second is 65%, above the first's 60%. Award one point each for both aggregates, standardized rate, composition explanation, and a limit on causal interpretation.

Lab artifact: standardize to the decision population

A common fifty-fifty mixture is convenient, but it is not automatically the deployment target. Suppose the intended workload contains twenty percent easy tasks and eighty percent hard tasks. Using the worked example's conditional rates, A's standardized rate is 0.2 × 0.9 + 0.8 × 0.1 = 0.26. B's is 0.2 × 0.95 + 0.8 × 0.2 = 0.35.
MixtureA standardized rateB standardized rate
50% easy, 50% hard50%57.5%
20% easy, 80% hard26%35%
The ranking agrees here, but the expected operational success rate changes greatly. A method that looks adequate under the convenient balanced mixture may miss a requirement under the real workload. Standardization is useful only if the subgroup rates are transportable to the target mixture and the groups are defined consistently. New tasks within the “hard” category may differ from the measured hard tasks.
A small calculation function makes the weighting rule inspectable:
python
1def standardized_rate(group_rates, target_weights):2    assert set(group_rates) == set(target_weights)3    assert abs(sum(target_weights.values()) - 1.0) < 1e-94    return sum(target_weights[g] * group_rates[g] for g in group_rates)5# Rates must be estimated for comparable groups; this is not a causal adjustment by itself.

Transfer calculation: a target group has no observations

A known target mixture does not supply missing group outcomes. Suppose the target contains 70% familiar tasks and 30% new tasks. A method succeeds on 80% of the familiar tasks, and no new tasks have been measured. If familiar-task performance transfers and the unknown success rate can be anywhere from zero to one, total target success is 0.7 × 0.8 + 0.3 × unknown. Its arithmetic bounds are 56% to 86%. These are bounds from missing information, not a confidence interval for sampling error.
Renormalizing to the observed group gives 80%, but that answers the familiar-task question. Assigning the unknown group a 50% rate gives 71%, but this adds an unsupported assumption. To compare two methods, construct each bound under explicit assumptions. If their bounds overlap, the supplied information may permit either ranking. A common unknown rate, or an assumption that one method cannot be worse in the new group, would constrain the comparison further. Neither assumption follows from observing the familiar group. Collect the missing target-group evidence before reporting a single full-target estimate unless the decision explicitly accepts a stated assumption.

A second failure case: conditioning on a changed outcome

Suppose a recommendation policy affects whether a session is classified as “high engagement.” Comparing success only among high-engagement sessions can select different people or tasks under the two policies. The group is partly a consequence of the policy, so the within-group comparison is not automatically the causal effect of the policy. A reversal does not tell you which conditioning set is causally correct.
Prefer pre-assignment characteristics when they match the intended question, and use the experiment's assigned population for a primary causal comparison under a valid design. If the scientific question is specifically about a post-treatment subgroup, it needs more careful assumptions or methods. The arithmetic of weighted averages alone cannot settle it.
Microsoft's experimentation team gives a concrete warning about segments based on activity during an experiment: the treatment can change a user's active days and thus change segment membership. Its first-party guidance favors segments whose membership is independent of treatment. This supports the selection warning above. It does not establish that balanced segment counts alone prove causal comparability.

Exercise: distinguish stable performance from changed mixture

A service has two task types. In both months its success rates remain ninety percent for type E and fifty percent for type H. Month one has eight hundred E and two hundred H tasks; month two has two hundred E and eight hundred H tasks. Calculate the two aggregates and a standardized comparison using month one's mixture.
Month one gives 720 + 100 = 820 successes out of one thousand, or eighty-two percent. Month two gives 180 + 400 = 580, or fifty-eight percent. At the fixed eighty-twenty mixture, both months give eighty-two percent. The operational aggregate worsened because the mixture changed, while the supplied within-type rates did not. Award one point for each aggregate, one for standardization, and two for separating operational impact from within-type regression.
The conclusion does not prove that the system is healthy for the new workload. The greater share of hard tasks can still require a better policy or more resources. Explaining the source of a metric change is not the same as deciding that the change is acceptable.

Misconceptions to correct

“Always report the aggregate” ignores mixture differences when the goal is a like-for-like method comparison. “Always adjust for every available group” ignores causal structure and treatment-induced membership. The correct denominator depends on the target question and design.
When an interviewer presents an impossible-looking table, calculate counts before speculating. Then ask whether groups are pre-existing, whether methods saw comparable cases within groups, and which mixture represents the decision. A clear numerical explanation plus a bounded causal statement is more useful than naming Simpson's paradox without checking the table.

Interview probe

Original practice: A model wins every subgroup but loses overall. Is the table impossible? A strong answer checks different subgroup weights and calculates counts. Follow up with treatment-induced subgroup membership. A weak answer averages percentages or assumes one table must be fabricated.

Sources

  1. 01Microsoft Research: stable and treatment-dependent experiment segments, opened 2026-09-26. Primary source for the membership warning. The numerical aggregation and missing-group examples are original derivations.
  2. 02scikit-learn: model evaluation metrics
docsMicrosoft Research: stable and treatment-dependent experiment segmentsmicrosoft.comdocsscikit-learn: model evaluation metricsscikit-learn.org

Checkpoint

How can a method win within each subgroup but lose in the pooled score?

AThe group labels alone force a reversal.BDifferent subgroup mixtures can change the weighted average.CEvery pooled rate is an unweighted mean.DA shared identical mixture necessarily reverses the ordering.
Sign up free to answer and see why

Checkpoint

At a 20/80 easy/hard target mixture, A has rates 90%/10%. Standardized rate?

A50%B82.7%C26%D18%
Sign up free to answer and see why

Checkpoint

The target mixture changes but within-group rates stay fixed. What can happen?

AThe operational aggregate can change without within-group regression.BThe aggregate must remain fixed.CEvery group rate must change by the same amount.DThe old mixture becomes the only valid operational denominator.
Sign up free to answer and see why

Checkpoint

A randomized policy changes whether users enter the post-assignment high-engagement group. What limits a causal comparison restricted to that group?

AOriginal random assignment guarantees that the selected high-engagement groups remain comparable.BMatching the two selected groups to the same size restores the original randomization.CAdding more users removes any bias caused by treatment-dependent group membership.DTreatment can change who enters the group, so the selected comparison need not preserve the randomized population contrast.
Sign up free to answer and see why

Checkpoint

What does standardization to a target mixture assume for its descriptive projection?

AThe target weights alone determine success regardless of group rates.BGroup definitions and estimated conditional rates are relevant to the target cases.CEvery subgroup has equal size in the observed sample.DThe adjustment by itself proves a causal effect.
Sign up free to answer and see why

Can you reconstruct aggregates and a declared target mixture, then identify whether a group is treatment-induced? Rate confidence from 1 to 5 and state the denominator for the intended decision.

Not yetGetting thereConfident

Wrap-up

  • Inspect denominators and mixtures before explaining an aggregate. Keep descriptive standardization separate from causal claims.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.