Lesson 2 of 4 · 35 min

Pair the evidence and respect the sampling unit

Analyze a small paired result without inflating the sample size.

Two methods evaluated on the same tasks can be compared through paired differences. Pairing removes some variation caused by task difficulty because each task acts as its own comparison. But the unit of independence still matters. Ten answers from one user are not automatically ten independent users, and ten checkpoints from one training run are not ten independent training runs.
Choose the resampling unit to match the estimand and dependence structure. If the claim concerns variability across trained models, resample runs rather than individual predictions alone. If the data contains user clusters, a user-level procedure may be needed. A bootstrap that resamples dependent rows as if independent can produce intervals that are too narrow.
A permutation test also needs a valid null and exchangeability assumptions. In a paired sign-flip test, the signs of differences are randomized under a suitable symmetry or exchangeability condition. The resulting p-value describes how unusual the observed statistic is under that null procedure. It is not the probability that the research hypothesis is true.
Small samples limit resolution. With four paired observations, there are only sixteen possible sign assignments in an exact sign-flip enumeration. No amount of confident prose turns four independent units into forty. Report the sample size and the test's assumptions alongside the result.

Worked example

The following invented paired differences are method B minus A on four independent experimental units.
UnitDifference
1+1
2+2
3+1
4+2
The mean is 1.5. Under an exact sign-flip null with all sixteen assignments equally valid, only all-positive and all-negative assignments have an absolute mean at least 1.5. The two-sided p-value is 2/16, or 0.125. All observed differences favor B, but this small test does not meet a conventional 0.05 threshold.
This does not prove there is no effect. It shows that this design has limited evidence under the stated test. The effect estimate and its practical meaning remain useful, and additional independent units may improve precision.

Exercise and solution

A researcher evaluates two models on 100 messages from each of five users. They report n = 500 independent observations for a user-level claim. What should be checked?
Check within-user dependence and the target population. A user-level resampling or hierarchical analysis may be appropriate, and the effective number of independent users is only five for that target. Award one point each for identifying the cluster, linking unit to estimand, avoiding a false independence claim, and reporting the resulting uncertainty. Do not select a procedure solely because it gives a smaller p-value.

Lab artifact: enumerate the exact toy null

The four-unit example is small enough to inspect without an asymptotic approximation. Under the stated sign-flip null, keep absolute differences [1,2,1,2] fixed and enumerate all sixteen sign vectors.
python
1from itertools import product2magnitudes = [1, 2, 1, 2]3observed = 1.54null_means = [5    sum(s * d for s, d in zip(signs, magnitudes)) / 46    for signs in product([-1, 1], repeat=4)7]8p_two_sided = sum(abs(v) >= observed for v in null_means) / len(null_means)9# 2 / 16 = 0.125, for this complete exact enumeration and stated null.
This uses an absolute-mean tail definition. Other software conventions for a two-sided p-value can differ in discrete or asymmetric settings, so state the statistic and tail rule. The enumeration includes the observed assignment among all possibilities. It is different from drawing a limited Monte Carlo sample, for which the estimator and any plus-one convention need their own specification.
The sign-flip procedure is justified only when the sign assignments are valid under the null model or randomization design. Paired observations alone do not automatically establish that assumption. For observational algorithm scores, symmetry of paired errors around zero may be an assumption rather than a consequence of randomized assignment. State that assumption and consider whether the design supports it.

A second failure case: a cluster changes the target

Two users produce different numbers of messages. User A has one hundred messages with a method difference of plus one on every message. User B has ten messages with a difference of negative one on every message. The message-weighted mean is ninety divided by one hundred ten, about 0.818. The equal-user mean is zero.
UserMessagesMean difference
A100+1
B10-1
If the target is a randomly selected user, the equal-user quantity is relevant. If the target is a randomly selected message from this traffic mixture, the message weighting may be relevant, but dependence still affects uncertainty. Choosing the resampling unit and choosing the metric weights are related but distinct decisions. A cluster bootstrap can resample users while retaining their messages and recomputing the desired weighted statistic. It does not automatically force equal-user weighting unless the statistic does so.

Exercise: define a paired bootstrap procedure

There are fifty independent documents, with two model scores per document. The claim concerns the mean within-document score difference for this document population. Describe one paired bootstrap and a mistake it avoids.
Form the fifty paired differences or keep the score pairs intact. Draw document indices with replacement, use the same sampled indices for both methods, and calculate the mean difference for each replicate. Use an interval method appropriate to the analysis plan, noting that bootstrap intervals can behave poorly with very small or degenerate samples. Sampling method A's documents independently from method B's would destroy the observed pairing. Award two points for the unit and paired resampling, one for the statistic, and two for the interval assumptions and avoided error.
If those documents are actually chapters from five books, the independence claim needs review. A book-level cluster procedure or a model of the hierarchy may be more appropriate. Resampling fifty chapters as independent books does not create information about fifty independent source populations.

Misconceptions to correct

“A p-value of 0.125 means a 12.5% probability the null is true” reverses the conditional interpretation. The calculation concerns data extremeness under the specified null procedure. “A non-significant result proves equivalence” ignores effect size, uncertainty and power. Equivalence needs a defined margin and suitable analysis, not merely failure to reject zero difference.
For an interview, write the independent unit, pairing rule, statistic, null or sampling assumptions, and conclusion before discussing software. A correct library call with the wrong unit gives a precise answer to the wrong question. The small table is valuable because it lets you see the target and dependence before trusting a large output report.

Interview probe

Original practice: Why can paired evaluation be more informative than comparing two separate means? A strong answer uses shared task difficulty and then states the correct independence unit. Follow up with repeated outputs from one model run. A weak answer treats every logged row as independent.

Sources

docsSciPy: bootstrap confidence intervalsdocs.scipy.orgdocsSciPy: permutation testsdocs.scipy.org

Checkpoint

One training run saves ten checkpoints. For across-run variability, how many independent training runs were observed?

ATen.BOne.CTen times the evaluation-set size.DThe count of optimizer steps.
Sign up free to answer and see why

Checkpoint

Four differences [1,2,1,2] have two of sixteen sign assignments as extreme in absolute mean. Exact two-sided p under the stated null?

A0.0625B0.25C0.125D0.5
Sign up free to answer and see why

Checkpoint

A paired bootstrap for a mean score difference on independently sampled documents should resample what?

AA's document scores and B's document scores with separate sampled indices.BScore pairs sampled without replacement into groups smaller than the original dataset, using their means as ordinary bootstrap replicates.CDocuments sampled with probability proportional to their observed improvement, while retaining each pair.DDocument indices with replacement, using the same sampled indices for both methods and retaining each pair.
Sign up free to answer and see why

Checkpoint

User A has 100 messages with difference +1; B has 10 with -1. Equal-user mean difference?

A0.B0.818.C1.D-1.
Sign up free to answer and see why

Checkpoint

What does a non-significant difference establish by itself?

AEquivalence within any useful margin.BThat the effect is exactly zero.COnly that the specified test did not reject; practical equivalence needs a suitable margin and analysis.DThat the sample should be extended until significance appears.
Sign up free to answer and see why

Can you identify the independent unit, retain pairing, and explain a test's assumptions and limits? Rate confidence from 1 to 5 and name a dependence pattern that would change your analysis.

Not yetGetting thereConfident

Wrap-up

  • Pair comparable observations and choose the correct unit. Report effect size, uncertainty, and the assumptions of the test.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.