Ablations need controls that answer the explanation
Use a two-factor comparison to identify an interaction.
Removing a component and observing a lower score shows that the component matters in that configuration. It does not automatically explain why it matters. The component may add parameters, compute, data, regularization, or a changed optimization path. A mechanism claim needs controls that distinguish these alternatives.
Start with the proposed explanation and its competing explanations. If a memory module supposedly helps through retrieving relevant history, compare against a parameter-matched control, a shuffled-history control, and a task slice where history should matter. These controls are informative only if their construction preserves the relevant conditions. A shuffled control that also changes sequence length confounds two factors.
Factorial designs examine combinations of changes. With two binary factors, evaluate neither, each alone, and both together. The joint improvement may differ from the sum of individual improvements. That interaction can be scientifically interesting, but it needs uncertainty estimates and a stable protocol. Do not infer it from one selected seed per cell.
Ablations also need fair tuning. If the full method receives extensive tuning but the ablated model uses an obviously poor default, the gap may reflect effort rather than the component. Decide whether the claim compares fixed implementations or reasonably tuned procedures, then allocate and report the selection budget accordingly.
Worked example
This invented table uses a fixed metric where higher is better.
Memory M
Augmentation A
Score
Off
Off
70
On
Off
72
Off
On
73
On
On
78
The memory-only gain is two points and augmentation-only gain is three. Adding both gives eight points over baseline. The interaction contrast is 78 minus 72 minus 73 plus 70, or three points. Under this simple table, the joint effect exceeds the additive expectation by three.
The table does not establish a causal mechanism by itself. Each cell needs comparable data, resources, tuning, and repeated runs. A follow-up can examine whether augmentation creates examples where memory is especially useful. That is a new testable explanation, not a conclusion already proven by the interaction arithmetic.
Exercise and solution
Four conditions score 80 for neither, 83 for factor X, 82 for factor Y, and 84 for both. Calculate the interaction contrast and interpret it.
The contrast is 84 minus 83 minus 82 plus 80, or negative one point. The joint gain is one point below the simple additive expectation. Award one point each for the arithmetic, non-additivity interpretation, uncertainty requirement, and a plausible control. Do not call the factors incompatible without checking variability and the metric's scale. An interaction on one metric scale may look different on another.
Lab artifact: derive the contrast from conditional effects
The interaction contrast can be understood without memorizing a formula. In the first table, adding memory when augmentation is off changes score from seventy to seventy-two, a gain of two. Adding memory when augmentation is on changes seventy-three to seventy-eight, a gain of five. The difference between these memory effects is three.
code
1Memory effect with augmentation off: 72 - 70 = 22Memory effect with augmentation on: 78 - 73 = 53Difference of effects: 5 - 2 = 34Raw interaction contrast: 78 - 72 - 73 + 70 = 3
This is a raw difference-of-differences contrast in score units. It is not automatically the coefficient of a regression term. Under a model with both factors coded as minus one and plus one, the interaction coefficient is one quarter of this raw contrast. Under zero/one coding, the interaction coefficient equals the raw contrast. State coding before comparing numbers from a statistical package. NIST's factorial-design and interaction materials give direct context for this distinction; these numbers are original teaching calculations.
The scale also matters. A three-point interaction in accuracy is not necessarily a three-unit interaction after a nonlinear transformation. Near a metric ceiling, additive gains may be constrained. Select the outcome scale based on the scientific question, not on which transformation produces the most interesting interaction.
A second failure case: a shuffled control changes difficulty
A memory system stores ten prior facts. A proposed control shuffles their order but also truncates the list to five. If performance drops, the result cannot isolate order sensitivity because information quantity also changed. Keep count, token budget, formatting, and relevant facts controlled when those are not the target of the intervention.
Control
Intended factor
Potential confound to audit
Shuffle order
Order dependence
Changed truncation or position distribution
Replace relevant history
Content relevance
Changed length or lexical cues
Parameter-matched network
Extra capacity
Different optimization or compute
Fixed retrieval count
Selection quality
Different total token budget
No control is perfect by its name alone. A parameter-matched model can still have different compute or inductive bias. The report should explain what alternative the control tests and what remains unmatched. Multiple complementary controls can narrow the explanation without pretending to prove every detail of an internal mechanism.
Exercise: calculate a replicated contrast
Two prespecified matched runs produce cell scores [60, 63, 62, 67] and [61, 62, 64, 66] in the order neither, memory only, augmentation only, both. The first raw interaction is 67 − 63 − 62 + 60 = 2. The second is 66 − 62 − 64 + 61 = 1. Their mean contrast is 1.5 points.
Report both contrasts and the small number of independent matched runs. Do not treat the eight cell scores as eight independent estimates of the interaction: the contrast combines four conditions within each matched unit. More units or a justified model are needed for a precise uncertainty claim. Award two points for arithmetic, one for the unit, and two for uncertainty and avoiding selected cells.
Misconceptions to correct
“A positive interaction means both factors are individually beneficial everywhere” fails because conditional effects can vary and the observed cells cover only specified settings. “An ablation with fewer parameters proves the missing mechanism matters” leaves capacity as a competing explanation. A useful ablation asks a specific question and records the remaining alternatives.
The final contrast sheet should contain every cell, its budget and tuning rule, each matched unit's contrast, and the declared metric scale. If one cell crashes, do not replace it with the best successful run from another seed and preserve the pairing label. Record the missing or failed condition and decide how the planned analysis handles it. This connects mechanism study design to the audit discipline required for trustworthy results.
Interview probe
Original practice: An ablation reduces performance. What has it proved? A strong answer limits the claim to the tested configuration and considers capacity, compute, tuning, and data confounds. Follow up with a factorial table. A weak answer treats any ablation gap as proof of the proposed mechanism.
A shuffled-history control also halves the history length. Which interpretation is supported by that comparison alone?
AIt measures the combined order-and-length change, but cannot attribute the score change specifically to order.BIt isolates order because the same underlying documents appear in both conditions.CIt isolates order if the model weights and random seed are identical.DIt isolates order once enough independent runs make the score difference precise.
The full method gets extensive tuning; an ablation gets one default. What needs stating?
AThe gap proves the component's mechanism.BSelection effort is a competing explanation unless the claim intentionally compares those fixed recipes.CTuning never affects ablation interpretation.DThe ablation result cannot be recorded at all.
Can you compute conditional effects and the raw interaction, then identify a competing explanation that remains? Rate confidence from 1 to 5 and state the factor coding.
Not yetGetting thereConfident
Wrap-up
Design controls around explanations. Use interactions as evidence to investigate, with comparable protocols and uncertainty.