Lesson 1 of 4 · 35 min

Name the claim before selecting the experiment

Write a hypothesis with a population, comparison, outcome, and failure condition.

A research question becomes testable when it identifies what changes, what remains fixed, and what observation would count against the idea. "The method is better" hides too many choices. Better on which population, under what budget, by which metric, and compared with which baseline? A precise claim makes both success and failure informative.
Define the estimand, the quantity the study intends to estimate. It might be the mean accuracy difference between two training procedures across a specified distribution of seeds and datasets. It might be the effect of a product intervention on user outcomes. These are different quantities, even if both produce a percentage. The design and analysis must match the chosen target.
Separate mechanism claims from performance claims. A method can improve a benchmark without proving the proposed explanation for its improvement. If the hypothesis says a component helps through better use of long-range information, an aggregate score alone is insufficient. A targeted intervention or diagnostic task should test the proposed mechanism while controlling obvious alternatives.
State practical relevance before seeing the result. A tiny metric gain may be too small to justify a large cost increase. Conversely, a similar score at half the compute may be useful. Report quality and resources under a defined comparison rather than forcing every result into a single leaderboard rank.

Worked example

An invented proposal says, "Our augmentation improves robustness". The refined claim is: "Under a fixed 20,000-step training budget, augmentation A improves mean classification accuracy on held-out camera devices relative to the same model without A, across five planned seeds, without reducing in-domain accuracy by more than one point."
The experiment now needs a held-out-device split, an in-domain set, paired training conditions, and a predeclared seed list. A result that improves only in-domain accuracy fails the stated robustness claim. A result that improves held-out devices but loses five in-domain points violates the guardrail. The protocol supports a bounded claim about these devices and conditions, not every distribution shift.

Exercise and solution

Rewrite "This retrieval method makes reasoning better" into a testable claim. You have two retrieval methods, 200 fixed questions, a verified answer rubric, and a fixed token budget.
A strong answer specifies the answer-quality difference on those questions under the same model and token budget, with retrieval as the changed factor. If the claim concerns reasoning mechanisms, it adds a separate diagnostic rather than equating final correctness with internal reasoning. Award one point each for population, controlled comparison, measurable outcome, budget, and a failure condition. A method that uses twice the context has changed both retrieval and resources unless that tradeoff is the stated target.

Lab artifact: a claim card

Use the following original template before a result exists. It separates the target quantity from the measurement procedure.
code
1Question: Does augmentation A help on new camera devices?2Target population: examples from the stated held-out camera family3Treatment: training procedure with augmentation A4Comparison: same procedure with A disabled5Primary outcome: accuracy difference, candidate minus baseline6Training budget: 20,000 completed updates at fixed global batch size7Variation: five planned paired initializations, all reported8Guardrail: in-domain accuracy loss no greater than one percentage point9Failure: no target gain, guardrail violation, or invalid data boundary10Mechanism claim: not established by accuracy difference alone
This card still needs details. “Camera family” must identify devices or a sampling process. Five initializations quantify one source of variation conditional on the chosen datasets; they do not create five independent camera populations. If held-out devices are selected for convenience, the population claim should be limited accordingly. The guardrail also needs an uncertainty rule: an observed loss below one point is not the same as evidence that the true loss is below one point. A confirmatory design should define how estimates and uncertainty enter the decision.
The card's fixed update count assumes equal global batch and input lengths. If augmentation changes the number of views per source image, equal updates may no longer imply equal compute or source exposure. Record the resource unit that matters to the research question. A practical recipe comparison can allow different resource use and report the tradeoff; a component-effect claim requires clearer control.

A second failure case: the metric weights the wrong population

Suppose device A contributes nine hundred images and device B one hundred. Candidate improves accuracy by one point on A and loses five points on B. An image-weighted difference is 0.9 × 1 + 0.1 × (−5) = 0.4 points. An equal-device difference is (1 − 5) / 2 = −2 points. Both calculations are correct for different quantities.
Target quantityWeight for AWeight for BDifference
Image mixture in this dataset0.90.1+0.4 points
Equal weight per device0.50.5-2 points
If the goal is average performance for a randomly chosen image from the current mixture, image weighting may be appropriate. If each device is intended to count equally, macro averaging answers that question. Neither weighting should be selected after observing which favors the method. State the estimand first, then calculate it. This numerical counterexample is original reasoning; the reporting checklist supplies context for explicit assumptions, not this specific formula.

Exercise: separate improvement and mechanism

A retrieval change improves answer accuracy at the same token budget, but a diagnostic that removes all long-range dependencies leaves the gain unchanged. The proposed mechanism was better use of long-range dependencies. What can be concluded?
The measured performance gain may remain supported under the evaluation conditions. The mechanism interpretation is weakened because the gain persists when the proposed dependency is absent. Alternative explanations include better local evidence selection or other changed properties, but those remain hypotheses. Design a comparison that varies the claimed mechanism while keeping obvious alternatives controlled. Award one point for preserving the valid performance observation, one for limiting the mechanism claim, one for alternatives labelled as hypotheses, and two for a discriminating design.

Misconceptions to correct

“A precise metric name defines the whole claim” fails because population, weighting, training budget and selection rule can all change its meaning. “A failed mechanism test erases every performance result” fails because performance and explanation are separate claims. Revise the unsupported interpretation while preserving observations that remain valid.
Before running the study, ask a colleague to restate the claim from the card. If they interpret “robustness” as all distribution shifts while you mean two held-out cameras, the wording is still too broad. Replace the umbrella term with the actual tested condition. This makes the eventual result easier to defend in a research interview and harder to overstate in a paper.

Interview probe

Original practice: What result would make you abandon your hypothesis? A strong answer names a predeclared observable failure, not merely lack of enthusiasm. Follow up with a performance gain that contradicts the proposed mechanism. A weak answer rewrites the claim after every outcome.

Sources

docsNeurIPS: paper checklistneurips.ccdocsGoogle DeepMind: roles and interview processdeepmind.google

Checkpoint

Which claim specifies a falsifiable comparison?

AThe model is more capable in general.BThe newer architecture should be preferred.CThe method improves a named held-out metric under a fixed budget and baseline.DThe method produces impressive examples.
Sign up free to answer and see why

Checkpoint

Device A has weight 0.9 and gain +1 point; B has weight 0.1 and gain -5 points. Image-weighted difference?

A-2 points.B+0.4 points.C+1 point.D-4 points.
Sign up free to answer and see why

Checkpoint

The equal-device mean in the same example is -2, while image weighting gives +0.4. Which is correct?

ABoth, for different stated quantities; choose weights from the target before results.BOnly the positive result because most images favor it.COnly the negative result because macro averaging is always required.DNeither because metrics cannot disagree.
Sign up free to answer and see why

Checkpoint

A gain persists after the proposed long-range mechanism is removed. What follows?

AThe mechanism is proven because performance remains high.BEvery performance observation is invalid.CThe diagnostic must be discarded.DThe performance result may remain, but the proposed mechanism needs revision or stronger evidence.
Sign up free to answer and see why

Checkpoint

An observed guardrail loss is 0.8 points with a predeclared one-point allowable true-loss limit. What analysis is still required before a confirmatory acceptance decision?

ACheck that the point estimate remains below one point after rounding to the reporting precision.BApply a predeclared uncertainty rule that compares an appropriate bound with the one-point margin.CTest whether the loss differs from zero, and accept whenever that test does not reject.DIncrease the allowable margin to the upper confidence bound calculated from this result.
Sign up free to answer and see why

Can you write the target quantity, weighting, comparison and falsification condition before seeing results? Rate confidence from 1 to 5 and separate one performance claim from its proposed mechanism.

Not yetGetting thereConfident

Wrap-up

  • Write the quantity and conditions you intend to estimate. Keep mechanism, quality, and resource claims distinct.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.