Lesson 1 of 4 · 35 min

Reproduction triage: find the first disagreement

Build a comparison ladder for a failed reproduction.

A reproduction can fail because the scientific idea is weak, because the implementation differs, or because the experimental conditions differ. These explanations require different next steps. Begin by making the reference and new run comparable, then find the earliest observable disagreement.
Use a comparison ladder. Check data identities and preprocessing first. Then compare input tensors, forward outputs, loss terms, gradients, and one optimizer update on a tiny shared batch. Only after these match should you compare long-run curves. If the first batch already differs, a week-long rerun is unlikely to provide a useful diagnosis.
Fix the variables that are not under investigation. Use the same seed and environment for a mechanical comparison where possible. Disable stochastic augmentation temporarily if it prevents alignment, then restore it and test its distribution separately. Keep this diagnostic simplification distinct from the final scientific protocol. A model that works only after removing a required component has not reproduced the original method.
Treat missing details as uncertainty. A published result may omit a preprocessing choice or exact selection rule. Record the ambiguity and test a small number of plausible interpretations. Do not silently choose whichever setting gives the best score and then call the result an exact reproduction.

Worked example

A fictional reference reports validation accuracy near 80%, but a new implementation reaches 62%. The first-batch investigation produces this record.
CheckReferenceNew implementation
Sample IDs11, 1811, 18
Pixel range0 to 10 to 255
Logit scaleabout 1about 100
First loss0.828
The first disagreement is preprocessing. Correcting the pixel scaling makes the forward outputs and first update align within tolerance. The engineer now reruns a short training window and later the full protocol. The 18-point gap should not be attributed to a new optimizer until the input mismatch is removed.

Exercise and solution

Two implementations use the same batch and produce matching logits, but one loss is exactly twice the other. The batch has two examples. What should you inspect next?
Inspect sum versus mean reduction, per-example weights, and whether regularization is counted differently. Compare each per-example loss before the final reduction. A strong answer tests these hypotheses with one- and two-example batches. Award one point each for localizing after logits, identifying reduction, a discriminating fixture, and avoiding premature long-run tuning. If gradients also differ by two, that supports but does not alone prove the reduction hypothesis. Record the confirmed cause only after examining the calculation.

Lab artifact: localize a gradient-only disagreement

Consider a second reproduction failure. Two programs use the same input tensor and initial weights. Their logits and per-example losses match. Their scalar losses also match, but one parameter's gradient differs. The investigation should stay at the first unexplained boundary rather than return immediately to data loading.
code
1shared batch hash: B122initial parameter hash: W03maximum logit difference: 0 within fixture precision4per-example losses: [0.3, 0.9] in both5reported scalar loss: 0.6 in both6gradient of parameter u:7  reference: -0.48  candidate:  0.09next check: graph connectivity, detach/in-place path, gradient reset/accumulation
A matching scalar does not imply matching derivatives. A detached tensor can retain the same numerical value while removing a path for gradients. An accidental stop-gradient operation is therefore plausible. So is stale or missing gradient accumulation. A finite-difference check of the actual parameterized forward-loss function can help distinguish whether the intended derivative is nonzero. Inspect the graph and the update rather than assuming any one cause from the trace.
The comparison must use the same objective, including regularization. A program can print a data-only loss while backpropagating data loss plus a penalty. If printed scalars match but gradients differ, ask which scalar was actually differentiated. Logging and optimization may use different expressions.

A second failure case: matching one batch after changing the protocol

To simplify diagnosis, the engineer disables random crops. The two programs then match on one batch. This establishes alignment for the deterministic simplified path. It does not establish that the original stochastic pipeline matches. Restore the crop transform and test its allowed ranges, label-preserving behavior, and random-stream design. Identical transformed images may not be required if the claim concerns the distribution of augmentation, but the comparison must specify that boundary.
Diagnostic modeHeld fixedWhat it can establish
Frozen tensor fixtureInput bytes and parametersLocal forward/loss/update agreement
Transform fixtureRaw input, transform parametersPreprocessing semantics
Stochastic auditSampling rule and planned seedsBehavior of intended randomness
Full comparisonFinal scientific protocolObserved method result under that protocol
These are complementary checks. A successful local test can remove a specific hypothesis without proving the final score.

Exercise: separate regularization from reduction

The reference and candidate have identical per-example data losses [1, 3]. The reference reports total objective 2.5; the candidate reports 2.0. The intended objective is mean data loss plus penalty 0.5. What is the first explanation to test, and what would distinguish it from sum reduction?
The data mean is two and the sum is four. A missing 0.5 penalty is directly consistent with the observed difference, while plain sum reduction would give four before penalty. Inspect the scalar used for backward and whether the penalty depends on the intended parameters. Compare its gradient separately. Award one point for mean and sum, one for the penalty hypothesis, and two for a discriminating check and cautious attribution. A numerical match to one hypothesis is not proof until the code path is inspected.

Misconceptions to correct

“Same loss value means same optimization” fails when gradient connectivity, regularization, or optimizer state differs. “A diagnostic simplification is a reproduced method” fails when the simplification removes a required part of the final procedure. Restore and validate every removed component before making the broader claim.
When documentation is incomplete, write an ambiguity register: missing choice, plausible values, reason for each, small test, and resulting interpretation. If the exact preprocessing cannot be recovered, label the result as an implementation under stated assumptions. This is useful research engineering, but it is a narrower claim than reproducing the original procedure exactly.
End the first-hour report with the earliest confirmed mismatch and the next bounded test. An efficient report might say that sample identities match, inputs differ by a unit conversion, and optimizer hypotheses remain untested. That is enough to guide work without inventing a complete explanation for the final accuracy gap.

Interview probe

Original practice: You cannot reproduce a paper's score. What do you do in the first hour? A strong answer aligns data and checks a tiny forward-loss-gradient-update chain. Follow up with missing preprocessing details. A weak answer runs more seeds before checking that the two programs compute the same thing.

Sources

docsPyTorch: reproducibilitydocs.pytorch.orgdocsscikit-learn: common pitfalls and data leakagescikit-learn.orgdocsNeurIPS: paper checklistneurips.ccdocsPyTorch: cross-entropy loss and reductiondocs.pytorch.org

Checkpoint

Shared logits match, but scalar loss differs by exactly the two-example batch size. Which check should come first?

AChange initialization and repeat the full run.BInspect sum-versus-mean reduction, example weights and the scalar used for backward.CIncrease the dataset size to reduce the discrepancy.DReduce the learning rate until the printed losses match.
Sign up free to answer and see why

Checkpoint

Logits and printed scalar loss match, but one parameter gradient is zero only in candidate. Which check is targeted?

AInspect graph connectivity, detaching, the differentiated scalar and gradient reset/accumulation.BIncrease evaluation size before inspecting backward.CTreat matching scalar values as proof that derivatives must match.DChange the heldout split to see whether the gradient discrepancy disappears.
Sign up free to answer and see why

Checkpoint

Per-example losses are [1,3]; intended objective is their mean plus penalty 0.5. Correct total?

A4.5B2C2.5D4
Sign up free to answer and see why

Checkpoint

Disabling augmentation makes a fixture match. What has been established?

AThe full stochastic method is reproduced.BAugmentation can be permanently removed without changing the claim.CEvery long-run score will now match.DThe simplified path matches; the original augmentation path still needs validation.
Sign up free to answer and see why

Checkpoint

A paper omits a preprocessing choice. What should a reproduction record do?

AChoose the best-scoring interpretation and describe it as the original procedure.BRecord plausible assumptions, the tests used to choose among them, and limits on exact-reproduction claims.CUse the library default without recording its version or value.DAverage scores across incompatible preprocessing choices and report one exact reproduction.
Sign up free to answer and see why

Can you locate the first mismatch in a forward-loss-gradient-update chain and design a test that distinguishes causes? Rate confidence from 1 to 5 and identify the simplification you must restore.

Not yetGetting thereConfident

Wrap-up

  • Find the earliest disagreement under controlled inputs. Separate implementation diagnosis from the final research comparison.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.