Write a manifest that identifies a result and its dependencies.
A result is produced by code, data, configuration, software, hardware, and a measurement protocol. Saving only weights preserves one output of that process. It does not explain how the output was produced or let another engineer verify a comparison. Start each experiment with a manifest that identifies the whole procedure.
Use immutable references where possible. A branch name can move, a dataset path can be overwritten, and a package constraint can resolve differently next week. Record the code revision, dataset version or content hash, split definition, preprocessing version, exact configuration, and environment versions. Include the command or entry point that produced the run and the metric implementation used to evaluate it.
Separate intended settings from observed execution. A configuration may request eight devices, but a failed allocation may run on four. Record the actual world size, device type, precision mode, and completed steps. Keep failure and resume history. If an experiment stops early, the result should say so rather than inheriting the planned token budget.
A manifest should be readable enough to support review. Large artifacts can live elsewhere with hashes and paths. Secrets do not belong in it. A useful record also includes the hypothesis and the comparison it is meant to support. Otherwise, an engineer may reproduce a number perfectly while misunderstanding the scientific claim.
Worked example
This invented manifest describes a small image-classification comparison.
The manifest makes the held-out camera explicit. It does not claim that one seed proves a robust augmentation gain. A baseline run needs the same split and protocol with the augmentation disabled. If the baseline instead uses another camera or fewer steps, the comparison changes more than one variable.
Exercise and solution
A colleague sends weights and says "This got 82%". List the missing fields needed before accepting the comparison.
A strong answer asks for the evaluation dataset and split, metric definition and denominator, code and preprocessing versions, training budget, configuration, seed, environment, and the baseline protocol. It asks whether 82% is one run, a selected best run, or an average. Award one point each for data identity, procedure identity, budget, selection history, and comparison validity. Do not infer missing settings from file names. Reproducibility starts by identifying the experiment, not by guessing the command that might recreate the score.
Lab artifact: compare procedures before comparing scores
A manifest becomes useful when it makes a difference visible. Inspect two supposedly matched runs:
Field
Baseline
Candidate
Interpretation
Data hash
D31
D31
Same recorded data
Split hash
S7
S7
Same membership
Augmentation
none
crop-v2
Intended treatment
Completed updates
2,000
2,000
Same update count
Global batch
32
64
Additional procedure change
Examples consumed
64,000
128,000
Different exposure budget
Metric code
M4
M4
Same recorded metric
Equal update counts do not mean equal data or compute exposure. The candidate's batch change doubled examples consumed in this simplified no-repeat accounting. If the scientific claim is the effect of crop augmentation under equal examples, this comparison does not isolate it. If the intended claim is the quality of two entire training recipes at a fixed wall-clock budget, a batch difference may be part of the treatment, but the manifest must state that different estimand. Reproducibility includes reproducing the claim's conditions, not only reproducing a score.
The examples-consumed field can itself be misleading when sampling with replacement or using repeated augmentations. Record whether it counts unique source items, transformed views, tokens, or forward-pass examples. A token budget also needs a rule for padding and masked tokens. Two teams can both say “one million tokens” while counting different work.
A second failure case: a hash with an incomplete boundary
A hash proves the identity of the bytes that were hashed. It does not establish that those bytes are the entire effective dataset. If a manifest hashes only a file listing paths, the underlying image files can change while the listing hash stays fixed. Likewise, hashing raw data without the label map or preprocessing can miss a semantic change.
Define the artifact boundary. A practical packet can include content hashes for data shards, the ordered shard manifest, label vocabulary, split membership, preprocessing configuration, and metric code. The procedure that makes the manifest should fail if a referenced artifact is absent rather than substituting a current default. This is an original engineering design, not a requirement imposed by the NeurIPS checklist; the checklist provides reporting context.
Exercise: repair a misleading comparison
A model scored eighty-two percent on yesterday's evaluation. Today the same weights score eighty-five percent. The evaluation path is unchanged, but the split membership file was overwritten to remove blurry images. The reported denominator falls from one thousand to eight hundred. Which artifacts and claims need correction?
Preserve both split versions and their denominators, name the filtering rule, and report that the population changed. Re-evaluate both candidate and baseline under the same fixed population if a comparative claim is needed. Do not interpret the three-point movement as a training improvement, because the weights did not change and the test population did. Award one point for population identity, one for denominator, one for common reevaluation, and two for preserving provenance and limiting the claim.
Misconceptions to correct
“Same weights means same experiment” fails when preprocessing, data, or metric semantics change. “A content hash makes a dataset valid” fails because identity is not correctness: mislabeled or leaked data can be perfectly immutable. Keep an identity audit separate from a validity audit.
The manifest should also state what cannot be reproduced. Proprietary data may be unavailable to an external reviewer; hardware may no longer exist; an external API may change. A bounded claim can still be useful: another authorized engineer can rerun the exact local artifact, while an outside reader can inspect a synthetic fixture and the method. Do not promise full external reproducibility merely because an internal file path is recorded.
Before leaving the lesson, write a two-sentence comparison claim that names the treatment, population, resource boundary, and result selection rule. This forces the manifest to serve the decision. If you cannot identify whether the result is a final checkpoint, best validation checkpoint, or best of multiple trials, the score is not yet ready for a scientific comparison.
Interview probe
Original practice: What belongs in an experiment manifest beyond hyperparameters? A strong answer includes immutable data and code references, actual execution, metric definition, and selection history. Follow up with a mutable dataset path. A weak answer says that setting a random seed is enough.
Which run identity permits a reviewer to recover the actual code and configuration?
AA branch name and current defaults.BAn immutable revision plus actual configuration and recorded local changes.CThe newest checkpoint's timestamp alone.DThe requested configuration without execution changes.
Both runs complete 2,000 updates; global batches are 32 and 64. Under full batches, what differs?
ANeither training exposure nor procedure.BOnly learning-rate scheduling; sample exposure stays equal.CExamples consumed are 64,000 versus 128,000.DThe larger batch necessarily has twice the accuracy.
Hashing a list of image paths guarantees which identity?
AThe bytes of that list, not necessarily the referenced image contents.BThe current contents of every image.CThe correctness of all labels.DThe validity of the train/test split.
The same weights score higher after difficult evaluation rows are removed. What should the report say?
AThe training method improved by the score difference.BThe population changed; a comparative claim needs a fixed shared evaluation.CThe old denominator can be retained because weights match.DThe filtering is irrelevant if both scores are percentages.
A run requested eight devices but used four. What should its packet preserve?
AOnly requested settings, because those define the intention.BOnly the final metric if training completed.COnly actual device count, deleting the failed request history.DBoth intended and actual execution, including the resource change.
Can you identify the effective experiment and distinguish identity from validity? Rate confidence from 1 to 5, then state a manifest field that would change your interpretation of a score.
Not yetGetting thereConfident
Wrap-up
Identify the full procedure behind a result. Record actual execution and the comparison the run is intended to support.