Lesson 3 of 4 · 35 min

A seed is one control, not a reproducibility guarantee

Design separate tests for repeatability and seed robustness.

Randomness enters experiments through initialization, data shuffling, augmentation, dropout, sampling, and library kernels. Setting one seed may control only part of that system. Even with all intended random generators controlled, a change in library version, hardware, or algorithm can change numerical results.
Distinguish repeatability from robustness. Repeatability asks whether the same procedure under the same stated environment produces the same result or an acceptable tolerance. Robustness asks whether the conclusion holds across reasonable changes such as initialization seeds. A single perfectly repeatable lucky run does not establish a reliable method advantage.
Control the random sources that matter. A data-loader worker may use its own generator. A NumPy generator object has state separate from a global seed. Distributed ranks need a deliberate seeding scheme so they do not accidentally use identical augmentation streams or overlapping samples. Record the scheme and test the resulting data identities.
Deterministic algorithms may have performance costs or lack support for some operations. PyTorch documents that complete reproducibility is not guaranteed across releases and platforms. State the environment boundary of a repeatability claim. If exact determinism is unavailable, use tolerances and repeated-run distributions suited to the question rather than hiding variation.

Worked example

An invented comparison evaluates baseline A and method B on the same five seeds. A scores 70, 72, 71, 73, and 69. B scores 71, 73, 72, 74, and 70. The paired differences are all plus one point. This small result is more informative about consistency than comparing A's mean with B's single best run of 74.
It still does not prove generalization across datasets, hardware, or all seeds. The record should include all runs, the pairing rule, and the planned seed set. If the team tried 50 seeds and reported only these five, the interpretation changes. Selection history belongs with the result.

Exercise and solution

A colleague reports that seed 42 gives a three-point gain, while seeds 1 and 2 give a two-point loss. They propose publishing only seed 42 because it is standard. What should the report do?
Report all planned runs and summarize the variability and paired differences. Investigate whether the method is unstable and whether the protocol has a defect. A standard seed has no privilege to represent the population of training outcomes. Award one point each for complete reporting, robustness distinction, pairing, and a justified follow-up. Do not average away a known implementation bug; first determine whether the runs measure the intended method.

Lab artifact: two tests with different questions

A repeatability test holds the intended configuration fixed. A robustness study changes a prespecified source of variation. Put both in the results packet rather than using one as a substitute for the other.
TestSeed planEnvironmentQuestion
R1Seed 17 repeated three timesFixed E9Does this execution repeat within tolerance?
R2Seeds 11, 17, 23, 31, 47Fixed E9Is the method difference stable across planned initializations?
R3Seed 17E9 and E10What changes across environments?
A different output in R3 does not automatically identify a defective random generator. Kernel selection, reduction order, compiler changes, and hardware precision can affect the trajectory. The intended tolerance and the scientific consequence matter. A tiny parameter difference may leave predictions unchanged, while an unstable training process can amplify it. Report what was compared: parameter bits, loss values, predicted labels, or a downstream metric.

A second failure case: pairing names without pairing conditions

Seed seventeen in two methods is not necessarily a meaningful pair if one method consumes random draws in a different order. Adding an augmentation can shift the random stream before initialization unless generators are separated. Shared seeds can still serve as a planned blocking scheme, but the meaning of the pairing should be stated. Do not assert identical initialization merely because the integer seed matches.
One design assigns separate random streams for parameter initialization, sampling, and augmentation. It can save the initial parameter state when identical initialization is part of the comparison. The manifest records those stream identities and the data order. This is an original experimental design choice, not a universal best practice: some studies intentionally compare full stochastic procedures without identical draw histories.
code
1seed_plan: [11, 17, 23, 31, 47]2pairing_unit: saved initial-parameter state plus fixed data order3augmentation_stream: separate per method and seed4reporting: all planned runs, including failures5selection: no replacement of losing seeds6repeatability_tolerance: stated per tensor/metric and environment

Exercise: calculate the complete result

Baseline scores across three planned seeds are [80, 82, 81]. Candidate scores are [83, 80, 80]. The paired differences are [3, -2, -1], whose mean is zero. Candidate's best score exceeds baseline's best, but the planned mean difference does not show an average gain. The observed spread warns that a single-run story is unstable.
What should the next experiment be? First inspect whether all runs used the intended method, data, and budget. If no defect is found, a larger prespecified set can estimate variability more precisely, subject to resource limits. Do not continue adding seeds only until the average turns positive. A fixed expansion plan or a valid sequential design is needed if inferential claims are made. Award one point for differences, one for mean, one for rejecting best-run comparison, and two for the bounded follow-up.
A failed run belongs in the record. If candidate seed twenty-three diverges because the method is numerically unstable under the planned configuration, silently replacing it with another seed hides part of the outcome. If it fails because of an unrelated infrastructure outage, document the reason and a prespecified retry rule. Distinguishing these cases requires logs, not retrospective preference for favorable scores.

Misconceptions to correct

“Same seed means same random draws everywhere” fails when generators, workers, or operation order differ. “Multiple seeds fix dataset selection bias” fails because repeating a biased split only measures variation conditional on that split. Initialization robustness, split robustness, and population transfer are separate questions.
For a research-engineering interview, explain the evidence boundary in one sentence: “Within environment E9 and these five planned paired initializations, the observed differences were this vector.” Then state what additional variation was not tested. This is more precise than calling the method universally stable or using a deterministic configuration as a substitute for empirical robustness.

Interview probe

Original practice: Does deterministic training remove the need for multiple seeds? A strong answer separates identical-run repeatability from robustness to initialization. Follow up with a hardware change. A weak answer treats a fixed seed as proof of a general scientific claim.

Sources

docsPyTorch: reproducibilitydocs.pytorch.orgdocsNeurIPS: paper checklistneurips.cc

Checkpoint

Which study examines seed robustness within a fixed protocol?

ARepeat one selected seed until results are identical.BReport all planned paired seed results and their variation.CSelect the best candidate seed and average baseline seeds.DUse the most common tutorial seed.
Sign up free to answer and see why

Checkpoint

Why can the same integer seed fail to imply identical initialization across two code paths?

AAny fixed seed prevents meaningful comparison.BOne path may consume random draws before parameter initialization.CA seed guarantees identical initialization even when generators differ.DDeterministic code never uses a generator.
Sign up free to answer and see why

Checkpoint

Differences across three planned seed pairs are [3,-2,-1]. Their mean is?

A3B-1C0D1
Sign up free to answer and see why

Checkpoint

A planned candidate run diverges because of its numerical method. What belongs in reporting?

AThe failure and its evidence, under the planned outcome definition.BA replacement favorable seed with no mention of the failure.COnly successful runs because failures have no scientific meaning.DA claim of hardware outage without supporting logs.
Sign up free to answer and see why

Checkpoint

Repeating the same biased split across many seeds addresses what?

APopulation-selection bias automatically.BOnly variation conditional on that split, not its population validity.CAll possible datasets because seeds are different.DThe exact deployment prevalence regardless of labels.
Sign up free to answer and see why

Can you distinguish repeatability, seed robustness and population transfer using the supplied tables? Rate confidence from 1 to 5 and identify a selection rule that would bias the report.

Not yetGetting thereConfident

Wrap-up

  • Control randomness and report variation. A reproducible run and a robust conclusion are different achievements.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.