Choose runs that distinguish competing explanations.
Research engineering often operates under a fixed compute or time budget. The best next run is not always the largest or most complete one. Choose experiments that distinguish plausible causes and preserve enough budget for confirmation. A run that could produce the same interpretation under every hypothesis has low diagnostic value.
Write the hypotheses and their predicted observations. Suppose a training failure may come from data corruption or unstable optimization. A tiny clean dataset test can isolate the learning loop. A data audit can inspect corrupted records without a full training run. A short run with unchanged data and one optimizer adjustment can test stability. These are different experiments with different costs.
Keep screening separate from confirmation. Small runs can reject obvious defects and rank ideas, but their results may not transfer to full scale. Once an idea survives screening, test it under the target protocol with matched controls. Record the selection process so the final comparison does not pretend that the candidate was chosen without looking at data.
A budget plan includes failed runs and overhead. Checkpointing, evaluation, data transfer, and debugging consume time. Reserve some capacity for reproducing the final result. If all compute goes into exploration, the project may end with an unconfirmed best run. State a stop rule for a hypothesis that repeatedly fails a necessary condition.
Worked example
A teaching team has 20 GPU-hours. A full run costs six. The candidate method is unstable, and two causes remain: incorrect masking or an aggressive learning rate. The plan spends one hour on deterministic mask-reference checks and one hour on short controlled learning-rate tests. It reserves 12 hours for one matched baseline/candidate full comparison and six for a confirmation run or failure recovery.
If the mask check fails, the team repairs it before spending the 12 hours. If both short tests pass but full-scale behavior differs, the small tests have not "proved" stability; they have removed specific simple causes. The run log retains all attempted settings and their costs.
Exercise and solution
You have eight hours, and each full run costs four. The new method's first batch produces NaN, while the baseline is finite. Should you launch two full runs?
No. First inspect finite inputs, loss terms, and gradients on a tiny batch, then test the first update. A full run that fails immediately cannot compare final quality. Award one point each for a necessary-condition test, local evidence, preserving confirmation budget, and an explicit stop condition. A valid plan may use much less than the full budget when the method fails a basic correctness check.
Lab artifact: make each run answer a question
An experiment plan should contain a branch after the result, not only a list of commands. Here is an original diagnostic plan for a training loss that fails only on padded sequences.
Hypothesis
Cheap discriminating test
Predicted observation
Next decision
Mask denominator wrong
Unequal valid-length reference fixture
Error scales with padding fraction
Repair objective before full run
Learning rate too high
Same correct fixture, small rate range
First update finite but later instability varies with rate
Bound a stability screen
Corrupt input row
Inspect failing batch and finite values
One record contains invalid values
Repair data contract
Capacity shortfall
Correct finite mechanics, controlled small task
Fails to fit clean task at tested capacity
Investigate representation
One test can support more than one hypothesis, so interpret it carefully. Lowering the learning rate may make a wrong denominator appear less unstable without repairing the intended objective. Conversely, a correct mask fixture does not prove that every production mask is valid. Record the input coverage of each check.
A second failure case: an informative screen that does not transfer
A candidate trains stably on sequence length sixty-four but diverges at length two thousand. The short screen ruled out certain local implementation defects; it did not establish stability at all shapes. Long sequences can change memory, normalization, numerical ranges, and effective token weighting. The confirmation plan needs representative target conditions.
A useful budget record includes both planned and actual expenditure:
code
1Total allocation: 20 GPU-hours2Local correctness and data audit: planned 2, actual 1.53Short stability screen: planned 2, actual 2.54Matched target comparison: planned 12, remaining reservation 125Confirmation/failure reserve: 46Admission rule: do not start work that consumes the reserved matched comparison7Stop rule: a failed necessary correctness condition blocks full-quality runs
This revised example allocates twenty hours as two plus two plus twelve plus four. The earlier example allocated one plus one plus twelve plus six. They are two possible plans, not costs to add together. Actual screening cost is four hours, so the reserve remains four. If the short screen overruns to six total, explicitly replan; do not pretend the original confirmation reserve still exists.
Exercise: choose between two equally priced tests
You have two hypotheses for a factor-of-two gradient mismatch: sum-versus-mean reduction or duplicated gradient accumulation. Test A runs another epoch on the same batch size. Test B measures gradients after one backward pass with cleared gradients, then after a second pass without clearing, for one- and two-example fixtures. Both fit within the available hour. Which is more informative?
Test B distinguishes how the gradient scales with example count and repeated backward calls. Test A may preserve the same mismatch without locating it. The expected mean-reduction gradient should average example contributions; repeated accumulation without clearing adds contributions again. Record both controls, then inspect the implementation once the pattern is known. Award two points for choosing the discriminating test, two for predicted observations, and one for confirming the code cause.
Misconceptions to correct
“The cheapest test is always best” ignores whether it can change the decision. A cheap unrelated benchmark may provide no evidence about the failure. “A full run is always stronger evidence” ignores prerequisites: an invalid objective can produce a precise final score for the wrong method. Evidence quality depends on the question and validity, not only compute spent.
Screening should not consume the final holdout. If candidate choices are made from a small development benchmark, keep that selection history and evaluate the selected procedure on the appropriate untouched data. A long-run confirmation can address target-scale behavior while an independent holdout addresses selection bias; these are different safeguards.
In an interview, state the maximum cost, the result that stops the branch, and the evidence retained on failure. This makes the plan operational. It also shows that “no full run yet” can be a correct engineering decision when the first batch already falsifies a necessary assumption.
Interview probe
Original practice: Which experiment would you run next with limited compute? A strong answer states competing hypotheses, predicted outcomes, cost, and what decision changes. Follow up with a small-scale result that may not transfer. A weak answer always selects the largest model.
A candidate fails the first-step finite-gradient check. What is the next useful test?
AMore full-scale seeds without changing the failed prerequisite.BA wider final-test search.CA local input-loss-gradient investigation.DA longer training schedule to average away NaNs.
Why can lowering learning rate mask but not repair a wrong loss denominator?
ASmaller steps may reduce instability while leaving the objective weighting wrong.BA stable training curve proves the denominator matches the scientific objective.CMultiplying learning rate always restores every per-example weight, even when valid lengths differ.DThe learning rate can replace a global valid-token count in distributed normalization.
A short-sequence screen passes but long sequences fail. What is the supported next interpretation?
APassing the small screen proves the full method, so exclude long failures.BThe failure must be due only to fewer seeds, without inspecting shapes.CAll input lengths share identical memory and numerical conditions.DThe screen removed some local hypotheses; target-shape behavior still needs a trace and controlled test.
To distinguish mean reduction from duplicated accumulation, which test is better?
AAnother full epoch at the same batch size.BOne- and two-example gradients with cleared state, then a repeated backward without clearing.COnly final accuracy after a new seed.DOnly a larger model under the same bug.
A 20-hour plan reserves 12 for matched comparison and 4 for confirmation. Screening consumes 6 rather than 4. What is required?
ATreat screening as sunk and keep a 4-hour reserve inside the same total.BCount only successful screening attempts against the budget.CRecord the now 2-hour remaining reserve and explicitly replan scope or resources.DShorten only the baseline run by 2 hours while keeping an equal-budget comparison claim.
Can you allocate finite compute to tests that change the next decision and retain confirmation capacity? Rate confidence from 1 to 5 and write the stop condition.
Not yetGetting thereConfident
Wrap-up
Choose runs that distinguish explanations. Reserve budget for a matched confirmation under the target protocol.