Distinguish absence of evidence, practical equivalence, and a failed mechanism.
A result that does not meet a significance threshold is not proof that two methods are equal. The study may be noisy, small, or poorly matched to the expected effect. Conversely, a well-designed result can rule out effects large enough to matter. Interpret the estimate and uncertainty against a stated practical threshold.
Before calling a result negative, check whether the experiment tested the intended method. An implementation defect, broken labels, or unequal budget can invalidate the comparison. Those are engineering failures to repair, not scientific evidence against the idea. Once the protocol is valid, report the result even if it is disappointing.
Practical equivalence needs a defined margin and an appropriate analysis. If differences within one point are operationally interchangeable, that margin should come from the use context rather than the observed interval. A wide interval spanning large gains and large losses cannot support equivalence. A narrow interval inside the predeclared margin may support a more useful conclusion than a binary non-significant label.
A mechanism can fail even when an aggregate score improves. Perhaps the proposed component helps through extra capacity rather than the claimed reasoning behavior. The next study should distinguish those explanations. This can redirect the research without pretending that the original story was confirmed.
Worked example
An invented valid comparison estimates a quality difference of plus 0.2 points with an interval from negative 2.0 to plus 2.4. The practical improvement threshold is one point. The result is inconclusive about a useful gain and also allows a meaningful loss. It does not establish equality.
A second study estimates plus 0.1 with an interval from negative 0.3 to plus 0.5. Under a predeclared equivalence margin of plus or minus one point and an appropriate equivalence procedure, this narrower evidence may support practical similarity. The interval alone should not be casually relabelled as a formal test result unless the analysis was designed for that purpose.
Exercise and solution
A method's average gain is zero, but it helps long inputs by four points and harms short inputs by four points. What should the report and next experiment do?
Report both preplanned or clearly exploratory slices with denominators and uncertainty. Investigate the interaction with input length using a controlled design. Do not conclude that the method has no effect merely because the aggregate cancels. Award one point each for recognizing heterogeneous effects, honest exploratory labelling, uncertainty, and a targeted next test. Avoid selecting only the favorable long-input slice and presenting it as the original universal claim.
Lab artifact: three different inferential questions
A difference test, an equivalence test, and a minimum-useful-gain decision can ask different questions about the same estimated effect.
Question
Target comparison
What a near-zero estimate alone cannot show
Any nonzero difference?
Difference versus zero
Whether the effect is useful
Practical equivalence?
Difference inside predeclared lower/upper margins
Whether uncertainty fits those margins
Useful superiority?
Difference above a useful-gain threshold
Whether the true gain exceeds the threshold
For paired mean differences, a two one-sided test approach can test a lower and an upper equivalence boundary. Both boundaries must be supported under the chosen significance level and assumptions. The statsmodels paired TOST documentation is direct authority for the paired mean-difference and margin contract. The illustrative intervals in this course are not outputs from an executed test.
Under the usual t-based TOST correspondence, two one-sided tests at alpha 0.05 correspond to checking a 90% two-sided confidence interval against the equivalence margins, with the relevant assumptions and strict containment rule. Do not automatically use an arbitrary reported interval or a one-sided noninferiority test and call it two-sided equivalence. State the interval level and procedure. More complex dependent data needs an analysis that respects its structure.
A second failure case: defining the margin after seeing the interval
A study's interval is [−0.8, 0.7]. The researcher proposes a margin of plus or minus 0.9 only because it contains the interval. That is a post-result choice, not a predeclared practical definition. The margin should come from the use context, measurement scale, and consequences of differences. A smaller margin of plus or minus 0.5 would not be established by the same interval.
code
1Negative-result record:2intended method verified: yes, reference checks passed3effect estimate and interval: recorded with unit and analysis method4practical margin: fixed from the application before result inspection5selection history: all planned runs and outcomes retained6conclusion: inconclusive / useful difference / practical similarity under stated test7next question: one specific validity, precision, mechanism or transfer gap
This record separates a valid inconclusive study from an invalid implementation. If a mask bug means the candidate never executed the intended objective, its score is not evidence against that intended method. Keep the failed attempt in the engineering history, repair it, and run the declared comparison. Do not hide its cost or treat it as a scientifically clean negative.
Exercise: choose the next study from heterogeneity
Suppose an aggregate gain is zero. Prespecified long-document tasks show plus four points and short tasks show minus four, with adequate counts for descriptive estimates but wide uncertainty. A researcher wants to deploy only on long documents. What evidence is needed next?
Report both slices, their counts, and uncertainty. Check that length was defined before seeing outcomes and is available at decision time. Test the length-by-method interaction under a controlled design and confirm the proposed routing policy on independent data. The routed system adds a policy choice; its overall utility depends on the actual length mixture and routing errors. Award one point for transparent slices, one for prespecified/available routing, one for interaction, and two for independent policy evaluation and mixture.
Misconceptions to correct
“Failure to reject zero proves the methods are equal” ignores the width of uncertainty. “A negative aggregate means nothing happened” ignores effects that cancel across groups. Both mistakes erase information that could guide a sharper experiment.
A useful final paragraph distinguishes what was ruled out from what remains plausible. If a valid narrow interval excludes a practically useful gain, the study can justify stopping that line under its conditions. If the interval is wide, the appropriate choice may be additional independent units, a more sensitive valid measurement, or a different question. More compute is not automatically the answer when uncertainty comes from too few independent datasets or a poorly defined outcome.
The aim is to make the next decision rational, including the decision to stop. An honest negative result can prevent a team from spending months on an unsupported mechanism, while a carefully qualified inconclusive result preserves a question that the study was not capable of resolving.
Interview probe
Original practice: Your hypothesis was not supported. What did you learn? A strong answer distinguishes validity checks, uncertainty, practical margins, and mechanism evidence. Follow up with offsetting subgroup effects. A weak answer either hides the result or declares the methods identical.
An interval spans -2 to +2.4 points with useful gain +1. What follows?
AThe methods are equivalent.BThe candidate is proved useful.CThe null is true.DUseful gain and meaningful harm remain compatible with the supplied uncertainty.
Why is choosing a margin just wide enough to contain the observed interval problematic?
AThe margin is selected from results rather than practical meaning fixed beforehand.BEquivalence margins must always equal zero.CEvery interval must cross a margin.DA larger sample makes post-result choice automatically valid.
A predeclared rule excludes runs with a verified implementation-contract failure from the valid-method score aggregate. A mask bug is confirmed in run R7: the executed loss omitted required tokens. R7 consumed six GPU-hours. What record and conclusion follow?
ARemove R7 and its cost because it did not execute the scientific method.BCount R7's score as a valid negative result for the intended method, but annotate the mask bug.CRetain R7, the verified cause, six GPU-hours, the applicable exclusion rule and a linked corrected rerun; distinguish the defective implementation from evidence about the intended method.DKeep R7 in the valid-method mean if its score is close to the other runs, but exclude it if it is an outlier.
Long inputs improve and short inputs worsen. A new long-only routing policy should be assessed how?
AReport only the favorable slice as the original universal result.BAssume routing has no errors or mixture effects.CUse the same selected slice as independent confirmation.DTest the length interaction and confirm the routed policy with its actual mixture and boundary.
Can you separate invalid implementation, uncertain difference and practical equivalence? Rate confidence from 1 to 5 and state the margin, procedure and next decision.
Not yetGetting thereConfident
Wrap-up
Interpret negative results against validity and practical margins. Use them to refine a specific next question.