Write a conclusion that separates internal validity, external validity, and practical effect.
A well-controlled experiment can still answer a narrow question. Internal validity concerns whether the comparison supports the claimed difference under the study conditions. External validity concerns whether that conclusion transfers to other populations, tasks, times, or environments. A single benchmark does not automatically represent the intended deployment world.
Describe the evaluation population and its exclusions. Were long inputs omitted? Were labels available only for easy cases? Were all tasks in one language or domain? These details define where the evidence applies. A limitation is useful when it names a specific untested condition and a plausible reason behavior could differ.
Practical importance is separate from statistical evidence. With a very large sample, a tiny effect may be precisely estimated but operationally irrelevant. With a small sample, a potentially useful effect may remain uncertain. Report the estimate, uncertainty, resource cost, and decision context. Avoid reducing the conclusion to whether a p-value crosses one threshold.
Consider harms and failure severity where the research could inform a deployed system. Aggregate performance can conceal concentrated errors. NIST's AI risk framework encourages attention to context and evaluation throughout a system's life. For this course, the practical lesson is to define the use context and inspect failure slices rather than treating a single average as a complete account.
Worked example
An invented method improves accuracy from 90.0% to 90.2% on one million synthetic tasks but doubles inference cost. On 500 human-authored tasks, the difference is uncertain. A careful conclusion reports a small measured synthetic-task gain and an unconfirmed transfer to human-authored tasks.
The result may still be useful for studying the mechanism or for a cost-insensitive application. It does not justify the broad claim that the method is better for all users. The next study should target the intended real task distribution and compare quality at a relevant resource budget. If failures concentrate in one language, that slice belongs in the report rather than being hidden by the million-task average.
Exercise and solution
A method wins on short English questions but has not been tested on long documents or other languages. Draft a defensible conclusion and one next test.
A strong conclusion states the gain under the short-English benchmark protocol and explicitly leaves long-document and multilingual transfer unresolved. The next test samples the intended deployment cases with a matching answer rubric and resource budget. Award one point each for population boundary, no universal claim, a relevant transfer test, and resource comparability. Do not replace an untested condition with a vague statement that "more research is needed"; name the condition and the decision it affects.
Lab artifact: write a transport boundary
A result transfers only when the target conditions are sufficiently related to the studied conditions for the intended claim. Use a boundary table to identify where evidence is missing.
Dimension
Studied condition
Intended use
Unresolved question
Source
Synthetic questions
Human-authored requests
Are task cues and ambiguity similar?
Length
Under 300 tokens
Up to 8,000 tokens
Does evidence remain accessible?
Language
English
Two additional languages
Do labels and failure modes transfer?
Cost
Doubled inference budget
Fixed service budget
Is the same gain feasible?
Labels
Automatic exact match
Human rubric
Does measured success mean the same thing?
The table does not prove failure in any target condition. It identifies assumptions that need evidence. A source change can alter both the input distribution and the labeling process. Testing more random seeds on the synthetic benchmark cannot answer whether human judges interpret success differently.
A second failure case: significance without useful magnitude
Suppose a method's estimated improvement is 0.05 percentage points, with an uncertainty interval from 0.03 to 0.07 under a valid planned analysis. The application has a predeclared minimum useful gain of 0.5 points because the method requires substantial extra cost. The result is precisely positive but far below that practical target.
Now consider another estimate of 0.6 points with interval [−0.2, 1.4]. Its point estimate exceeds the target, but the interval includes no improvement and gains below the target. The decision may require more evidence rather than treating the positive point estimate as established practical value. Neither example can be reduced to “significant is good, non-significant is useless.”
code
1Decision rule example:2primary gain must exceed the minimum useful margin under the planned analysis3service cost must remain within the stated budget4predefined failure slices must pass their separate constraints5otherwise: report the evidence and choose stop, redesign, or further study
This is an original decision-rule example, not a universal statistical prescription. An application can rationally use expected utility or another framework instead. What matters is that the rule and its tradeoffs are explicit before the result drives the choice.
Exercise: compare error severity
Method A makes twenty harmless formatting errors and one severe unsupported factual claim in a fixed task set. Method B makes ten formatting errors and three severe unsupported claims. A single unweighted error count favors B, but the intended use places high cost on unsupported factual claims. What should the study report?
Report the error taxonomy and counts, explain the predeclared weighting or guardrail, and avoid hiding severe failures inside the aggregate. If no weights were specified beforehand, present the disaggregated results and state the decision uncertainty instead of inventing favorable weights afterward. A follow-up evaluation can test the severe-error rate with enough relevant cases. Award one point for disaggregation, one for severity relevance, one for predeclared rules, and two for uncertainty and a targeted next study.
The task set's size and sampling matter. Three severe errors in a deliberately adversarial set do not directly estimate population frequency without the appropriate sampling design. The stress test may still be useful for finding failure modes. State whether the experiment measures frequency, capability under a challenge, or examples of possible failure.
Misconceptions to correct
“A large sample fixes external validity” fails when all observations come from the wrong population or label process. “A narrow interval makes an effect important” confuses precision and value. Cost, failure severity, and the target decision still matter.
A strong conclusion can have three sentences: the observed result under the tested protocol; the practical tradeoff under the stated decision rule; and the specific untested condition most likely to change the recommendation. Avoid a vague list of every possible limitation. Choose the boundary that affects the next decision and name the evidence needed to cross it.
NIST's risk framework supports evaluating systems in their use context. The original tables here turn that broad requirement into a research exercise; the framework does not validate these particular numerical thresholds or guarantee transfer from one benchmark to another.
Interview probe
Original practice: Your result is statistically significant. Should it ship? A strong answer asks about effect size, cost, validity, deployment population, and failure severity. Follow up with a tiny gain on a huge synthetic dataset. A weak answer treats significance as a deployment rule.
A gain is measured on one synthetic benchmark. Which conclusion follows directly?
AThe method improves all intended human tasks.BThe proposed mechanism is established.CThe method improved that measured benchmark under the stated protocol.DEvery deployment slice is safer.
A gain is 0.05 points with interval [0.03,0.07]; minimum useful gain is 0.5. What is the practical reading?
APrecisely positive but below the stated useful margin.BAutomatically valuable because zero is excluded.CEvidence that the true gain is exactly 0.05.DNo empirical result exists because the gain is small.
One million synthetic examples give a narrow uncertainty interval. No human-authored tasks were evaluated. Which unresolved condition is necessary before extending the result to those tasks?
AThe synthetic interval must exclude zero by a larger numerical distance.BThe synthetic evaluation must use more random seeds while retaining the same input and label process.CThe two model scores must be paired on every synthetic example, after which transfer follows.DThe inputs, label meanings and task conditions must be sufficiently comparable, supported by target-domain evidence.
An aggregate error count improves but severe factual failures increase. What report supports a decision?
AOnly the lower total error count.BDisaggregated error types and the predeclared severity rule or stated decision uncertainty.CWeights chosen after results to favor the candidate.DThe severe cases omitted as outliers without a rule.
An adversarial stress set finds three failures. What can it directly establish?
ATheir population frequency without any sampling assumptions.BThat every real user will encounter a failure.CPossible failures under the tested challenges, with frequency requiring a sampling design.DThat a larger random seed set eliminates the failures.
Can you state the tested population, practical margin and most consequential transfer gap? Rate confidence from 1 to 5 and specify the next evidence needed for the decision.
Not yetGetting thereConfident
Wrap-up
State the tested boundary and the practical tradeoff. A precise limited claim is stronger than an unsupported universal one.