Critique a result without substituting your own story
Build an evidence audit with claims, support, and missing tests.
A research interview may ask you to critique a paper or result. A useful critique identifies which claim lacks which evidence. It does not list every possible limitation or dismiss work because it is imperfect. Start with the central claim, then inspect whether the experiment actually estimates it.
Use a claim-evidence table. Separate the observation, the authors' interpretation, and your alternative explanations. For each alternative, name a test that would distinguish it. This turns criticism into a research plan. If no possible result would change your view, you may be defending a preference rather than analyzing evidence.
Check the baseline, data split, resource budget, selection process, uncertainty, and failure slices. These are not ceremonial boxes. Each can change the conclusion. For example, a baseline trained with fewer tokens cannot isolate an architecture effect. A test set used for tuning cannot independently estimate the selected model's performance. An average without seed variation may conceal unstable training.
Acknowledge what the evidence does support. A method can demonstrate an interesting benchmark gain while leaving its explanation unresolved. Recognizing that distinction makes the critique more precise. Avoid turning missing evidence into proof of wrongdoing. The proper claim is that the conclusion is not yet established, unless there is direct evidence of an error.
Worked example
An invented report says a new memory module improves reasoning because it raises accuracy from 70% to 75%. The candidate uses twice the parameters and context tokens of the baseline. Only the best of ten seeds is shown.
Claim
Available support
Missing discrimination
This selected run scored 75%
Reported benchmark result
Reproducible predictions and protocol
Memory mechanism caused the gain
Aggregate difference
Capacity/context-matched controls
Gain is robust
Best seed only
All planned seed results
Method transfers
One benchmark
Relevant held-out task families
The next study should first verify the reported comparison, then use matched controls and complete seed reporting. It need not begin with a brand-new architecture. The critique has identified a smaller path to a stronger conclusion.
Exercise and solution
A study claims lower inference cost but reports only batch throughput at batch size 128. Deployment requires single-request latency under 100 ms. Write one supported claim, one unsupported claim, and one useful next measurement.
The supported claim concerns throughput under the measured batch setup. The deployment latency claim is untested. Measure latency and cost on the target request distribution and hardware under the required service constraint. Award one point each for preserving the valid result, naming the scope mismatch, choosing a relevant measurement, and avoiding a blanket dismissal of the study.
Lab artifact: prioritize the flaw that can change the conclusion
Imagine three concerns about the memory study: the plot font is hard to read, the candidate has twice the context budget, and the code uses an older package version. Only one directly threatens the claim that the memory mechanism caused the observed improvement without further evidence: the changed context budget. The package version might matter if it changes behavior, but age alone is not a defect.
code
1Central claim: memory selection, not extra context, causes the gain2Observation: selected candidate run scores 75%, baseline 70%3Competing explanation: candidate sees twice as many evidence tokens4Discriminating test: match available context and selection effort5Result that weakens concern: gain persists under valid matched controls6Result that strengthens concern: gain disappears when token budget matches7Remaining issue: uncertainty and transfer still need separate evidence
A prioritized critique is conditional. It says what evidence would change your mind. If the matched comparison preserves the gain, the context explanation becomes less compelling, though it does not establish the entire mechanism. A good critique can be answered; it does not move to a new objection every time the previous one is resolved.
A second failure case: the reported metric is correct but answers another question
A system reports average latency of thirty milliseconds while the application requires p95 latency below one hundred. In a toy set of one hundred requests, ninety-five take ten milliseconds and five take four hundred ten. Mean latency is thirty milliseconds: (950 + 2,050) / 100. Under a nearest-rank p95 convention, the 95th observation is ten milliseconds, while p99 is four hundred ten. This example also shows why quantile definition and exact service requirement matter.
If instead six requests take four hundred ten and ninety-four take ten, the mean is thirty-four milliseconds and nearest-rank p95 is four hundred ten. A low average can coexist with a failed tail requirement. Do not infer a tail quantile from the mean. Use the declared quantile convention and enough representative requests to estimate the tail with uncertainty.
Evidence
Directly supports
Does not directly support
Batch-128 throughput
Work completed per time in that setup
Single-request tail latency
Mean latency
Average under measured workload
Every-request deadline
Nearest-rank p95
That sample quantile
An exact population guarantee
One selected seed
That run's observed result
Robustness across runs
Exercise: turn a vague objection into a study
The critique is “The benchmark is too easy.” Rewrite it so a researcher can respond.
A stronger critique names the target claim and a missing difficulty condition. For example: the claim concerns multi-source synthesis, but all evaluation questions can be answered from one passage. Add verified held-out questions that require combining two independent sources, with a control that removes one necessary source, and compare under the same answer rubric and resource contract. If the method succeeds only when one passage contains the full answer, the broad synthesis claim remains unsupported. Award one point for specificity, one for a relevant test, one for a discriminating control, and two for preserving the valid existing result and limiting the broader claim.
Misconceptions to correct
“Missing evidence proves misconduct” substitutes an accusation for an evidence gap. “Every limitation deserves equal attention” makes the critique unusable. Prioritize the issue most likely to change the central conclusion and the cheapest valid test that addresses it.
The critique should also identify a strength. This is not politeness padding: it states which observation remains available for future reasoning. If the reported speed measurement is reproducible and properly bounded, preserve it even when the deployment claim fails. A research program can build on a valid narrow result while repairing an unsupported interpretation.
Interview probe
Original practice: What is the strongest weakness in this result? A strong answer selects the limitation most likely to change the central claim and proposes a discriminating test. Follow up with evidence that would remove the concern. A weak answer recites generic complaints without prioritization.
An observed gain has an untested context-budget confound. What is the precise critique?
AThe gain cannot have occurred.BEvery reported metric is unusable.CThe mechanism attribution needs a matched control.DThe proposed mechanism is disproved by missing evidence alone.
Which evidence would directly weaken the concern that extra context, rather than the proposed mechanism, explains a method's gain?
AA gain that persists in valid context-matched comparisons with the relevant resource and selection controls.BMore independent seeds for the same unequal-context recipes.CA higher score after increasing the candidate's context further while leaving the baseline unchanged.DA post hoc regression on context length when the observed data contains no matched or overlapping context conditions.
A multi-source claim is tested only on one-passage-answerable questions. Which next test is relevant?
AA new plot of the same single-source scores.BVerified held-out multi-source tasks with a necessary-source removal control.CMore replicas of the same one-passage questions only.DA different model size without changing task structure.
How should a critique treat a valid narrow speed result when a deployment claim lacks evidence?
AErase the speed result.BTreat speed as proof of deployment suitability.CPreserve the measured result and specify the missing target-workload measurement.DAssume deployment failure is certain.
Can you identify the central claim's most consequential evidence gap and a test that could remove your concern? Rate confidence from 1 to 5 and preserve one valid narrow observation.
Not yetGetting thereConfident
Wrap-up
Tie each criticism to a claim and a missing test. Preserve what the study does establish.