Define a small research proposal with alternatives and a stop rule.
A proposal interview tests whether you can turn an interesting question into useful experiments. Start with the problem and why the current evidence is insufficient. Then specify a hypothesis that can fail, a baseline, and the observations that distinguish success from alternative explanations.
Keep the first study small enough to execute and strong enough to answer a real question. A tiny experiment that cannot exhibit the proposed mechanism is not useful merely because it is cheap. A huge experiment with no discriminating control is not useful merely because it is impressive. Choose the scale where the necessary behavior can appear and the main confounds can be controlled.
Write a decision table before execution. If the method improves only the target slice, what does that suggest? If it improves every slice equally, could extra capacity explain the result? If it harms the target slice, what hypothesis is weakened? These predictions need not cover every possible outcome, but they should prevent the interpretation from changing opportunistically.
A proposal should include data provenance, evaluation isolation, resource accounting, and a plan for reporting failures. If the study needs human annotations, define the rubric and disagreement process. If it touches sensitive use cases, define the relevant safeguards and limits. Do not add broad claims about societal benefit without a path from the measured outcome to that benefit.
Worked example
A fictional proposal asks whether a retrieval filter reduces irrelevant-context errors. The baseline uses unfiltered retrieved passages. All methods start from the same retrieved pool and use the same answer-generation limit. The candidate filters passages; a random-removal control retains the same number and matched total token length as the candidate. The unfiltered baseline may use more input tokens within a declared maximum, so its resource difference is reported. Questions are grouped by underlying task family before development and test separation.
The predicted pattern is candidate improvement over both baseline and random removal on distractor-heavy tasks, without a major loss on tasks needing multiple sources. If candidate and random removal perform similarly, less context may explain the gain. If the candidate loses on multi-source tasks, the filter may remove necessary evidence. The study reports all three outcomes rather than only the favorable slice.
Exercise and solution
Propose a first study for "Longer context improves answers". You can vary context length and whether added material is relevant. Include a control and an informative failure.
A strong design crosses short versus long context with relevant versus irrelevant additions, keeps model and answer rubric fixed, and evaluates held-out question families. An informative failure is worse performance with longer irrelevant context, which limits the claim that length alone helps. Award one point each for factors, controlled budget interpretation, valid split, predicted outcomes, and a stop rule. If longer context increases compute, report that tradeoff instead of pretending length changed without cost.
Lab artifact: make competing predictions explicit
The retrieval-filter proposal needs separate comparisons for selection quality and context quantity. Candidate versus matched random removal addresses whether semantic selection helps beyond merely shortening the input. Candidate versus unfiltered context measures the combined effect of selection and the changed retained context, under the stated resource record.
Observed pattern
Interpretation to investigate
Next discriminating check
Candidate beats random and unfiltered on distractors
Selection may add value
Test independent task families
Candidate matches random; both beat unfiltered
Less context may explain gain
Vary retained length with matched random controls
Candidate loses only on multi-source cases
Necessary evidence may be removed
Audit passage retention against verified evidence
All methods improve after rubric change
Measurement changed
Re-score under the fixed original rubric
These interpretations are hypotheses, not automatic causal proofs. The methods must see comparable question sets, retrieval pools, and answer-generation conditions. A filtering method that secretly sees the final answer label has an invalid advantage even if its retained token count is matched.
A second failure case: a cheap experiment cannot exhibit the mechanism
A proposal claims to improve handling of conflicting evidence but screens only on tasks with one unambiguous fact. A null result on that screen does not refute the claimed conflict-resolution mechanism because the necessary condition is absent. Conversely, success may measure ordinary extraction rather than conflict handling. The first study should include verified conflict cases and a no-conflict control, while remaining small enough to inspect.
code
1Minimum informative study:2task families: conflict, distractor, multi-source, simple extraction3development: tune filter only on designated families/examples4final evaluation: held-out structures and verified evidence requirements5primary comparison: candidate versus length-matched random removal6secondary comparison: unfiltered baseline with actual input cost reported7stop: reference or answer-rubric failure blocks method ranking8report: all conditions, failed runs, excluded cases and reasons
The primary comparison addresses semantic selection under matched retained length. The secondary comparison addresses a different practical choice. Keeping both is useful as long as their interpretations are not merged into one mechanism claim.
Exercise: plan annotation without letting disagreement disappear
Two reviewers independently score forty proposed evaluation answers. They disagree on eight. The study author wants to keep only the thirty-two agreements to produce a clean benchmark. What should a better plan do?
First inspect whether disagreement reflects ambiguous questions, unclear rubric, or reviewer error. Use a documented adjudication process and preserve the original labels and reasons. If ambiguous cases are excluded, state the exclusion rule and how it changes the target population. Excluding difficult or ambiguous cases can make the final benchmark unlike intended use. Do not resolve disagreements by selecting the label that favors the candidate. Award one point for the ambiguity audit, one for independent/adjudicated records, one for consistent exclusions, and two for target-population and method-blinding concerns.
A revised rubric should be checked on development examples before final scoring. If final outcomes guide the rubric revision, acknowledge that the evaluation has become partly adaptive and seek a new confirmation set where needed.
Misconceptions to correct
“A small study is useful because it is cheap” fails when it cannot distinguish the hypothesis from alternatives. “A control is fair because its name says random” fails when it retains different token lengths, removes a different number of sources, or samples from a different pool. Inspect the actual intervention.
A strong proposal ends with branches rather than a promise of success. If the semantic filter matches random removal, investigate context quantity and stop claiming unique selection benefit. If it preserves needed sources and wins on independent structures, scale the study with a predeclared protocol. If answer verification fails, repair measurement before ranking methods. Each branch should retain enough budget and time to make its next result interpretable.
This makes failure informative. The project can discover that a simpler random reduction is sufficient, that a method harms multi-source tasks, or that the original benchmark never tested the target skill. Those outcomes can all improve the research decision without producing the hoped-for leaderboard gain.
Interview probe
Original practice: What is the cheapest experiment that could change your mind? A strong answer preserves the mechanism and controls while reducing unnecessary scale. Follow up with a result where random removal matches the new filter. A weak answer proposes a leaderboard run with no alternatives.
The semantic filter matches length-matched random removal and both beat unfiltered context. What should be investigated?
AReduced context quantity as a competing explanation.BThe semantic mechanism as already proven.CRandom removal as optimal for every task.DA guaranteed answer-rubric defect.
Semantic removal and random removal are compared to test selection beyond quantity. Which control set addresses the quantity alternative?
AMatch the initial pool and number of retained documents, without checking their token lengths.BMatch total retained tokens, but let each method draw from a different initial evidence pool.CMatch final answer accuracy and compare the retained contexts only among those equally successful cases.DMatch the initial pool, retained quantity/token length, and answer conditions relevant to the claim.
A conflict-resolution method is screened only on unambiguous single facts. What is the main limitation?
AThe cheap screen proves the mechanism fails.BThe necessary conflict condition is absent, so that mechanism is not tested.CSingle-fact extraction is impossible to evaluate.DA larger model automatically supplies conflicting evidence.
Reviewers disagree on 8 of 40 answers. What should the protocol do?
AKeep whichever label favors the candidate.BDelete all disagreement without reporting the scope change.CAudit ambiguity, adjudicate under a recorded rule, and preserve label/exclusion history.DTreat agreement on 32 as proof the other 8 are irrelevant.
Which proposal branch is justified if verified multi-source cases fail because necessary passages are removed?
AInspect retention against evidence requirements and revise the filter before claiming broad benefit.BReport only distractor tasks as the original universal claim.CIncrease answer length without examining retained sources.DAssume random seeds will restore missing evidence.
Can you write predictions under the preferred and competing explanations, with controls that actually match? Rate confidence from 1 to 5 and name an informative stopping result.
Not yetGetting thereConfident
Wrap-up
Design experiments whose outcomes change the next decision. Include a control that can defeat your preferred explanation.