Lesson 3 of 4 · 35 min

Keep the evaluation separate from the method's memory

Identify contamination paths and design a cleaner benchmark.

Evaluation data can enter a model or workflow through more paths than explicit training. It may appear in pretraining, retrieval indexes, prompt examples, tuning notes, or repeated error analysis. An evaluation can therefore reward memorization or adaptation to known cases while being described as generalization.
Map the data path. Identify where questions, answers, and near-duplicates can enter training, retrieval, development, and evaluation. Exact duplicate removal is a useful first check, but paraphrases and shared templates can preserve the same information. A benchmark with many variants of one underlying problem may have fewer independent units than its row count suggests.
A fresh benchmark is not automatically representative. Newly authored tasks can avoid known public contamination but introduce unrealistic wording or difficulty. Use a clear task specification, answer verification, and a sampling plan tied to the target use. Keep benchmark construction separate from tuning the candidate on its answers. If the research team studies failures, designate those cases as development data afterward.
Do not claim that contamination is absent merely because no exact match was found. State the checks performed and their limits. If training data is unknown, a complete guarantee may be impossible. You can still strengthen evidence through temporal separation, private or newly authored evaluation where appropriate, and tests that vary the underlying task rather than only surface wording.

Worked example

A teaching benchmark contains 200 questions. An exact-match audit finds 20 in the retrieval index, including their answers. Removing those leaves 180. A template audit then finds 60 questions derived from ten shared problem templates. The team groups by underlying template when constructing development and final splits.
The final report states both checks. It does not claim 180 fully independent, uncontaminated tasks. A cleaner confirmation set contains newly authored problems from held-out templates, with answers checked by a separate process. The method's retrieval index is frozen before that set is evaluated. These controls reduce specific contamination paths while leaving pretraining overlap uncertain.

Exercise and solution

A model's prompt contains three worked examples, and the test set includes paraphrases of those examples with only names changed. What should the study do?
Treat the paraphrases as related to development examples, remove or group them out of final evaluation, and construct genuinely different cases for the target skill. Award one point each for identifying semantic overlap, separating development and test, preserving a record of the change, and avoiding a guarantee based on exact string matching alone. Changing names does not create a new reasoning problem.

Lab artifact: map exposure by stage

Build a matrix for question content, answer content, and underlying templates. A “no” means the audit found no known exposure through that recorded path; it is not a universal guarantee about unknown pretraining.
Evaluation materialPrompt examplesRetrieval indexTuning notesKnown training corpus
Exact question Q17NoYesNoUnknown
Verified answer Q17NoYesNoUnknown
Template family T4YesNo known matchYesUnknown
New family T9No known matchNo known matchNo known matchUnknown
Q17 cannot support an answer-generalization claim when the deployed workflow can retrieve its answer from the evaluation index, unless that retrieval is explicitly the intended task. A benchmark for locating a known document is different from one claiming unseen problem solving. Exposure is not always a flaw; it becomes a flaw when it conflicts with the claimed evaluation boundary.
The NAACL contamination paper describes retrieval-based overlap investigation and probing approaches for models whose training data may be opaque. This supports treating contamination detection as a methodological problem rather than assuming string deduplication settles it. This course does not reproduce those probes or claim that any one test proves absence of exposure.

A second failure case: answer rubric enters development

A team keeps final questions private but shares the exact answer rubric with the prompt-tuning process. If the rubric contains task-specific answers or distinctive clues, the final evaluation may still leak. Separate general scoring criteria from answer-bearing material. General requirements such as “cite the supplied source” can be part of the task specification; specific hidden conclusions from the final cases cannot quietly become development examples.
code
1Frozen final packet:2question set hash: Q93answer/rubric hash: A94template-family assignment: T95retrieval snapshot hash: R4, frozen before final packet use6access record: who used answer-bearing material and for what purpose7post-evaluation rule: any case used to revise the method becomes development evidence
Hashes identify versions; they do not prove no person or system saw the answers. The access and use record supports that separate question. An evaluation team can still make mistakes or create unrealistic tasks, so private construction needs answer verification and representativeness checks as well as secrecy.

Exercise: design a counterexample that tests structure

A prompt example solves a scheduling puzzle by ordering three tasks with one dependency. The final set changes names and durations but preserves the same dependency pattern and solution steps. Design a stronger transfer set for the stated skill of reasoning over dependencies.
Include held-out graph structures such as branching dependencies, two prerequisites for one task, and cases where resource constraints interact with precedence, if those belong to the target task. Keep the instructions and answer verification consistent. Avoid merely making every final problem longer or harder without a sampling rationale. Group related structures during development/final separation and record the relationship. Award one point for identifying shared structure, one for meaningful variation, one for preserved task scope, and two for grouping and verified answers.

Misconceptions to correct

“New wording means a new independent task” fails when the solution structure and answer-bearing cues are unchanged. “Private means representative” fails when authors create convenient or artificial cases unlike intended use. Contamination controls and population validity solve different problems.
Report the effect of cleaning transparently. If removing exposed questions changes the score, provide the original and cleaned denominators, rules, and scope. Do not remove only items the candidate got wrong or keep only exposed items it got right. A rule based on exposure evidence should be applied consistently across compared methods.
The final conclusion should list the inspected paths, detection methods, and remaining unknowns. A strong statement is “No exact matches were found in the frozen retrieval snapshot; template overlap was grouped; pretraining exposure remains unknown.” That statement is narrower than “uncontaminated,” but it is testable and useful for deciding the next evaluation.

Interview probe

Original practice: How would you test whether a benchmark gain is memorization? A strong answer maps exposure paths, checks exact and semantic overlap, and evaluates held-out problem structures. Follow up with unknown pretraining data. A weak answer says a new filename makes the set fresh.

Sources

docsscikit-learn: common pitfalls and data leakagescikit-learn.orgdocsscikit-learn: cross-validationscikit-learn.orgdocsNeurIPS: paper checklistneurips.ccdocsDeng et al., NAACL 2024: investigating benchmark data contaminationaclanthology.org

Checkpoint

No exact duplicate is found in a known retrieval snapshot. What can be claimed?

ANo exposure through any possible path.BNo exact match was found there; semantic and unknown training paths remain.CEvery task is independent.DThe model cannot have memorized any answer.
Sign up free to answer and see why

Checkpoint

The final questions remain private, but an answer-bearing rubric with task-specific conclusions enters prompt tuning. What is the correct evidence boundary?

AThe final score remains independent if no exact question string entered development.BIt remains independent if the rubric was called documentation rather than labelled training data.CIt remains independent if the team edited prompts manually rather than using gradients.DThe method can adapt to final-case information through the rubric; the untouched-final claim is no longer justified.
Sign up free to answer and see why

Checkpoint

A benchmark is intended to test finding a known document. Retrieval includes that document. Is exposure automatically invalid?

ANo; validity depends on whether that access is part of the declared task.BYes, every retrieval task must exclude all answer sources.CYes, even if the task is explicitly lookup.DNo, therefore every unseen-reasoning claim is also valid.
Sign up free to answer and see why

Checkpoint

Changing puzzle names preserves the same dependency graph and solution steps. Which additional evaluation most directly tests structural transfer within the declared task scope?

AGenerate more entity names on the same graph and treat each as a new independent problem structure.BRewrite each original puzzle in several sentence styles while keeping the graph and solution unchanged.CUse held-out dependency structures within scope, verify their answers, and report performance separately from familiar-template cases.DPut variants of each original graph in both development and final sets, then stratify the split by answer length.
Sign up free to answer and see why

Checkpoint

An exposure audit removes questions. What reporting is needed?

AOnly the cleaned score if it improves.BOriginal and cleaned denominators, consistent removal rules, and remaining unknown paths.COnly removed cases the candidate answered incorrectly.DA guarantee of clean pretraining from the retrieval audit.
Sign up free to answer and see why

Can you map evaluation exposure through prompts, retrieval, rubrics and development use? Rate confidence from 1 to 5 and name one path your audit cannot rule out.

Not yetGetting thereConfident

Wrap-up

  • Trace how evaluation information can enter development. Report contamination checks and their remaining limits.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.