Define protected correctness conditions for a performance task.
A performance benchmark is a contract about work and correctness. An optimization is valid only if it performs the same required computation under the same allowed resources. Changing tests, reducing input difficulty, or increasing a prohibited resource can make a number look better while invalidating the result.
Read the benchmark contract before editing the implementation. Identify inputs, outputs, allowed transformations, resource limits, and the authoritative correctness tests. Keep a baseline result and a clean copy of the test definitions. Measure speed only after the candidate passes the required checks. A fast wrong answer is not a partial performance improvement.
This lesson has real first-party hiring evidence. Anthropic published a version of its historical performance take-home, based on optimization for a simulated accelerator. The repository explicitly states that newer time-limited assessments use a different base. It is useful practice evidence, not a claim about a current interview. The repository also warns that changing protected tests or intentionally restricted machine settings invalidates apparent improvements.
Our worked case below is original and smaller. It teaches the same general discipline without reproducing the employer task or claiming to solve it. The full public task remains an optional external exercise under its own instructions.
Worked example
A teaching benchmark sums 1,024 integers and checks the exact integer result. The baseline processes one value per loop iteration. Candidate A unrolls the loop and reduces overhead while preserving all inputs. Candidate B processes only the first 512 values and changes the test to compare that partial sum.
A records a valid speed change if it passes the original tests. B records no valid gain because it changed the problem. Candidate C uses eight workers when the task requires one. Even if its answer is correct, it violates the resource condition. Correctness and allowed resources are separate checks.
A useful result table includes baseline time, candidate time, input size, correctness status, allowed-resource status, and environment. If any condition changes, record a new experiment instead of presenting it as a direct speedup.
Exercise and solution
An optimized kernel is twice as fast on one input shape but fails on odd lengths. The benchmark includes odd lengths. The engineer proposes excluding them as "rare". What is the valid report?
The candidate is incomplete for the benchmark contract. Report the speed result only as a limited diagnostic on the passing subset, fix the odd-length path, and rerun the full protected suite. Award one point each for respecting the input domain, separating limited evidence, preserving tests, and requiring full correctness before a general claim. The frequency of odd lengths does not authorize removing them from a stated requirement.
Lab artifact: write the benchmark's protected boundary
An optimization task often permits changes in internal order while protecting outputs and resource limits. Make those boundaries explicit before comparing times.
code
1Original teaching benchmark:2input domain: integer arrays, lengths 0 through 1,0253output: exact mathematical sum within the stated no-overflow domain4allowed: loop transformations and internal buffer layout5fixed: one worker, same input set, same timing boundary6protected: reference outputs and domain checks7measurement: repeated completed runs after the stated warm-up8report: passing domain, latency distribution, environment
The empty input and odd maximum length are deliberate. They test loop boundaries that a vector-width-aligned case can miss. If the implementation uses fixed-width integers, define the allowed values so overflow is impossible or define the overflow semantics explicitly. An optimization that changes accumulation width can change answers even when ordinary small fixtures pass.
The historical Anthropic repository is a different task with its own simulator and constraints. Its public README is first-party evidence that the historical task was used, not a license to infer current question wording, duration, thresholds, or scoring. The employer's engineering article provides context for why preserving test integrity matters. Learners should read the linked repository's actual instructions before attempting it.
A second failure case: a timing change disguises a speedup
Candidate D caches the correct output for the fixed benchmark inputs and measures only cache lookup, while the baseline includes computation. If the task requires computing arbitrary allowed inputs, this changes the effective work. If caching is explicitly allowed and the production claim concerns repeated inputs, it may be a valid separate policy, but cache warm-up, memory, invalidation, and hit-rate assumptions become part of the benchmark.
Candidate E starts its timer after input transfer while the baseline starts before transfer. Both compute correct results, but the reported speed ratio mixes boundaries. Correctness passing is necessary, not sufficient, for a valid performance comparison.
Candidate
Correct output
Allowed resources
Same work/timing
Claim
Loop unroll with tail handling
Pass
Pass
Pass
Candidate speed measurement valid
Cached known outputs only
Pass on fixed set
Depends
Fails arbitrary-input task
No general speedup established
Excludes transfer only in candidate
Pass
Pass
Fails
Timing comparison invalid
Exercise: detect a changed numeric contract
A reference sums signed integers in a wide accumulator. A candidate uses a narrow accumulator and passes all small-value tests but overflows on a valid large-value input. It is thirty percent faster on the small cases. What should the report and next change be?
Report a limited small-case timing with an explicit correctness failure in the full domain, not a valid full-benchmark gain. Preserve the failing input as a regression case and fix accumulation semantics or use a valid algorithm that respects the original domain. The benchmark owner could define a narrower task, but that would be a new contract and cannot be imposed silently by the optimizer. Award one point for each scope distinction, one for the regression fixture, and two for preserving the original requirements.
Misconceptions to correct
“If outputs pass the visible examples, every allowed input is handled” confuses finite tests with a full domain argument. Use boundary cases and reasoning about indexing and numeric range. “Any change that gives the same answers is a valid optimization” ignores resource and timing conditions. A larger worker pool or precomputed cache can change the problem even while returning correct values.
For an AI-assisted code change, inspect modifications to the evaluation harness separately from implementation changes. A proposed test adjustment may be a legitimate bug fix, but it requires its own rationale and should not be counted as an implementation speed improvement. Keep an untouched baseline and record the exact candidate diff so another reviewer can check whether the work remained the same.
Interview probe
Original practice: How do you ensure an AI-assisted optimization did not cheat? A strong answer preserves tests, checks allowed resources, compares against a reference, and reviews changes affecting workload. Follow up with an altered core count. A weak answer trusts the improved timing output alone.
A candidate changes protected expected outputs to match its faster result. What follows?
AThe comparison no longer measures the original task.BA larger speedup compensates for correctness changes.COnly the baseline needs rerunning under changed tests to preserve the original claim.DOutput changes are valid whenever tensor shape is unchanged.
A sum kernel fails on odd lengths included in the domain. Which report is valid?
AA general speedup because most inputs are even.BA full pass after excluding odd lengths from the protected suite.CA limited passing-subset timing with explicit full-domain failure.DA successful optimization if the average error is small.
Candidate timing excludes transfer; baseline includes it. What is wrong?
AThe candidate must be numerically incorrect.BThe baseline should be compared only on its slowest run.CTransfer is irrelevant to every performance claim.DThe ratio compares different work boundaries.
An optimization task fixes one worker and exact output checks. Candidate preserves outputs but uses eight workers. Which comparison is valid?
AReport eight-worker speed as the one-worker optimization result.BTreat output agreement as proof every task constraint passed.CEvaluate within the one-worker limit, or clearly report the eight-worker run as a separate out-of-contract measurement.DDivide the eight-worker time by eight to estimate the permitted runtime.
A narrow accumulator overflows on a valid large input but passes small examples. What must change?
AThe numeric implementation must meet the protected domain before a full valid speedup claim.BDelete the large case because small examples pass.CReport only speed and omit the overflow.DAssume all integer accumulation widths are equivalent.
Can you defend a benchmark's output, input-domain, resource and timing contract? Rate confidence from 1 to 5 and name a change that would require a new benchmark claim.
Not yetGetting thereConfident
Wrap-up
Protect the benchmark contract. Use the historical employer task as attributed practice, not as a current hiring promise.