Recommend release, rollback, or further testing from a mixed evaluation.
A release interview may present one improved average and several troubling examples. Read the denominator before celebrating the average. A model revision can improve common tasks while making a rare but consequential action unsafe. Define release criteria before seeing the candidate results so the decision does not change to favor a preferred model.
Use a task set with meaningful slices. For an agent, useful slices include ambiguous identity, missing permission, transient failure, stale state, long context, and budget exhaustion. Record both task outcome and invariant violations. A task-level pass rate should not conceal severity. An unauthorized action may be a blocking issue even if it occurs once in a large set.
Compare candidates on the same tasks and preserve their traces. When outcomes are stochastic, repeat selected cases and report the number of runs. A single success after several failures does not establish reliability. Separate changes in the model from changes in tools, prompts, and environment. If all change together, the result describes the combined system but does not identify the cause of improvement.
Your recommendation should include what is known, what remains uncertain, and the next smallest informative test. A rollback can be justified by an invariant breach even before you prove the root cause. A limited release may be appropriate for a read-only path while write paths remain disabled. Do not imply that a small test proves safety for every future input.
Worked example
A fictional evaluation has 200 tasks. Version A passes 170 and version B passes 184. On 30 timeout-recovery cases, A passes 25 and B passes 20. B also creates two duplicate external writes; A creates none in this set. The overall rate improves from 85% to 92%, but the write invariant regresses.
The decision is to block B on write-enabled traffic, preserve the duplicate traces, and investigate recovery logic. A read-only experiment could be considered separately if its scope and evidence support it. The team should not call B generally safer or more reliable based on the average. The two duplicate cases are evidence of a known failure mechanism, not merely statistical noise to average away.
Exercise and solution
A new agent resolves 96 of 100 tasks versus 91 for the old agent. All four new failures are unauthorized cross-tenant reads; the old failures are harmless incomplete summaries. Make a recommendation and name the missing evidence.
A strong answer blocks the new agent, investigates the authorization boundary, and adds tenant-isolation regression cases. It notes that 100 tasks are too few to estimate a rare-event rate precisely, but the observed breach is already actionable. Award one point each for severity-aware decision, containment, causal investigation, and uncertainty about frequency. Do not award credit for claiming a 96% score makes the release acceptable.
Read the release table before the headline
This expanded invented evaluation separates ordinary completion from consequential failure.
Slice
Tasks
A passes
B passes
B unauthorized effects
Routine reads
120
105
116
0
Ambiguous identity
30
25
28
0
Timeout recovery
30
25
20
2
Budget/permission limits
20
15
20
0
Total
200
170
184
2
The table is consistent with the original overall rates. It shows that B's gain comes mainly from routine reads and limit handling while recovery regresses. The two unauthorized effects are not an extra denominator; they are a separate severity measure associated with tasks in the evaluation. Define whether a task with an unauthorized effect is automatically a failed task, as it is in this teaching contract.
A release memo should make the requirement explicit rather than rely on an emotional description of two bad examples. If the requirement forbids duplicate writes, the observed duplicates block the write path. A team could separately evaluate read-only use because its effect boundary differs. That is a new scoped decision, not a way to relabel the failing write release as successful.
A second worked case: zero observed failures
Suppose a patched candidate has zero unauthorized writes in 100 test tasks. The result is encouraging, but it does not prove the true failure probability is zero. Under a simple independent Bernoulli model, the probability of seeing zero failures at a true rate of 3% is 0.97 raised to 100, about 4.76%.
This calculation illustrates why a small clean sample cannot establish impossibility. It assumes independent, representative trials and a stable system. Those assumptions can fail for agent tasks with shared templates or correlated outages. Report the actual test conditions and use targeted regressions for known mechanisms alongside broader sampled evaluation.
code
1release record:2 known duplicate-write regression: fixed and reproduced as passing3 broader test: 0 observed unauthorized writes in 100 tasks4 inference limit: does not prove zero population risk5 scope: read/write tool version T8, model M3, policy P46 remaining review: correlated outage and key-expiry cases7 decision: do not exceed the scope supported by these checks
The exact deployment decision depends on the agreed risk requirement and available evidence. This workbook does not invent a universal sample size that makes every agent safe. It teaches how to avoid converting a finite test into a guarantee.
Keep the comparison attributable
If a candidate changes the model, tool schema, retry controller, and provider environment simultaneously, the evaluation measures their combined effect. That may be enough for a release decision, but it is weak evidence for which component caused the improvement. Preserve versions and use controlled replays to narrow the cause of a regression.
Also distinguish exploratory failure analysis from final evaluation. Cases used to tune the patch remain valuable regressions, but they are no longer untouched tests of generalization. Keep a fresh set for broader estimation and state how it was selected. A known failure test passing is necessary evidence of the repair, not sufficient evidence of universal reliability.
Misconceptions to reject
"Zero failures in a test means the failure cannot occur" confuses an observed count with a population guarantee. The sample and dependence assumptions limit the inference.
"An average improvement permits any rare regression" ignores requirements and severity. Some failures violate a hard action boundary even when the overall score rises.
Transfer exercise
A patch passes all three known duplicate-write cases and 50 new routine-read tasks. The team wants to re-enable every write tool. What evidence is missing?
The new tasks do not test write recovery, key expiry, concurrent workers, or relevant provider errors. The repair has evidence for the known cases, but broad write-path coverage remains limited. Propose fresh write scenarios matched to the intended scope and preserve the successful regressions. Score one point for scope mismatch, one for new recovery cases, one for preserving known tests, and one for a conclusion that neither discards the repair evidence nor overstates it.
The Bernoulli zero-failure calculation is original teaching arithmetic under the stated independent and representative trial model. No cited provider source is claimed to report that rate or establish the independence assumption.
Interview probe
Original practice: How do you explain a release block when average success improved? A strong answer names the violated requirement, affected scope, evidence, and next test. Follow up with a read-only subset that did improve. A weak answer hides the regression in the aggregate.
Overall pass rate rises while duplicate writes appear. What should a release decision prioritize?
AThe average alone.BThe required write invariant and the affected failure slice.CThe candidate's best single trace.DA subset chosen after excluding every failure.
A patch has zero failures in 100 representative independent trials. What can be claimed?
AEvery untested provider is covered.BThe true failure probability is zero.CNo failures were observed under that finite test; population risk remains uncertain.DThe failure is mathematically impossible.
A patch passes known write regressions and 50 new read-only tasks. What remains untested for broad write release?
AThe existence of any readable output.BNothing because the new tasks passed.CAll known regressions by definition.DFresh write recovery, concurrency, and provider-boundary cases.
Failure cases used repeatedly to tune a patch should be treated as what?
AProof of universal reliability once they pass.BUntouched final evaluation.CDevelopment/regression cases, with fresh estimation evidence still needed.DInvalid data to delete.
Using the supplied evidence, explain how you would make a scoped release decision without treating zero observed failures as zero risk. Name one observation that would change your conclusion. Rate confidence from 1 to 5.
Not yetGetting thereConfident
Wrap-up
Release decisions need task slices and severity. State the evidence that supports the decision and the limits of the test.