Build an outcome verifier that cannot be fooled by narration
Define success, partial success, and invalid shortcuts for a task.
An evaluation is useful only if its verifier reflects the user's goal. For an action task, a fluent final answer is weak evidence. For a research task, an exact string match may reject an equally correct explanation. Start by identifying the observable properties that make the result useful.
Separate required outcomes from optional path choices. A calendar task may require one event at the chosen time with the specified attendees. The agent can reach that state through different read sequences. Scoring one exact sequence can penalize safe recovery. However, some path constraints are essential, such as never accessing another tenant's data or never sending before approval. These need explicit invariant checks even if the final state looks correct.
Use several verifier types where needed. A deterministic check can count provider objects and inspect fields. A rubric can assess explanation completeness. A human review may resolve ambiguous evidence. A model judge can assist with prose evaluation, but it needs a clear rubric, calibration examples, and review of disagreements. Do not ask the same ungrounded model merely whether its own answer is good.
Hidden failures matter. A task that creates the requested file and also modifies unrelated files should fail a constrained coding task. A research agent that supplies the right answer with invented citations should fail evidence quality. A high aggregate success rate can hide a small number of severe unauthorized effects. Record these separately rather than averaging them into a harmless-looking score.
Worked example
A teaching task asks an agent to create a draft event for two attendees, without sending invitations. The verifier checks that exactly one draft exists, its time and attendees match, and the provider send count is zero. A successful final sentence alone earns no outcome credit.
Candidate A creates the correct draft and sends two notifications. Candidate B creates the correct draft but uses a different valid read sequence. Candidate C creates nothing and gives a perfect verbal description. The result is A fail for unauthorized effect, B pass, C fail for missing artifact. This demonstrates why both final-state checks and path invariants belong in the evaluation.
Exercise and solution
Design a verifier for "Summarize three supplied documents and cite each factual claim". The agent may paraphrase and use a different section order.
A model solution checks coverage of all three document IDs, maps claims to source passages, tests that quotations or paraphrases support the claims, and marks unsupported statements. It does not require identical wording or section order. It treats missing documents as incomplete and fabricated sources as a failure. Award one point each for coverage, claim support, flexible phrasing, and explicit failure rules. A source URL existing is insufficient if its content does not support the statement.
Make the verifier inspectable
A verifier should expose the rule that turns evidence into a score. This teaching specification grades a draft-event task with a prohibited send effect.
code
1required:2 event_count_for_operation == 13 event.status == "draft"4 event.start == approved_start5 set(event.attendees) == set(requested_attendees)6forbidden:7 provider_send_count_for_task > 08 reads_outside_authorized_tenant > 09observation_coverage:10 provider_effect_query_complete_for_task == true11 authorized_read_audit_complete_for_task == true12result:13 fail if a required condition is false or a forbidden effect is confirmed14 inconclusive if required observation coverage is missing15 pass only if all required conditions and coverage checks hold16 and no forbidden effect is present
This specification intentionally does not require a fixed number of search calls or a particular order of harmless reads. Those path details can be diagnostic metrics, but they are not the user's objective. By contrast, unauthorized reads belong in the failure conditions because a correct final draft does not undo disclosure.
The verifier also needs trustworthy instrumentation. If provider sends are logged only when the application receives a successful response, a lost response can hide a real send. Use evidence from the service that owns the effect where available, or state the observation limit. A zero count in incomplete logs is not the same as verified absence. The verifier must return inconclusive or needs-review when required observation coverage is missing. It must not convert missing telemetry into a pass. A confirmed violation can still fail immediately even when other evidence is incomplete.
A second worked case: a persuasive but unsupported report
A research task asks for three claims with citations to supplied documents. Candidate R produces the right-looking conclusion, cites real URLs, but one URL supports a different claim. Candidate S covers two claims correctly and marks the third unresolved. A single prose-quality score might rank R higher because it appears complete.
Candidate
Supported claims
Unsupported claims
Coverage status
R
2
1
Claims complete, evidence invalid
S
2
0
Explicitly incomplete
T
3
0
Complete and supported
Under a strict complete-and-supported requirement, only T passes. S is a useful partial result and should be represented as such. R fails evidence quality even if its unsupported claim later happens to be true. The task required a justified answer, not a lucky guess.
This distinction matters when a model judge is used. Give the judge the relevant source passages and a claim-support rubric. Ask it to identify unsupported statements, not merely rate confidence or style. Calibrate the judge on examples with known support and review disagreement cases. A judge can make mistakes, so its output is one component of the evidence system rather than a replacement for all checks.
Build adversarial cases for the verifier itself
A verifier can be wrong. Test it with a correct artifact using different harmless formatting, an incorrect artifact with copied expected wording, a missing source disguised as a citation, and an unauthorized effect hidden behind a correct final state. These are tests of the grading mechanism, separate from tests of the agent.
A useful verifier review asks both false-pass and false-fail questions. Does it accept a wrong result that imitates the expected answer? Does it reject a valid alternative because of punctuation or ordering? Measuring both prevents the team from making a grader stricter in ways that reduce validity.
Misconceptions to reject
"A real URL proves a claim is sourced" confuses source existence with source support. The passage must substantiate the associated statement.
"A stricter exact-match grader is always more reliable" can reject valid paraphrases or alternative safe paths. Strictness should apply to required meaning and effects, not irrelevant surface form.
Transfer exercise
An agent must create two files with specified content and leave all other files unchanged. It creates the correct two files but also edits a configuration file. Write a verifier and expected grade.
The model solution checks the two required contents and compares the bounded changed-file set against the allowed set. The extra configuration change fails the task even if tests pass. Award one point for each content check, one for the unchanged-files invariant, and one for a case that catches an incorrect answer with correct-looking final narration.
Interview probe
Original practice: When should you avoid exact-match grading? A strong answer distinguishes flexible expression from fixed state invariants. Follow up with a correct answer produced through an unauthorized tool call. A weak answer replaces every check with a single model rating.
An agent creates the correct draft but sends invitations despite a no-send requirement. How should it score?
APass because the artifact is correct.BFail the prohibited-effect invariant.CPass if the final answer says draft.DPass with only a formatting deduction.
A real URL is attached to a claim but its page discusses a different fact. What is the evidence status?
ASupported if another claim on the page is true.BVerified by matching the source's title.CSupported because the URL resolves.DUnsupported by that citation.
A correct answer uses different harmless wording. Which verifier behavior is preferable?
AScore only the number of citations.BReject every non-identical string.CCheck semantic requirements and required effects while allowing valid expression.DRemove all deterministic checks.
Provider sends are logged only when a successful response reaches the app. What limit does a zero log count have?
AIt proves no send occurred.BIt may miss committed sends whose responses were lost.CIt proves the provider rejected the request.DIt establishes all sends were authorized.
A coding task allows changes to two files. Both are correct, but config also changed. Which verifier catches the extra effect?
ACheck only the two required files' contents.BCompare the bounded changed-file set with the allowed set as well as checking required content.CAccept any change if the test suite passes.DInspect only the final assistant's listed filenames.
Using the supplied evidence, explain how you would test a verifier for false passes, false failures, and incomplete effect observation. Name one observation that would change your conclusion. Rate confidence from 1 to 5.
Not yetGetting thereConfident
Wrap-up
Verify the artifact and the required constraints. Allow valid alternative paths while detecting harmful shortcuts.