An incident interview often gives you a symptom before it gives you a trace. Resist the urge to name a model problem immediately. A wrong final answer may result from stale retrieval, a malformed tool argument, a service timeout, a hidden authorization failure, or an incorrect summary of a correct result. Your first task is to locate the earliest transition where observed behavior diverged from the contract.
Build a timeline with task ID, step ID, action, inputs, result class, and verified effect. Preserve the model and tool versions, but do not expose secrets or unnecessary user content. A trace should let you compare what the model requested with what the service did. If the service returns a structured error that the model reports as success, the failure is in interpretation or orchestration. If the service itself reports success without committing, investigate the service contract.
Use counterfactual tests after identifying a candidate cause. Replace the suspect tool result with a known correct result and replay the remaining read-only interpretation step. If the failure disappears, you have narrowed the problem. This is not proof that all future tasks are fixed. Build a regression case that reproduces the exact causal boundary, then test related cases.
Worked example
This invented trace concerns a duplicate meeting invitation.
Step
Action
Result
Provider state
1
Create invitation, key K1
Timeout
Invitation I1 exists
2
Create invitation, key K2
Completed
I1 and I2 exist
3
Summarize
"Invitation sent"
Two invitations
The first unsafe transition is step 2. The timeout in step 1 is an uncertain outcome, not a confirmed failure. Step 2 changes operation identity and permits another effect. The final summary also hides the duplicate, but fixing the wording alone would leave the incident mechanism intact. The proposed correction reconciles K1 before any new invitation.
Containment is to stop new writes for affected tasks and identify duplicate provider records. A compensating cancellation needs its own authorization and evidence. The regression test commits I1, drops the response, restarts the worker, and expects one invitation. This tests the distributed failure window rather than simply mocking every call as successful.
Exercise and solution
A tool returns permission_denied for document D9. The next model message says "I updated D9". Name the first failure, a safe immediate response, and a regression test.
The service correctly rejected the write. The first wrong transition is the model or controller converting rejection into completion. The immediate response states that the update did not occur and gives the access failure. The test supplies the structured denial and asserts that no completion claim is shown and no alternate unauthorized tool is tried. Award one point each for correct component attribution, truthful state, no permission bypass, and a regression assertion tied to the user-visible result.
Distinguish the incident from its report
An agent incident has at least two observable layers: the service effects and the account shown to the user. They can fail separately. The invitation trace above contains both a duplicate effect and an incomplete summary. A repair that changes only the summary leaves the duplicate mechanism. A repair that prevents duplicates but still reports denials as success leaves a trust problem.
Use a minimal incident record that distinguishes observation from hypothesis.
code
1incident: schedule-duplicate-72observed:3 provider invitation IDs: I1, I24 same intended meeting: E95 operation keys: K1, K26 first response: timeout7hypothesis:8 retry controller interpreted timeout as confirmed failure9containment:10 pause fresh writes for unresolved operation tasks11next_test:12 commit first effect, drop response, restart, inspect effect count13unknown:14 number of affected tasks outside the inspected trace
The unknown field matters. Finding one duplicate proves the mechanism can occur in that case. It does not identify the full affected population. Search the bounded relevant logs for the same signature and preserve the query definition. Do not estimate the impact by multiplying a guessed rate across every user.
A second worked case: callbacks arrive out of order
Now inspect a different failure. A job service sends status notifications with monotonically increasing event sequence numbers.
Arrival order
Provider sequence
Status
Local action
First
12
Completed
Mark task completed
Second
11
Running
Mark task running
Third
12
Completed duplicate
Send a second completion notice
The provider's execution may be correct. The local event consumer has two defects. It accepts an older sequence after a newer one, regressing the task state. It also treats a duplicate completion as a new user-notification event. The repair needs an ordered state update and an idempotent notification boundary.
A conditional update can accept an event only when its sequence is newer than the stored sequence under the provider's documented ordering contract. The state transaction should also create one durable notification intention keyed by execution and terminal event meaning. A sender processes that intention with a stable external operation identity and reconciles uncertain provider sends. A duplicate terminal event returns the existing state and notification intention rather than creating another one. Updating state once is insufficient if a separate sender can duplicate the external notification after a crash. If the provider does not promise monotonic sequences, use its authoritative status query or another documented version field. Do not invent ordering from client arrival time.
Design the discriminating test
The test supplies events in the order 12, 11, 12. It expects final state completed, stored provider sequence 12, and one durable completion-notification intention. A separate effect test commits the notification at the provider, loses the response, and restarts the sender; the stable identity and provider reconciliation must yield one external notice. Without a suitable provider contract, the result remains uncertain rather than claiming exactly one delivery. A second test supplies 11 then 12 and expects the same terminal state. These tests distinguish correct event handling from a workaround that simply ignores every event after the first.
A real provider may send a later reversal or correction. If the contract permits that, terminal state is not necessarily immutable forever. Model the correction explicitly with its own event version and user-facing meaning. The teaching example assumes that completed cannot regress to running for the same execution; that assumption belongs in the test fixture and contract.
Misconceptions to reject
"The last message received is the latest state" confuses delivery order with event order. Networks and queues can duplicate or reorder notifications.
"Fixing the final text fixes the incident" addresses reporting but may leave incorrect external effects. Conversely, correct effects do not excuse a misleading completion claim.
Transfer exercise
A document-export job reports failed at sequence 8, then completed at sequence 9 after a documented retry, and an old failed event arrives again. The service contract permits the retry transition. Describe the final state and notification behavior.
The consumer accepts sequence 9 as current and ignores the older repeated sequence 8 for state updates. It records the retry history and sends at most the intended completion notice for sequence 9. It does not enforce a blanket rule that any failed state is permanently terminal when the provider permits recovery. Award one point for ordered comparison, one for contract-aware transitions, one for duplicate handling, and one for preserving evidence of the earlier failure.
Interview probe
Original practice: The agent says it succeeded, but the user reports no change. Where do you start? A strong answer compares the proposed action, service result, and resource state, then identifies the first incorrect transition. Follow up with a lost response. A weak answer changes the prompt before preserving evidence.
A trace shows K1 timed out and K2 then created a second invitation. Which transition first permits the duplicate?
AThe user's original valid request.BThe provider returning I1 after commit.CThe final summary omitting the duplicate.DThe fresh K2 write before K1 is reconciled.
Provider event sequence 12 completed arrives before sequence 11 running. What should an ordered consumer do under this contract?
AUse arrival time because it reflects the latest delivery.BKeep sequence 12 completed as current and retain the older event as history.CAccept sequence 11 because running is a safer conservative status.DCreate a new execution for sequence 11 so both states remain current.
A duplicate terminal event arrives with the same provider sequence. What should happen to the user notification?
AReturn existing terminal handling without another notice.BReset the task so the event can run normally.CCreate a new task ID for the duplicate.DSend again because the event was delivered again.
A valid permission denial becomes an assistant success claim. Which component boundary should be tested?
AThe network timeout duration alone.BOnly the provider's permission policy.CThe conversion of structured tool result into user-visible completion.DOnly the model's training dataset.
One duplicate incident is confirmed. Which impact claim is justified before a broader log audit?
AThe inspected task duplicated; affected population remains unknown.BNo other task can be affected.CThe population rate equals one divided by all registered users.DEvery task using that tool duplicated.
Using the supplied evidence, explain how you would identify the first wrong transition and handle out-of-order status events. Name one observation that would change your conclusion. Rate confidence from 1 to 5.
Not yetGetting thereConfident
Wrap-up
Investigate transitions before blaming a component. Fix the earliest causal failure and separately correct misleading reporting.