Resolve a timeout without applying an operation twice.
A network timeout does not tell you whether a write happened. The request may have failed before it reached the server. The server may also have completed the write while the response was lost. An agent that treats both cases as failure can repeat a payment, invitation, or reservation. This is a distributed-systems problem even when the user sees only a chat window.
Define a logical operation before sending the first attempt. Its identity must survive retries and process restarts. A new random key for every retry defeats the purpose because the service sees each attempt as a different action. The operation record binds the key to the target, action, and relevant payload. Reusing the same key with a different payload must fail rather than return a misleading prior result.
A local flag such as sent = true does not solve the problem. The process can crash after the remote effect and before it stores the flag. Prefer a provider-supported idempotency contract or a transactional operation record at the service that owns the effect. Where the provider offers neither, use a queryable business identifier and reconciliation. If the effect cannot be reliably identified, pause uncertain operations for review. Do not advertise exactly-once behavior that the system cannot prove.
Read operations can often retry with bounded backoff. Writes need more care. Distinguish a rejected input, an expired authorization, a rate limit, a transient transport problem, and an unknown outcome. Retry policy should use structured error types. A model reading the word "error" is too weak a controller for deciding whether money moves twice.
Worked example
A fictional booking tool receives operation K44 for seat S12. The service stores K44 as pending, books the seat, and records reservation R88 as completed. The response is lost. The agent process restarts with K44 in its persisted state.
The next action queries K44. The service returns completed with reservation R88. The agent reports one reservation and does not issue a fresh booking. If K44 is pending, it polls within a deadline or escalates. If the service authoritatively confirms that K44 never executed and the same key remains valid, the controller can retry K44. The key's retention period matters. A retry after the provider has forgotten the key may create another reservation.
Exercise and solution
An email provider accepts client operation M5, sends one email, and stores its message ID. Your application loses the response. A worker proposes retrying with M6 because M5 "failed". Explain the correction and a test.
The correction retains M5 and asks for its status or retries under the provider's documented idempotency semantics. It stores the returned message ID. The test drops the first response after the provider commits, then restarts the worker. The expected result is one provider message and one completed local operation. Award one point each for stable identity, reconciliation before a new effect, restart coverage, and a provider-side effect count. A test that counts local function calls alone is insufficient.
Follow the write through a crash
Use this state table to decide where recovery can proceed without guessing.
Local state
Provider evidence
Safe interpretation
Prepared
No request sent
The operation can start
Sent
No final response
Outcome is unknown
Sent
Provider says completed
Reconcile local completion
Completed
Provider result ID stored
Report the stored result
Sent
Provider authoritatively rejects before effect
Correct the cause within the contract
The difficult row is the second. "Sent" does not establish whether the provider committed. A local database transaction can make the local operation record durable, but it cannot atomically include an unrelated provider unless the provider participates in a suitable protocol. This is why the write boundary needs cooperation or a reconciliation path.
The following controller is illustrative pseudocode. It omits authentication, locks, backoff, and provider-specific retention handling.
code
1operation = load_or_create(user_intent_id, canonical_payload)2if operation.payload_hash != hash(canonical_payload):3 reject("same operation identity, different intent")4if operation.status == "completed":5 return operation.provider_result6if operation.status == "uncertain":7 result = provider.lookup(operation.id)8 if result.is_completed:9 persist_completed(operation.id, result)10 return result11 if not result.authoritatively_safe_to_retry:12 return needs_reconciliation(operation.id)13return provider.execute(operation.id, canonical_payload)
The final line is not sufficient by itself. The caller must persist the returned result, and a timeout must move the operation into an uncertain state. Two workers must not independently conclude that they own the same transition. Use an atomic claim or another concurrency mechanism that matches the service contract.
Payload identity is part of intent
A client may accidentally reuse an operation ID after changing the amount or destination. Returning the old result without detecting the mismatch would tell the user that the new request succeeded when it did not. Executing the changed payload would defeat idempotency. The correct behavior is to reject the conflicting reuse.
Stripe's documented implementation compares parameters for reused keys and describes key retention. Those are provider-specific guarantees. A different service may implement a different scope, duration, or error policy. Read its contract before treating the word "idempotent" as sufficient. In particular, reusing an expired key can become a new operation.
A second worked case
A teaching service retains operation keys for 24 hours. Request K7 created order O7 at 09:00 Monday, but the caller lost the result. On Wednesday, the caller retries K7 without a business-record lookup. The key may have been pruned, so the same text key no longer guarantees deduplication. The controller should use an authoritative order lookup by a stable business identifier or send the case to reconciliation if no reliable lookup exists.
The conclusion is limited. Key expiry does not prove that the retry will duplicate an order; it removes the guarantee that prevents it. Good incident language distinguishes known duplication from an unsafe unresolved state.
Misconceptions to reject
"Retrying the same bytes is always safe" ignores the provider's interpretation, key scope, and retention. Byte equality alone does not supply a deduplication mechanism.
"A timeout means the server rolled back" assumes that client observation controls remote commit. The server can finish after the client stops waiting, and the response can be lost after commit.
Transfer exercise
Two workers receive the same logical refund task. Both read local status prepared. One provider request completes before either worker stores completion. Specify a protection and its limit. The model solution atomically claims the task or uses a stable provider idempotency key for both attempts, preferably both. It binds the key to the refund payload and records the provider result. The local claim reduces concurrent work; provider idempotency covers the remote response-loss window. Award one point for each boundary and one for rejecting the claim that a local lock alone guarantees an external effect exactly once.
Interview probe
Original practice: Can you guarantee exactly one external action with a local database and an arbitrary HTTP API? A strong answer identifies the commit-response gap and the need for provider cooperation or reconciliation. Follow up with idempotency-key expiry. A weak answer says to retry three times or set temperature to zero.
A write times out after the provider may have committed. What is the best next step?
ARetry the original key after its retention period without checking provider state.BRetry as a new logical operation.CAssume the provider rolled back.DResolve the original operation using its stable identity.
Two workers read the same prepared operation. Which design covers both local concurrency and response loss?
AA local in-memory flag only.BA stable provider key only, with no payload binding.CAn atomic local claim plus a payload-bound provider operation identity.DA new random key per worker.
A reused key now carries a different amount. What should the service do under this contract?
AReturn the old result without checking the payload.BReject the key-payload mismatch.CExecute the new amount under the old key.DGenerate a replacement key automatically and execute the changed amount.
A provider may prune keys after 24 hours. An uncertain request is retried after two days. What is true?
AA new key is safer than the old key.BThe same key guarantees no new effect forever.CThe previous effect certainly failed.DThe key-retention guarantee may no longer protect this retry.
A local lock prevents concurrent workers. A process then crashes after remote commit but before storing completion. Which statement explains the remaining recovery need?
AThe lock serialized local workers but did not make local state and remote commit atomic.BCommitting the local prepared record before sending proves the provider completed the write.CRenewing the lock lease removes uncertainty about a lost provider response.DHolding the lock until the provider's retention period expires will establish whether the write happened.
Without looking at the worked solution, explain how you would recover a timed-out write across a worker crash and provider-key expiry. Name the service evidence you need. Rate confidence from 1 to 5 and identify the boundary you still cannot justify.
Not yetGetting thereConfident
Wrap-up
A retry belongs to the same logical operation. Keep its identity and reconcile unknown outcomes before creating another effect.