Lesson 1 of 4 · 60 min

An asynchronous API needs an honest contract

Specify a job resource that survives disconnects and supports cancellation races.

Moving a slow operation to a worker shortens the request, but it also changes what success means. The initial response can prove that the service accepted responsibility for work. It cannot prove that the result exists. A useful asynchronous API makes that distinction visible through a durable job resource with a stable identifier.
Describe the job lifecycle before choosing a queue. A document export might move through accepted, running, succeeded, failed, and cancelled. Accepted means the request and its work intention are durable. Running means a worker holds a current lease or claim. Succeeded means a complete result is available at the declared location. Failed should distinguish a permanent input problem from exhausted transient attempts. A cancelled state needs a precise boundary, particularly when an external effect cannot be reversed.
Queue presence is not a good user-facing state. Messages can be redelivered, hidden temporarily, or removed after acknowledgement. Keep the authoritative job state in a durable record. A message tells a worker which job needs attention. A worker may receive that message after another worker already completed the job, so it checks the record before acting.
Cancellation is a request, not a time machine. For a pure computation, a worker can check a cancellation flag between bounded chunks and discard partial output. For an irreversible send, cancellation can succeed only before the send begins. If completion wins the race, return the completed result. Do not label a job cancelled while leaving a delivered business effect unexplained.

Worked example

A fictional export contains 12,000 rows. The client submits operation X4. The API transaction inserts job J4 in accepted state and its outbox intention, then returns J4. A worker claims J4 and writes a complete immutable attempt result that users cannot access directly. One guarded job-row transaction checks current ownership and cancellation, then binds that result pointer and succeeded together. Only the committed job pointer makes output user-visible. A browser that disconnected after submission can recover J4 using the original operation identity.
At row 8,000, a cancellation request arrives. The worker or cancellation transaction observes it before the guarded result-pointer commit, leaves attempt output inaccessible for cleanup, and marks cancelled. In a second schedule, the guarded result-pointer and succeeded transaction commits first. The cancellation request returns already completed. Both schedules have a truthful terminal state.

Specify each state as a contract

A job status should describe evidence the system can defend. The following table gives one original export contract. It deliberately separates execution state from the browser connection and from the lifetime of a download URL.
StateRequired evidenceUser action
acceptedJob and work intention committedWait or request cancellation
runningCurrent worker claim recordedInspect progress or request cancellation
succeededComplete result published and linkedRetrieve authorized result
failedTerminal reason and attempt history recordedCorrect input or request controlled retry
cancelledCancellation won before publicationSubmit a new operation if desired
The table is not a claim that all job systems use these exact states. Its value is consistency. If running means a worker claimed the job, the application must also handle an expired claim after a crash. A lease timestamp can identify when recovery may claim work again. It should not automatically prove that the original worker stopped executing. Long-running work needs a fencing or equivalent ownership strategy if two workers could otherwise publish conflicting results.
Here is a teaching record shape:
code
1jobId: J42operationKey: X43accountId: A24state: running5stateVersion: 36inputFingerprint: weekly-report-september7attemptId: W78leaseExpiresAt: 10:05:009resultId: null10cancelRequested: false
The lease and attempt identity belong to execution control. The operation key belongs to the user's logical request. A recovery attempt gets a new attempt ID while retaining job and operation identity. This lets the audit show multiple attempts without claiming multiple exports.

A crash after the worker claim

Suppose W7 claims J4 and crashes before writing output. The job remains running until the recovery policy detects the expired lease. A new worker W8 claims the same job under a higher state version or ownership token. It resumes or restarts the bounded computation according to the job contract. If W7 was only paused and later resumes, it must not overwrite W8's final result using stale ownership.
Use immutable attempt-specific output objects. They remain inaccessible through the user API until a single guarded database transaction checks current owner, state version, and cancellation, then stores the result pointer and succeeded state together. The user API serves only the committed pointer after authorization. A stale worker cannot change that pointer because the conditional transaction fails. Unreferenced attempt objects remain private and are cleaned up under a retention policy. Object storage writes and database status are not one atomic operation; the design makes the object write happen first and treats the database pointer commit as publication.
Cancellation intersects with this recovery schedule. A cancellation request records intent in the durable job. W8 sees it before starting new chunks and changes the job to cancelled if no final publication occurred. The browser must not send a cancellation request automatically on every navigation unless that is the explicit product contract. Users often close a tab precisely because they expect a long export to continue.

Misconceptions to correct

The first misconception is that a queue message's visibility timeout is the job state. Visibility controls delivery to workers. It does not tell the user whether a complete result exists, whether input was invalid, or whether cancellation won. Keep execution and user-facing truth connected through the job record.
The second misconception is that a 202 response means the worker has started. It means the API accepted the request under the stated contract. A job can remain accepted while workers are unavailable. The interface should represent that wait and the operator should monitor its age.

Extend the exercise

Add a worker lease expiring at 10:05. W8 claims at 10:06, then W7 resumes at 10:07. Draw the allowed publication rule. The model answer permits only the current owner to publish or finalize the result. A stale worker can discard its temporary output and report its obsolete attempt. Award one point for stable job identity, one for rejecting stale publication, and one for preserving cancellation intent across recovery. This exercise tests ownership, a guarantee that a simple queue-and-spinner design leaves unstated.

Exercise and solution

Design responses for three observations: job accepted but not claimed, job succeeded with an expired download link, and job failed because an input field is invalid. The model answer returns pending progress for the first, refreshes authorized result access without rerunning the export for the second, and reports a permanent field error for the third. Award one point per correct distinction and one for keeping job identity stable when a download link changes.

Interview probe and wrap-up

Would you return 200 or 202 for initial acceptance? A strong answer starts with the documented semantics and client expectations, then separates initial acceptance from result retrieval. Follow up with duplicate submissions during a client reconnect. A weak answer treats a successful enqueue call as completed export delivery. The contract must let a user recover the truth without knowing which worker processed the job.

Sources

docsAWS transactional outbox patterndocs.aws.amazon.comdocsStripe idempotent request contractdocs.stripe.comdocsIETF RFC 9110, section 15.3.3: 202 Acceptedietf.org

Checkpoint

What evidence supports accepted in this contract?

AThe final file exists.BThe request passed input validation but its job insert has not committed.CJob and delivery intention committed.DA worker has necessarily begun.
Sign up free to answer and see why

Checkpoint

A worker lease expires and a replacement claims the same job. The old worker may resume. Which publication rule protects the final result?

AFinalize the result pointer only through a conditional transaction that verifies current ownership and job state.BChange the user's operation key each time a worker is replaced.CPermit the old attempt to reuse its expired lease as proof of ownership.DLet both attempts publish and retain whichever result arrives last.
Sign up free to answer and see why

Checkpoint

Which identifier should change when recovery starts another attempt?

AThe user's operation key.BThe job ID.CThe final business request identity.DThe worker attempt ID.
Sign up free to answer and see why

Checkpoint

A completed file's URL expires. What is appropriate?

ARestore queue visibility for the old message.BRefresh authorized access to the existing result.CRerun the computation automatically.DMark the job failed.
Sign up free to answer and see why

Checkpoint

Cancellation arrives after final publication. What should the API report?

AAccepted as a second job.BCancelled because the request arrived eventually.CAlready completed under the terminal-state contract.DRunning until the user retries cancellation.
Sign up free to answer and see why

Can you explain what evidence supports each job state and how a recovered worker prevents a stale attempt from publishing over its result? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.