Lesson 3 of 4 · 35 min

Retries can turn an outage into overload

Design bounded retry and fairness rules for mixed tenant workloads.

A failed attempt contains information. A malformed payload probably needs correction. A rate-limit response suggests waiting according to a service contract. A connection interruption may be transient, but the previous effect can be unknown. A worker that retries all three cases identically wastes capacity and can repeat a business action.
Classify errors at the boundary that understands them. Transport failures need safe operation identity. Explicit permanent validation errors need a terminal result that a user can fix. Rate limits need coordinated throttling. Authentication failures may require operator action rather than immediate repetition. Keep the original error and attempt history; a final generic retry-exhausted message can hide the cause that matters.
Backoff reduces the rate of attempts over time. Randomness spreads callers across the waiting interval. Neither provides a total retry budget by itself. Set an attempt limit or elapsed-time budget, and identify which layer owns retries. Three layers each making three attempts can produce up to twenty-seven downstream calls for one original request. SDK retries count toward the system budget.
A poison job repeatedly fails for a reason that ordinary retries cannot repair. Isolate it with its input identity, error, and replay policy. Replaying unchanged input retains its operation identity. A correction needs an explicitly authorized input version or new operation under the service contract, while preserving audit correlation. A dead-letter queue is an investigation mechanism, not proof that the business workflow completed.
Fairness adds another constraint. A single customer can fill a shared queue with expensive work and delay everyone else. Use per-tenant admission limits, bounded active jobs, or a scheduler that alternates eligible tenants. Priority queues need starvation controls. A permanently busy premium queue must not make ordinary accepted work impossible to complete.

Worked example

Tenant A submits 900 ten-second jobs. Tenant B submits 10 one-second jobs. A strict FIFO queue with five workers can keep B waiting roughly thirty minutes if A occupies the front. A scheduler with a per-tenant active cap of four leaves one worker available for B. B's jobs finish in roughly ten seconds of service after admission, assuming no other bottleneck. This is an invented policy, not a universal allocation recommendation.
One A job returns a permanent schema error. It becomes failed after one attempt. Another receives 429 and waits under the shared rate limiter. A timed-out external write uses the original operation key during recovery.

Make retry ownership explicit

This original policy table is for a fictional export worker. The provider's documented semantics determine the final classification. Do not infer every 4xx or 5xx response has identical behavior.
ObservationClassificationNext action
Invalid required fieldPermanent input errorFail with correction detail
Documented 429Capacity or rate restrictionRespect delay and shared limiter
Connection lost after write may commitUnknown effectRecover with stable operation identity
Temporary unavailable responsePotentially transientBounded backoff if safe
Expired credentialOperator or user actionStop automatic repetition and surface cause
The table separates eligibility from timing. Backoff answers when another attempt may occur. It does not decide whether repeating the operation is safe or useful. A permanent input error can wait an hour and remain invalid. An unknown external write can be unsafe even if the retry occurs politely.
Count attempts across every layer. Consider this teaching trace:
code
1user operation P12  worker attempt W1 -> SDK attempts S1,S2,S33  worker attempt W2 -> SDK attempts S4,S5,S64  worker attempt W3 -> SDK attempts S7,S8,S95maximum provider attempts = 3 * 3 = 9
The worker must know whether the SDK's configured maximum counts total attempts or retries after the first. Those conventions differ. Record the chosen interpretation in configuration and telemetry. A setting named retries=3 can otherwise be mistaken for three total attempts when it actually permits four.

An elapsed budget can expire before an attempt budget

Suppose a job has at most five total provider attempts and a thirty-second recovery window. Each call can consume eight seconds. After three slow attempts, only six seconds remain, so another eight-second attempt would exceed the overall budget. The scheduler should use the remaining deadline rather than independently honoring every local maximum. A timeout on one call should also account for connection setup and response work according to the client library's behavior.
When a budget expires, preserve the final error category and operation identity. Do not convert an unknown payment or export-side effect into a definite rejection merely because attempts ran out. The terminal workflow might be waiting for review rather than failed. Error taxonomy should describe what is known, while execution policy describes why automatic work stopped.

Inspect fairness as a scheduling artifact

A tenant-aware schedule can make capacity allocation visible:
SlotTenantJob cost estimatePolicy reason
1A10 secondsEligible bulk work
2A10 secondsWithin A's active cap
3A10 secondsWithin A's active cap
4A10 secondsA reaches cap four
5B1 secondCapacity reserved for other tenant
This policy improves B's access in the supplied case, but it is not automatically optimal for every workload. If A is the only tenant, a work-conserving scheduler may let it borrow the unused slot. When B arrives, the scheduler can stop assigning new borrowed work without interrupting an irreversible effect already in progress. State whether borrowing is allowed and how long a new tenant might wait for a slot.

Misconceptions to correct

The first misconception is that dead-lettering completes the user's work. It stops ordinary processing and preserves a failed message for investigation. The business request still needs a visible outcome, owner, and replay decision.
The second misconception is that priority alone provides fairness. A permanently busy high-priority class can starve lower classes. Add a guaranteed share, aging, or another explicit policy when accepted work must eventually receive service.
A third tempting repair is replaying poison jobs with new event identities. That can hide the connection to previous attempts and bypass deduplication. Replay of unchanged input preserves the payload-bound operation identity and records a new attempt. Corrected input must not silently reuse an old payload-bound idempotency key. Preserve audit/business correlation, but create an explicitly authorized new input version or operation under the service's contract, with its own key where required.

Extend the exercise

A job has a twenty-second deadline. Two attempts consume seven seconds each and backoff consumes three seconds total. Only three seconds remain. Explain whether a six-second next attempt fits. It does not. Award one point for the remaining budget, one for stopping or shortening according to the client contract, and one for retaining the correct unknown or failed business state.

Exercise and solution

A client retries twice after the initial attempt. The worker also retries twice. How many remote attempts are possible if neither layer knows about the other? Nine. Propose a correction. Let one layer own the retry budget and account for any SDK attempts. Award one point for nine, one for ownership, and one for preserving safe operation identity.

Interview probe and wrap-up

Why not give every failed job maximum priority? A strong answer explains how retry traffic can starve healthy new work and overload a recovering dependency. Follow up with replay after an input correction. A weak answer says exponential backoff guarantees safety. A retry policy must specify eligibility, timing, budget, identity, and who waits when capacity is scarce.

Sources

docsAWS exponential backoff and jitteraws.amazon.comdocsMDN HTTP 429 statusdeveloper.mozilla.orgdocsStripe webhook delivery and verificationdocs.stripe.com

Checkpoint

A documented invalid-field error repeats with unchanged input. Best action?

ARecord a permanent input failure and correction path.BDouble the elapsed retry budget.CChange the event ID and replay.DRetry through another worker pool.
Sign up free to answer and see why

Checkpoint

Worker permits three total attempts; SDK permits three per worker attempt. Maximum calls?

ATwelve.BThree.CSix.DNine.
Sign up free to answer and see why

Checkpoint

Twenty-second deadline; calls used fourteen seconds and backoff three. Can a six-second attempt fit?

AYes, because backoff does not count.BYes, if the queue is short.CNo, only three seconds remain.DYes, because attempt count remains.
Sign up free to answer and see why

Checkpoint

What does dead-letter placement prove?

AOrdinary processing stopped and investigation/replay is needed.BThe customer's requested effect completed.CEvery duplicate was removed.DThe input is safe to discard.
Sign up free to answer and see why

Checkpoint

A high-priority queue is continuously busy. What protects accepted lower-priority work?

AUnlimited high-priority preference.BA defined share or aging policy.COnly FIFO inside each queue.DA larger shared worker pool without any tenant allocation rule.
Sign up free to answer and see why

Can you classify a failure, count nested attempts, and defend a retry/fairness policy without confusing stopped execution with completed business work? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.