Design bounded retry and fairness rules for mixed tenant workloads.
A failed attempt contains information. A malformed payload probably needs correction. A rate-limit response suggests waiting according to a service contract. A connection interruption may be transient, but the previous effect can be unknown. A worker that retries all three cases identically wastes capacity and can repeat a business action.
Classify errors at the boundary that understands them. Transport failures need safe operation identity. Explicit permanent validation errors need a terminal result that a user can fix. Rate limits need coordinated throttling. Authentication failures may require operator action rather than immediate repetition. Keep the original error and attempt history; a final generic retry-exhausted message can hide the cause that matters.
Backoff reduces the rate of attempts over time. Randomness spreads callers across the waiting interval. Neither provides a total retry budget by itself. Set an attempt limit or elapsed-time budget, and identify which layer owns retries. Three layers each making three attempts can produce up to twenty-seven downstream calls for one original request. SDK retries count toward the system budget.
A poison job repeatedly fails for a reason that ordinary retries cannot repair. Isolate it with its input identity, error, and replay policy. Replaying unchanged input retains its operation identity. A correction needs an explicitly authorized input version or new operation under the service contract, while preserving audit correlation. A dead-letter queue is an investigation mechanism, not proof that the business workflow completed.
Fairness adds another constraint. A single customer can fill a shared queue with expensive work and delay everyone else. Use per-tenant admission limits, bounded active jobs, or a scheduler that alternates eligible tenants. Priority queues need starvation controls. A permanently busy premium queue must not make ordinary accepted work impossible to complete.
Worked example
Tenant A submits 900 ten-second jobs. Tenant B submits 10 one-second jobs. A strict FIFO queue with five workers can keep B waiting roughly thirty minutes if A occupies the front. A scheduler with a per-tenant active cap of four leaves one worker available for B. B's jobs finish in roughly ten seconds of service after admission, assuming no other bottleneck. This is an invented policy, not a universal allocation recommendation.
One A job returns a permanent schema error. It becomes failed after one attempt. Another receives 429 and waits under the shared rate limiter. A timed-out external write uses the original operation key during recovery.
Make retry ownership explicit
This original policy table is for a fictional export worker. The provider's documented semantics determine the final classification. Do not infer every 4xx or 5xx response has identical behavior.
Observation
Classification
Next action
Invalid required field
Permanent input error
Fail with correction detail
Documented 429
Capacity or rate restriction
Respect delay and shared limiter
Connection lost after write may commit
Unknown effect
Recover with stable operation identity
Temporary unavailable response
Potentially transient
Bounded backoff if safe
Expired credential
Operator or user action
Stop automatic repetition and surface cause
The table separates eligibility from timing. Backoff answers when another attempt may occur. It does not decide whether repeating the operation is safe or useful. A permanent input error can wait an hour and remain invalid. An unknown external write can be unsafe even if the retry occurs politely.
Count attempts across every layer. Consider this teaching trace:
The worker must know whether the SDK's configured maximum counts total attempts or retries after the first. Those conventions differ. Record the chosen interpretation in configuration and telemetry. A setting named retries=3 can otherwise be mistaken for three total attempts when it actually permits four.
An elapsed budget can expire before an attempt budget
Suppose a job has at most five total provider attempts and a thirty-second recovery window. Each call can consume eight seconds. After three slow attempts, only six seconds remain, so another eight-second attempt would exceed the overall budget. The scheduler should use the remaining deadline rather than independently honoring every local maximum. A timeout on one call should also account for connection setup and response work according to the client library's behavior.
When a budget expires, preserve the final error category and operation identity. Do not convert an unknown payment or export-side effect into a definite rejection merely because attempts ran out. The terminal workflow might be waiting for review rather than failed. Error taxonomy should describe what is known, while execution policy describes why automatic work stopped.
Inspect fairness as a scheduling artifact
A tenant-aware schedule can make capacity allocation visible:
Slot
Tenant
Job cost estimate
Policy reason
1
A
10 seconds
Eligible bulk work
2
A
10 seconds
Within A's active cap
3
A
10 seconds
Within A's active cap
4
A
10 seconds
A reaches cap four
5
B
1 second
Capacity reserved for other tenant
This policy improves B's access in the supplied case, but it is not automatically optimal for every workload. If A is the only tenant, a work-conserving scheduler may let it borrow the unused slot. When B arrives, the scheduler can stop assigning new borrowed work without interrupting an irreversible effect already in progress. State whether borrowing is allowed and how long a new tenant might wait for a slot.
Misconceptions to correct
The first misconception is that dead-lettering completes the user's work. It stops ordinary processing and preserves a failed message for investigation. The business request still needs a visible outcome, owner, and replay decision.
The second misconception is that priority alone provides fairness. A permanently busy high-priority class can starve lower classes. Add a guaranteed share, aging, or another explicit policy when accepted work must eventually receive service.
A third tempting repair is replaying poison jobs with new event identities. That can hide the connection to previous attempts and bypass deduplication. Replay of unchanged input preserves the payload-bound operation identity and records a new attempt. Corrected input must not silently reuse an old payload-bound idempotency key. Preserve audit/business correlation, but create an explicitly authorized new input version or operation under the service's contract, with its own key where required.
Extend the exercise
A job has a twenty-second deadline. Two attempts consume seven seconds each and backoff consumes three seconds total. Only three seconds remain. Explain whether a six-second next attempt fits. It does not. Award one point for the remaining budget, one for stopping or shortening according to the client contract, and one for retaining the correct unknown or failed business state.
Exercise and solution
A client retries twice after the initial attempt. The worker also retries twice. How many remote attempts are possible if neither layer knows about the other? Nine. Propose a correction. Let one layer own the retry budget and account for any SDK attempts. Award one point for nine, one for ownership, and one for preserving safe operation identity.
Interview probe and wrap-up
Why not give every failed job maximum priority? A strong answer explains how retry traffic can starve healthy new work and overload a recovering dependency. Follow up with replay after an input correction. A weak answer says exponential backoff guarantees safety. A retry policy must specify eligibility, timing, budget, identity, and who waits when capacity is scarce.
A documented invalid-field error repeats with unchanged input. Best action?
ARecord a permanent input failure and correction path.BDouble the elapsed retry budget.CChange the event ID and replay.DRetry through another worker pool.
AOrdinary processing stopped and investigation/replay is needed.BThe customer's requested effect completed.CEvery duplicate was removed.DThe input is safe to discard.
A high-priority queue is continuously busy. What protects accepted lower-priority work?
AUnlimited high-priority preference.BA defined share or aging policy.COnly FIFO inside each queue.DA larger shared worker pool without any tenant allocation rule.
Can you classify a failure, count nested attempts, and defend a retry/fairness policy without confusing stopped execution with completed business work? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.