Operate the incident before writing the postmortem
Produce a short incident record with mitigation, evidence, and a verified recovery condition.
A founding engineer may be the only person available when a workflow fails. The first task is to reduce user harm without destroying the evidence needed for diagnosis. Determine which users and actions are affected, whether harmful work is continuing, and which reversible control can contain it.
Use a short incident timeline with known facts and hypotheses separated. A queue backlog might result from provider throttling, a poisoned job, lost worker capacity, or increased arrivals. Restarting every process may briefly change the symptom while erasing useful state. A targeted mitigation, such as pausing new expensive jobs or reducing concurrency, can protect users while preserving the investigation.
Communicate the actual user effect. A report delay differs from a wrong report. An unknown payment outcome differs from a confirmed rejection. Avoid saying all data is safe unless you have verified the relevant boundary. State what is known, what is being checked, and when the next useful update will be available.
Recovery needs a condition, not just a service status. A worker process being healthy does not prove the backlog has drained or that users can retrieve correct results. Confirm a representative end-to-end path, inspect outstanding work, and account for duplicate or partial effects created during the incident. Then decide which follow-up removes the cause or improves detection.
Worked example
A fictional export service receives 4 jobs per second, but the provider begins throttling and completion falls to 1. Backlog grows by 3 per second. The operator pauses nonessential bulk submissions and lowers provider concurrency to the documented limit. Existing jobs retain their identities. A small valid export completes and its content matches the fixture.
The incident record says valid exports were delayed, no evidence yet of incorrect content, and 900 jobs remain. With restored completion at 5 per second and admitted arrivals at 2, the remaining backlog should drain in about 300 seconds. The operator measures actual oldest-job age instead of treating the estimate as proof.
Write facts before a cause label
An incident note should support action even while the cause is uncertain. In the export case, reduced completion and growing age are observations. Provider throttling is a hypothesis until request outcomes or documented limits support it. Keep alternatives visible long enough to choose a discriminating check.
Time
Observation
Interpretation or next check
10:00
Arrival 4/s, completion 1/s
Backlog growth about 3/s
10:02
Many provider 429 responses
Check concurrency and retry behavior
10:03
Worker processes healthy
Process liveness is not enough
10:05
One account has repeated large jobs
Inspect fairness and per-account demand
10:07
Admission reduced to 2/s
Measure completion and oldest age after mitigation
This is an invented timeline. The backlog arithmetic is derived from the supplied rates under a constant-rate approximation. Actual arrivals, job sizes, and service times vary, so the estimate guides observation rather than replacing it.
A useful first question is whether users are waiting for correct results or receiving wrong results. Wrong output can require stopping delivery immediately even when throughput looks healthy. Delay can permit a drain strategy with clear status. Treating both as a generic degraded service label hides different harms.
Separate containment from repair
Containment can pause new expensive jobs or limit one noisy account. Repair addresses the cause, such as nested retries exceeding the provider's intended concurrency. A temporary limit may be correct while investigation continues, but it is not evidence that the defect has been removed.
Use one retry owner for a given dependency call chain. If the SDK allows three total attempts per invocation, including the first call, and the outer loop allows three total invocations, full exhaustion can produce 3×3=9 provider calls. These limits mean two retries at each layer. If both instead allow three retries after the first attempt, the bound is 4×4=16 calls. Count total attempts, identify retryable errors, and inspect the implementation before applying the multiplication to a real system.
Preserve operation identity while changing execution policy. Replaying a failed export with a new logical key can create duplicate results or charges when the old outcome is unknown. Classify terminal rejection separately from uncertain completion. A provider timeout does not establish that no paid work occurred.
Make recovery observable
code
1Recovery record2admittedArrivalRate:2 jobs/s3sustainedCompletionRate:5 jobs/s4startingBacklog:900 jobs5modeledDrainTime:900/(5-2)=300s6requiredChecks:7 oldest pending age stops increasing and trends down8 known valid fixture content matches9 no isolated account remains stuck10 unknown outcomes reconciled under operation contract
The 300-second estimate assumes rates remain stable and completion serves the relevant backlog. A high aggregate completion rate can hide one poison job or an unfair queue that starves a small customer. Inspect age and outcomes by relevant segment, not only total depth.
A successful new test job does not prove old jobs recovered. It might bypass the bad input or a stuck partition. Include at least one representative affected job, using a safe fixture or authorized recovery record, and confirm its actual retrievable content. Close the incident only when the agreed user recovery condition is met.
Choose follow-up by causal value
A useful follow-up has a failure mode, owner, completion check, and reason. For nested retries, the change could make one layer own retries and add an attempt-budget assertion. For undetected starvation, add oldest-age visibility by account or work class. For an unknown outcome, improve durable identity and reconciliation.
Avoid a list of unrelated infrastructure upgrades. More replicas do not necessarily change provider capacity. A new dashboard does not prevent an unbounded retry loop unless someone acts before the harmful boundary. Separate preventative controls from diagnostic improvements so their value is clear.
Misconceptions and a second exercise
One misconception is that a green process means users have recovered. It proves only a limited health condition. Another is that the first plausible cause justifies an irreversible cleanup. Deleting queued work before identifying its state can destroy recoverable user intent and evidence.
Exercise: total backlog falls, but one account's oldest job remains at two hours while new work completes. Can the incident close for every user? No. Inspect routing, fairness, input classification, and that job's durable state. Report partial recovery accurately. Award one point for recognizing aggregate masking, one for a discriminating account-level observation, one for preserving job identity, and one for a specific closure condition.
The final note should let another operator continue: current user effect, containment in place, evidence collected, unresolved jobs, and next verification. It should not require reconstructing the whole incident from chat or a process restart history.
Exercise and solution
The service process is healthy but oldest pending age continues rising. Should the incident close? No. User recovery has not occurred. Check arrival and completion rates, downstream failures, and whether a subset of jobs is stuck. Award one point for keeping the incident open, one for a discriminating observation, and one for an end-to-end recovery condition.
Interview probe and wrap-up
How do you choose between rollback and reducing load? A strong answer connects the symptom to the recent change, compatibility limits, and the harm of each action. Follow up with an irreversible data migration. A weak answer says restart is always the fastest fix. A useful incident response leaves users safer, the remaining uncertainty smaller, and the next operator able to continue from written facts.
An SDK permits three total attempts per invocation, including the first. The outer layer permits three total invocations, including the first. Every invocation exhausts its attempts. What is the maximum provider-call count?
Total backlog falls while one account stays stuck. What follows?
AThe account must have invalid data.BThe queue must be empty.CAggregate recovery does not prove that account recovered; inspect its routing and durable state.DAll users recovered.
Which follow-up prevents the observed nested-retry amplification most directly?
ADefine one retry owner and enforce an attempt budget at the paid boundary.BAdd more dashboard panels only.CRaise every timeout without counting attempts.DCreate new operation keys on each retry.
Can you separate containment from cause repair and prove recovery for affected users rather than only healthy processes? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.