Lesson 3 of 4 · 60 min

Operate the incident before writing the postmortem

Produce a short incident record with mitigation, evidence, and a verified recovery condition.

A founding engineer may be the only person available when a workflow fails. The first task is to reduce user harm without destroying the evidence needed for diagnosis. Determine which users and actions are affected, whether harmful work is continuing, and which reversible control can contain it.
Use a short incident timeline with known facts and hypotheses separated. A queue backlog might result from provider throttling, a poisoned job, lost worker capacity, or increased arrivals. Restarting every process may briefly change the symptom while erasing useful state. A targeted mitigation, such as pausing new expensive jobs or reducing concurrency, can protect users while preserving the investigation.
Communicate the actual user effect. A report delay differs from a wrong report. An unknown payment outcome differs from a confirmed rejection. Avoid saying all data is safe unless you have verified the relevant boundary. State what is known, what is being checked, and when the next useful update will be available.
Recovery needs a condition, not just a service status. A worker process being healthy does not prove the backlog has drained or that users can retrieve correct results. Confirm a representative end-to-end path, inspect outstanding work, and account for duplicate or partial effects created during the incident. Then decide which follow-up removes the cause or improves detection.

Worked example

A fictional export service receives 4 jobs per second, but the provider begins throttling and completion falls to 1. Backlog grows by 3 per second. The operator pauses nonessential bulk submissions and lowers provider concurrency to the documented limit. Existing jobs retain their identities. A small valid export completes and its content matches the fixture.
The incident record says valid exports were delayed, no evidence yet of incorrect content, and 900 jobs remain. With restored completion at 5 per second and admitted arrivals at 2, the remaining backlog should drain in about 300 seconds. The operator measures actual oldest-job age instead of treating the estimate as proof.

Write facts before a cause label

An incident note should support action even while the cause is uncertain. In the export case, reduced completion and growing age are observations. Provider throttling is a hypothesis until request outcomes or documented limits support it. Keep alternatives visible long enough to choose a discriminating check.
TimeObservationInterpretation or next check
10:00Arrival 4/s, completion 1/sBacklog growth about 3/s
10:02Many provider 429 responsesCheck concurrency and retry behavior
10:03Worker processes healthyProcess liveness is not enough
10:05One account has repeated large jobsInspect fairness and per-account demand
10:07Admission reduced to 2/sMeasure completion and oldest age after mitigation
This is an invented timeline. The backlog arithmetic is derived from the supplied rates under a constant-rate approximation. Actual arrivals, job sizes, and service times vary, so the estimate guides observation rather than replacing it.
A useful first question is whether users are waiting for correct results or receiving wrong results. Wrong output can require stopping delivery immediately even when throughput looks healthy. Delay can permit a drain strategy with clear status. Treating both as a generic degraded service label hides different harms.

Separate containment from repair

Containment can pause new expensive jobs or limit one noisy account. Repair addresses the cause, such as nested retries exceeding the provider's intended concurrency. A temporary limit may be correct while investigation continues, but it is not evidence that the defect has been removed.
Use one retry owner for a given dependency call chain. If the SDK allows three total attempts per invocation, including the first call, and the outer loop allows three total invocations, full exhaustion can produce 3×3=9 provider calls. These limits mean two retries at each layer. If both instead allow three retries after the first attempt, the bound is 4×4=16 calls. Count total attempts, identify retryable errors, and inspect the implementation before applying the multiplication to a real system.
Preserve operation identity while changing execution policy. Replaying a failed export with a new logical key can create duplicate results or charges when the old outcome is unknown. Classify terminal rejection separately from uncertain completion. A provider timeout does not establish that no paid work occurred.

Make recovery observable

code
1Recovery record2admittedArrivalRate:2 jobs/s3sustainedCompletionRate:5 jobs/s4startingBacklog:900 jobs5modeledDrainTime:900/(5-2)=300s6requiredChecks:7  oldest pending age stops increasing and trends down8  known valid fixture content matches9  no isolated account remains stuck10  unknown outcomes reconciled under operation contract
The 300-second estimate assumes rates remain stable and completion serves the relevant backlog. A high aggregate completion rate can hide one poison job or an unfair queue that starves a small customer. Inspect age and outcomes by relevant segment, not only total depth.
A successful new test job does not prove old jobs recovered. It might bypass the bad input or a stuck partition. Include at least one representative affected job, using a safe fixture or authorized recovery record, and confirm its actual retrievable content. Close the incident only when the agreed user recovery condition is met.

Choose follow-up by causal value

A useful follow-up has a failure mode, owner, completion check, and reason. For nested retries, the change could make one layer own retries and add an attempt-budget assertion. For undetected starvation, add oldest-age visibility by account or work class. For an unknown outcome, improve durable identity and reconciliation.
Avoid a list of unrelated infrastructure upgrades. More replicas do not necessarily change provider capacity. A new dashboard does not prevent an unbounded retry loop unless someone acts before the harmful boundary. Separate preventative controls from diagnostic improvements so their value is clear.

Misconceptions and a second exercise

One misconception is that a green process means users have recovered. It proves only a limited health condition. Another is that the first plausible cause justifies an irreversible cleanup. Deleting queued work before identifying its state can destroy recoverable user intent and evidence.
Exercise: total backlog falls, but one account's oldest job remains at two hours while new work completes. Can the incident close for every user? No. Inspect routing, fairness, input classification, and that job's durable state. Report partial recovery accurately. Award one point for recognizing aggregate masking, one for a discriminating account-level observation, one for preserving job identity, and one for a specific closure condition.
The final note should let another operator continue: current user effect, containment in place, evidence collected, unresolved jobs, and next verification. It should not require reconstructing the whole incident from chat or a process restart history.

Exercise and solution

The service process is healthy but oldest pending age continues rising. Should the incident close? No. User recovery has not occurred. Check arrival and completion rates, downstream failures, and whether a subset of jobs is stuck. Award one point for keeping the incident open, one for a discriminating observation, and one for an end-to-end recovery condition.

Interview probe and wrap-up

How do you choose between rollback and reducing load? A strong answer connects the symptom to the recent change, compatibility limits, and the harm of each action. Follow up with an irreversible data migration. A weak answer says restart is always the fastest fix. A useful incident response leaves users safer, the remaining uncertainty smaller, and the next operator able to continue from written facts.

Sources

docsGoogle SRE service-level objectivessre.googledocsAWS Builders Library: avoiding queue backlogsd1.awsstatic.comdocsStripe idempotent request contractdocs.stripe.comdocsAWS transactional outbox patterndocs.aws.amazon.com

Checkpoint

Arrival 4/s and completion 1/s remain constant. Backlog growth?

AThree jobs/s.BFour jobs/s.CFive jobs/s.DOne job/s.
Sign up free to answer and see why

Checkpoint

An SDK permits three total attempts per invocation, including the first. The outer layer permits three total invocations, including the first. Every invocation exhausts its attempts. What is the maximum provider-call count?

ASixteen.BThree.CSix.DNine.
Sign up free to answer and see why

Checkpoint

Total backlog falls while one account stays stuck. What follows?

AThe account must have invalid data.BThe queue must be empty.CAggregate recovery does not prove that account recovered; inspect its routing and durable state.DAll users recovered.
Sign up free to answer and see why

Checkpoint

Which follow-up prevents the observed nested-retry amplification most directly?

ADefine one retry owner and enforce an attempt budget at the paid boundary.BAdd more dashboard panels only.CRaise every timeout without counting attempts.DCreate new operation keys on each retry.
Sign up free to answer and see why

Checkpoint

A900 job backlog drains with completion 5/s and admitted arrivals 2/s. Modeled time?

A180 seconds.B300 seconds.C450 seconds.D900 seconds.
Sign up free to answer and see why

Can you separate containment from cause repair and prove recovery for affected users rather than only healthy processes? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.