Lesson 4 of 4 · 25 min

Present an incident review that changes the next response

Build a causal incident account with measurable prevention.

Mechanism and reasoning

An incident review should explain how the system produced the outcome and what will change. It is not a list of commands or a search for a person to blame. Start with user impact, timeline and the evidence that supports each causal link. Keep unknowns visible.
Distinguish trigger, contributing condition and failed protection. A release can trigger an incident because a readiness check is too weak and spare capacity is missing. Reverting the release may restore service, but the prevention should also address why the rollout admitted bad capacity or lacked a safe stop.
Use counterfactual reasoning carefully. Ask whether the proposed protection would have interrupted the actual failure path. A generic action such as 'improve monitoring' is not sufficient. Name the signal, threshold, owner and expected response. A test should demonstrate that the new rule catches the supplied trace without excessive false alarms.
Separate immediate repair from durable work. The service can recover before every prevention is implemented. Track each follow-up with an owner and verification condition. Avoid making the incident remain vaguely open until an endless list of improvements is complete.
The interview answer should be brief enough to follow and concrete enough to challenge. Explain one mistake you would own in the hypothetical response, what evidence changed your view and how the revised process would behave. Do not invent personal production experience. Use the teaching case as a simulation unless describing real work you actually performed.

Incident evidence packet

Teaching service has four replicas and a 99.9% request-success objective. At 11:00 a rollout begins. The new image starts its HTTP listener before loading a required index. Readiness checks only that the listener responds. At 11:02 the new Pod becomes ready and an old healthy Pod drains. At 11:03 another old Pod fails independently. Requests sent to the new Pod return application errors until index loading finishes at 11:08.
TimeObservationEvidence type
11:00Release R2 beginsDeployment audit
11:02New Pod marked readyProbe and condition record
11:03Old Pod failsContainer termination record
11:03–11:08Error rate risesRequest metrics by version
11:08Index completes and errors fallApplication event plus metrics
The trigger is the rollout exposing a partially initialized instance. A contributing condition is the concurrent loss of another healthy replica. The failed protection is readiness that does not represent the ability to serve required requests. A second contributing condition may be insufficient capacity reserve during rollout, depending on the measured workload. State that inference separately from the observed readiness defect.
Do not call index-loading duration the root cause by itself. Slow initialization can be legitimate. The system should avoid routing required work until it is ready, or it should support a declared degraded mode. The question is why traffic reached an instance that could not fulfill the contract.

Prevention test artifact

Proposed changeTest inputExpected result
Readiness checks required index stateListener up, index absentPod remains unready
Startup handling allows measured initializationIndex takes 90secondsNo premature liveness restart
Rollout reserve includes one extra failureOne old Pod fails during canaryRequired ready capacity remains
Alert uses version-specific request errorsR2 fails while R1 is healthyActionable release signal appears
Each row can be tested in a disposable environment. A check that always returns unready would pass the first negative case but fail the positive case, so also test index-loaded readiness. The prevention test must include both expected denial and successful service. For the rollout reserve, use the same workload that established per-replica capacity.
The review should quantify impact from eligible requests, not just incident minutes. If the five-minute interval has ten thousand requests and eight hundred fail, the incident error fraction is eight percent for that interval. That does not directly equal the thirty-day SLO result; the rolling window has its own denominator. Record affected routes and whether retries caused repeated operations.

Response reflection

Suppose the operator initially restarted every replica because they suspected memory pressure. This made recovery slower by restarting index loads. The review can state that the action was based on an untested hypothesis and lacked a predicted recovery effect. A better next response checks version-specific errors and readiness/index state before broad restarts, unless immediate evidence demands a different mitigation.
Avoid hindsight certainty. During the incident, the operator may not yet have had the index event or per-version metric. That absence is part of the system problem. Add the evidence needed for the next response and make it discoverable in the runbook. The goal is a response that works under the information available at the time.
Close with verification, not an action count. Four completed tickets do not prove the failure path is blocked. Replay the teaching trace and show that the new Pod remains out of normal traffic until ready, the old capacity remains sufficient and the alert points to the failing version. This is a checkable improvement rather than a promise to be more careful.

Worked example

Teaching impact calculation: over five minutes, 10,000 eligible requests occur and 800 fail, so the interval success rate is 92%. If the thirty-day window contains one million eligible requests and the total bad count after this incident is 1,200, window success is 99.88%, below a 99.9% objective. The incident added 800 failures but the prior 400 also count. Keep interval impact and rolling-window compliance as separate figures.

Exercise

A review proposes only 'add more monitoring' after an index-readiness failure. Rewrite it as one measurable prevention and one verification. Then name a separate capacity question.

Model solution and rubric

Add a readiness condition that stays false until the required index is loaded and validated. Test listener-up/index-absent and index-ready cases, then replay a rollout to verify traffic stays on useful replicas. Separately calculate whether the service meets demand after one additional replica fails during rollout. Monitoring alone does not prevent routing to an unusable instance, and readiness alone does not create spare capacity.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

Misconception 1: the last change is the complete cause. The trigger can combine with weak readiness and missing reserve. Misconception 2: more alerts prove prevention. A useful change must interrupt or expose the actual failure path and pass a concrete replay test.

Interview probe

Evidence class: recommended. Original practice.
What would you include in a five-minute incident review explanation?
Strong answer: User impact, the observed timeline, trigger and failed protections, the mitigation and its evidence, then one or two verified changes. I would separate what the trace proves from hypotheses and avoid inventing experience.
Follow-up: How would you test that the proposed alert is useful rather than simply noisy?
Weak answer indicators: Blaming one person; listing commands without causal links; vague monitoring tasks; mixing incident and SLO denominators.

Sources

Technical references: Kubernetes probes; Kubernetes Deployments; Google SRE alerting on SLOs. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsKubernetes probeskubernetes.iodocsKubernetes Deploymentskubernetes.iodocsGoogle SRE alerting on SLOssre.google

Checkpoint

New Pods receive traffic before their index loads. An old node also fails during that interval. Which causal account is best supported?

AThe rollout alone explains every failure without capacity evidenceBPremature readiness exposed incomplete Pods; the node loss may add a capacity contribution that needs separate evidenceCThe node failure proves the new Pods were ready enoughDThe temporal sequence proves that the new image caused the old node to fail
Sign up free to answer and see why

Checkpoint

In a complete interval, 800 of 10,000 eligible requests fail. What is the success rate for that interval?

A99.92%B99.2%C92%D8%
Sign up free to answer and see why

Checkpoint

Which test most directly checks the proposed index-readiness repair?

AAssert unready before a valid index is available, then ready after a validated load, and exercise a representative indexed requestBAssert the TCP listener opens within the startup deadlineCAssert desired replicas equal the rollout targetDAssert the new image reaches Ready once on an already-warm cache
Sign up free to answer and see why

Checkpoint

Why record the premature-readiness defect separately from the simultaneous old-node loss?

ATo choose the earliest timestamp as the only root causeBTo infer that fixing either condition guarantees all future availabilityCTo assign each failed request equally to both conditions without more dataDTo test each contribution and identify which protections interrupt the observed failure path
Sign up free to answer and see why

Checkpoint

The team closes the readiness and capacity tickets. Which evidence best supports the claim that these protections address the incident?

AOne quiet hour without a page under low trafficBA failure replay with cold index load and the stated node loss, plus explicit readiness and capacity assertionsCBoth code reviews have approval and the rollout succeededDThe readiness dashboard has a new panel showing desired replicas
Sign up free to answer and see why

Explain how you would build a causal incident account with measurable prevention without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Explain the failure path and verify a change that blocks it. Keep impact, evidence and remaining uncertainty distinct.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.