Present an incident review that changes the next response
Build a causal incident account with measurable prevention.
Mechanism and reasoning
An incident review should explain how the system produced the outcome and what will change. It is not a list of commands or a search for a person to blame. Start with user impact, timeline and the evidence that supports each causal link. Keep unknowns visible.
Distinguish trigger, contributing condition and failed protection. A release can trigger an incident because a readiness check is too weak and spare capacity is missing. Reverting the release may restore service, but the prevention should also address why the rollout admitted bad capacity or lacked a safe stop.
Use counterfactual reasoning carefully. Ask whether the proposed protection would have interrupted the actual failure path. A generic action such as 'improve monitoring' is not sufficient. Name the signal, threshold, owner and expected response. A test should demonstrate that the new rule catches the supplied trace without excessive false alarms.
Separate immediate repair from durable work. The service can recover before every prevention is implemented. Track each follow-up with an owner and verification condition. Avoid making the incident remain vaguely open until an endless list of improvements is complete.
The interview answer should be brief enough to follow and concrete enough to challenge. Explain one mistake you would own in the hypothetical response, what evidence changed your view and how the revised process would behave. Do not invent personal production experience. Use the teaching case as a simulation unless describing real work you actually performed.
Incident evidence packet
Teaching service has four replicas and a 99.9% request-success objective. At 11:00 a rollout begins. The new image starts its HTTP listener before loading a required index. Readiness checks only that the listener responds. At 11:02 the new Pod becomes ready and an old healthy Pod drains. At 11:03 another old Pod fails independently. Requests sent to the new Pod return application errors until index loading finishes at 11:08.
Time
Observation
Evidence type
11:00
Release R2 begins
Deployment audit
11:02
New Pod marked ready
Probe and condition record
11:03
Old Pod fails
Container termination record
11:03–11:08
Error rate rises
Request metrics by version
11:08
Index completes and errors fall
Application event plus metrics
The trigger is the rollout exposing a partially initialized instance. A contributing condition is the concurrent loss of another healthy replica. The failed protection is readiness that does not represent the ability to serve required requests. A second contributing condition may be insufficient capacity reserve during rollout, depending on the measured workload. State that inference separately from the observed readiness defect.
Do not call index-loading duration the root cause by itself. Slow initialization can be legitimate. The system should avoid routing required work until it is ready, or it should support a declared degraded mode. The question is why traffic reached an instance that could not fulfill the contract.
Prevention test artifact
Proposed change
Test input
Expected result
Readiness checks required index state
Listener up, index absent
Pod remains unready
Startup handling allows measured initialization
Index takes 90seconds
No premature liveness restart
Rollout reserve includes one extra failure
One old Pod fails during canary
Required ready capacity remains
Alert uses version-specific request errors
R2 fails while R1 is healthy
Actionable release signal appears
Each row can be tested in a disposable environment. A check that always returns unready would pass the first negative case but fail the positive case, so also test index-loaded readiness. The prevention test must include both expected denial and successful service. For the rollout reserve, use the same workload that established per-replica capacity.
The review should quantify impact from eligible requests, not just incident minutes. If the five-minute interval has ten thousand requests and eight hundred fail, the incident error fraction is eight percent for that interval. That does not directly equal the thirty-day SLO result; the rolling window has its own denominator. Record affected routes and whether retries caused repeated operations.
Response reflection
Suppose the operator initially restarted every replica because they suspected memory pressure. This made recovery slower by restarting index loads. The review can state that the action was based on an untested hypothesis and lacked a predicted recovery effect. A better next response checks version-specific errors and readiness/index state before broad restarts, unless immediate evidence demands a different mitigation.
Avoid hindsight certainty. During the incident, the operator may not yet have had the index event or per-version metric. That absence is part of the system problem. Add the evidence needed for the next response and make it discoverable in the runbook. The goal is a response that works under the information available at the time.
Close with verification, not an action count. Four completed tickets do not prove the failure path is blocked. Replay the teaching trace and show that the new Pod remains out of normal traffic until ready, the old capacity remains sufficient and the alert points to the failing version. This is a checkable improvement rather than a promise to be more careful.
Worked example
Teaching impact calculation: over five minutes, 10,000 eligible requests occur and 800 fail, so the interval success rate is 92%. If the thirty-day window contains one million eligible requests and the total bad count after this incident is 1,200, window success is 99.88%, below a 99.9% objective. The incident added 800 failures but the prior 400 also count. Keep interval impact and rolling-window compliance as separate figures.
Exercise
A review proposes only 'add more monitoring' after an index-readiness failure. Rewrite it as one measurable prevention and one verification. Then name a separate capacity question.
Model solution and rubric
Add a readiness condition that stays false until the required index is loaded and validated. Test listener-up/index-absent and index-ready cases, then replay a rollout to verify traffic stays on useful replicas. Separately calculate whether the service meets demand after one additional replica fails during rollout. Monitoring alone does not prevent routing to an unusable instance, and readiness alone does not create spare capacity.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
Misconception 1: the last change is the complete cause. The trigger can combine with weak readiness and missing reserve. Misconception 2: more alerts prove prevention. A useful change must interrupt or expose the actual failure path and pass a concrete replay test.
Interview probe
Evidence class: recommended. Original practice.
What would you include in a five-minute incident review explanation?
Strong answer: User impact, the observed timeline, trigger and failed protections, the mitigation and its evidence, then one or two verified changes. I would separate what the trace proves from hypotheses and avoid inventing experience.
Follow-up: How would you test that the proposed alert is useful rather than simply noisy?
Weak answer indicators: Blaming one person; listing commands without causal links; vague monitoring tasks; mixing incident and SLO denominators.
Sources
Technical references: Kubernetes probes; Kubernetes Deployments; Google SRE alerting on SLOs. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
New Pods receive traffic before their index loads. An old node also fails during that interval. Which causal account is best supported?
AThe rollout alone explains every failure without capacity evidenceBPremature readiness exposed incomplete Pods; the node loss may add a capacity contribution that needs separate evidenceCThe node failure proves the new Pods were ready enoughDThe temporal sequence proves that the new image caused the old node to fail
Which test most directly checks the proposed index-readiness repair?
AAssert unready before a valid index is available, then ready after a validated load, and exercise a representative indexed requestBAssert the TCP listener opens within the startup deadlineCAssert desired replicas equal the rollout targetDAssert the new image reaches Ready once on an already-warm cache
Why record the premature-readiness defect separately from the simultaneous old-node loss?
ATo choose the earliest timestamp as the only root causeBTo infer that fixing either condition guarantees all future availabilityCTo assign each failed request equally to both conditions without more dataDTo test each contribution and identify which protections interrupt the observed failure path
The team closes the readiness and capacity tickets. Which evidence best supports the claim that these protections address the incident?
AOne quiet hour without a page under low trafficBA failure replay with cold index load and the stated node loss, plus explicit readiness and capacity assertionsCBoth code reviews have approval and the rollout succeededDThe readiness dashboard has a new panel showing desired replicas
Explain how you would build a causal incident account with measurable prevention without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Explain the failure path and verify a change that blocks it. Keep impact, evidence and remaining uncertainty distinct.