Lesson 5 of 8 · 48 min
Failure, incidents, and learning without theater
Postmortem-style failure spine, worked incident story, weak vs strong endings, systems root cause, deadline misses, feedback stories, and prevention artifacts that survive probes.
Failure stories are trust stories
Failure story spine (postmortem format)
1FAILURE SPINE23 1. Context + your role4 2. What broke (user/business impact) with magnitude5 3. Your contribution (honest — not false total blame or zero blame)6 4. Detection (how you found out; how slow)7 5. Mitigation (what you did in the window)8 6. Root cause (systems, not only 'human error')9 7. Prevention (still live today)10 8. Personal change (habit/checklist — operational)1112 Mirror a blameless postmortem voice: systems thinking + personal agency.Key idea
Worked incident story
1PROMPT: Tell me about a production failure you were involved in.23 Context: I owned config for feature flags on checkout.4 Break: bad default flag flipped off payments for 0.8% of users in one region (~12 min).5 Contribution: I merged a flag default change without region soak; assumed global default safe.6 Detection: revenue dashboard + support spike; paged late because no flag-diff alert.7 Mitigation: reverted flag; confirmed region recovery; customer comms with support lead.8 Root cause: defaults coupled to global; no canary on flag defaults; alert gap.9 Prevention: region canary for flag defaults; alert on payment success by region;10 checklist item in flag RFC template (still used).11 Personal: I never ship flag default changes without the canary section filled.Weak vs strong failure answers
1WEAK2 "I'm a perfectionist so my failure is caring too much. Once a project slipped a week3 because requirements changed. I learned to communicate."45ALSO WEAK6 "The PM failed to give requirements and ops was asleep. I did everything right."78STRONG9 "I shipped X; impact Y; my miss was Z; we mitigated in T; root cause; prevention10 artifact A still live; I now personally do B."Common mistake
I should pick a tiny failure so I do not look bad.
Root cause: systems language
1ROOT CAUSE LADDER23 Too shallow: human error / typo4 Better: no test for config default5 Strong: flag defaults global + no region canary + no payment-by-region alert6 Senior: strong + why the process made the wrong default easyKey idea
Deadline misses and non-incident failures
1DEADLINE-MISS SPINE23 Goal + date → early signals you ignored or raised → decision to cut/slip →4 stakeholder comms quality → final outcome → planning change (buffer, milestone,5 risk list) you still use.Critical feedback stories
- 01State the feedback in their words, not your spin.
- 02Own the part that was true even if delivery was rough.
- 03Show a behavior change with an example weeks later.
- 04Avoid 'feedback was wrong but I listened anyway' as the whole story.
- 05Avoid fake failures: 'I work too hard'.
- 06Keep one feedback story and one incident story in the bank.
Integrity edges
Common mistake
Blameless culture means I should not admit personal mistakes.
Probe pack
1FAILURE PROBES23 - What did you personally do wrong (not the system only)?4 - How did customers find out before you did?5 - What would detection look like if it happened tonight?6 - Who was angry and how did you handle that conversation?7 - What still worries you about that class of failure?Company flavor
A good failure story makes the interviewer believe production is safer because you were there — not because you are flawless.
1PREVENTION ARTIFACT IDEAS (from real eng work)23 - Canary + auto rollback on 5xx4 - Alert on regional payment success5 - RFC checklist section6 - Load test in CI for hot path7 - Feature flag soak requirement8 - Postmortem template field for customer detection lagCheckpoint
Which failure ending best signals senior ownership?
Checkpoint
Best way to state your contribution in an incident you partially caused?
Checkpoint
Why is 'my biggest weakness is perfectionism' a bad failure answer?
Checkpoint
Detection lag was 40 minutes and customers tweeted first. How should you handle that in the story?
Checkpoint
You are asked what still worries you about that failure class. Best response shape?
Can you tell one incident failure and one non-incident miss with prevention artifacts still live?
Failure locked
- Failure spine mirrors a blameless postmortem with personal agency.
- Real stake beats safe clichés; fabrication is forbidden.
- Prevention artifacts in production beat 'I'll be careful'.
- Root cause should make the wrong action harder or louder.
- Own detection lag; keep residual risk honesty for probes.
Next: technical storytelling — a system you designed or shipped.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.