Lesson 08 of 10 · Part 3 — People and outcomes
Leading a major incident and the review after it
Lead a major incident and the learning after it: take the incident commander role, stabilise before diagnosing, communicate on a rhythm, hand off cleanly, then run a blameless review that finds contributing factors and produces a few actions that actually get done.
During the incident
| Role | Does |
|---|---|
| Incident commander | Owns the incident: priorities, decisions, roles, timeline. Does not debug |
| Operations lead(s) | Investigate and change systems, one set of hands per system |
| Communications | Status updates to stakeholders and users on a fixed rhythm |
| Scribe | Timeline of observations, decisions and actions |
In a small team one person may hold two roles; in a big incident, keeping commander and debugger separate is what keeps it under control.
In a house fire, the fire chief doesn't hold a hose. They decide where the hoses go, keep everyone safe, and tell the neighbours what's happening. Someone has to see the whole picture, and you can't do that with your hands full.
The commander's loop
- Declare the incident and its severity early; it's cheaper to downgrade than to start late.
- Stabilise first: roll back the last change, fail over, scale up, shed load. Find the root cause after users are safe.
- Assign work explicitly ("Asha: check the last deploy; Ben: DNS"), with times to report back.
- Communicate on a rhythm (every 15–30 minutes): impact, what we know, what we're doing, next update time. Even "no change" is an update.
- Decide under uncertainty: pick the safest reversible action, announce it, watch the effect.
- Hand off cleanly if it runs long: written state, owner of each thread, who is now commander.
- Close when stable, with monitoring in place, and schedule the review.
After the incident: the blameless review
Within a few days, while memories are fresh:
- Timeline: what happened, what people saw, what they thought and did, and why it made sense at the time.
- Contributing factors: not one "root cause" but the conditions that lined up (a risky default, a missing alert, an unclear runbook, a deploy during a peak).
- What went well: so it stays that way.
- Actions: few, specific, owned and dated. Prefer actions that change the system (a guard-rail, an automated check, a smaller blast radius) over "be more careful".
Blameless doesn't mean nobody is accountable: people are accountable for following up, not punished for honest mistakes.
Make the actions happen
Track review actions with the same priority as roadmap work, review their status weekly, and report completion rate. A review whose actions are never done teaches the team that reviews are theatre.
Across many incidents
Look for patterns every quarter: repeated contributing factors (change management, capacity, dependencies) are roadmap material (lesson 05).
Try it: a game-day incident
- In a test environment, break something realistic (a bad config rollout, a full disk, an expired certificate).
- Run it as a real incident with assigned roles and 10-minute status updates.
- Afterwards, write a one-page review: timeline, contributing factors, what went well, three actions.
- Check each action: does it change the system, or only ask people to be careful?
- Track the actions to completion and report back to the team.
Going deeper: incident culture
Thank people who raise incidents early, even false alarms. Share reviews widely (redacted if needed) so other teams learn. Measure time to declare, time to mitigate and action completion, not the number of incidents alone.
Recap
- Clear roles: commander coordinates, operators fix, someone communicates, someone writes the timeline.
- Stabilise first, communicate on a rhythm, decide with reversible actions, hand off cleanly.
- Blameless reviews find contributing factors; actions change the system.
- Track actions to completion; look for patterns across incidents.
This site is a public version of my personal engineering knowledge hub. It intentionally excludes confidential company information and internal operational details.