Writing an Incident Response Runbook People Actually Follow at 3 A.M.
A runbook written calmly, in detail, during business hours, often fails the one test that actually matters: whether someone half asleep at three in the morning, stressed and reading fast, can follow it without getting confused. That's a different writing problem than most documentation, and it needs a different process to get right.
The test for every runbook step isn't "is this accurate," it's "can someone who isn't thinking clearly execute this exact step without having to stop and figure out what it means."
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How should an incident runbook start?
Every runbook needs a clear answer to "how do we know this is happening" before it gets to "what do we do about it": the specific alert, dashboard, or signal that triggers this runbook, and what it looks like when the situation has actually resolved. Without a clear detection and resolution signal, an on-call engineer can execute every response step correctly and still not know whether the incident is actually over.
Write the detection criteria specifically enough that two different people looking at the same dashboard would agree on whether this runbook applies right now, not a vague description that requires judgment to interpret under pressure.
How do you write runbook steps people can follow?
"Restart the affected service" is a summary that requires the reader to remember or look up the actual command, which is exactly the kind of extra step that goes wrong under stress. "Run this exact command against this exact service" is a literal instruction that removes a decision point entirely. Every step in the response section should be copy-paste executable, not a description of an action that still requires expertise to translate into something you can actually run.
At 99.999% uptime, the whole year's downtime allowance is about 5.26 minutes1, which is exactly why the runbook has to work on the first read during an incident, not the third attempt after two failed guesses at the right command.
Include the decision points, not just the happy path
A real incident rarely follows the runbook's first branch cleanly; the primary fix doesn't work, or the situation looks slightly different than the scenario the runbook describes. Include explicit "if this doesn't work, try this" branches for the most likely deviations, and a clear point at which the runbook says to escalate rather than keep trying variations alone.
A runbook that only covers the ideal path leaves the on-call engineer improvising exactly when they're least equipped to improvise well, tired, under pressure, and aware that people are waiting on a fix.
Store it somewhere it actually gets found during an incident
A runbook in a wiki page twelve clicks deep, or in a repository the on-call engineer doesn't have open at three in the morning, might as well not exist in the moment it's needed. Host the runbook somewhere your team already checks for procedures, such as an SOP tool like Trainual, and link directly to it from the alert itself, so finding it isn't a separate task competing with actually responding.
Test this specifically: during a drill, time how long it takes someone to actually locate the right runbook, not just how long the response takes once they have it open in front of them.
Update it every time it's actually used
The single best source of runbook improvements is the last time it was actually followed during a real incident: what step was confusing, what assumption turned out to be wrong, what the runbook didn't cover that it should have. Build a short review into every postmortem specifically asking whether the runbook held up, and update it immediately while the specific gap is still fresh, rather than filing it away as a future improvement that never quite happens.
A runbook that's never updated after an incident is a runbook that's slowly drifting out of sync with how the system actually behaves now, one small architecture change at a time.
Assign the update itself to whoever actually ran the incident, while it's still fresh, rather than to the runbook's original author who may not have been in the room. The person who just followed the steps under real pressure knows exactly which line confused them or turned out to be wrong, and that detail is easy to lose within a day or two once the adrenaline wears off.
A runbook that holds up under pressure meets these checks:
- It names the alert or signal that triggers it and describes what the situation looks like once resolved.
- Each response step is a literal, copy-paste command instead of a summary.
- It includes branches for the likely deviations and a clear point at which to escalate.
- It is linked directly from the alert and hosted where the team already looks for procedures.
- It gets reviewed in every postmortem and updated while the specific gap is still fresh.
What Good Looks Like
Incident response is working when an on-call engineer, tired and under pressure, can follow the runbook's literal steps to resolution without needing to interpret or look anything up separately.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How detailed should an incident runbook be?
Detailed enough that a specific literal command replaces every summary, but focused only on the specific scenario it covers. A runbook trying to cover every possible variation becomes too long to read quickly under pressure; write a focused one per scenario instead of one sprawling document.
Who should be responsible for keeping runbooks up to date?
The team that owns the system the runbook covers, with an explicit review step added to every postmortem that touches it. Ownership without a trigger to actually revisit the document tends to mean it just doesn't happen until the next incident reveals it's stale.
Should we run drills against our runbooks even without a real incident?
Yes, a scheduled drill, ideally with someone unfamiliar with the system running it cold, is the most reliable way to find the gaps a runbook's author can't see themselves, since they already know what they meant by every step.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
What Actually Belongs in an Incident Response Runbook
What a useful incident response runbook actually contains: the first five minutes, a named commander, a communication cadence, and a scheduled postmortem.
Writing an Incident Runbook People Actually Follow at 2 A.M.
How to write an incident response runbook that a half-awake, stressed engineer can actually follow, instead of one that only reads well in review.
Writing an Incident Runbook Someone Can Actually Follow at 3 AM
A step by step way to write a streaming pipeline incident runbook that a half-awake on-call engineer can actually follow, not just a policy document.
Writing an Incident Runbook People Will Actually Follow
How to write an incident response runbook engineers actually reach for during a real outage, instead of one that sits unread until the next audit.
An On-Call Runbook for When Retrieval Quality Drops
A concrete triage order for a RAG on-call incident: outage versus quality drop, the three most common causes to check first, and a scoped kill switch.
Writing an Incident Runbook for When Agents Misbehave
How to build an incident response runbook specifically for agent failures, since a misbehaving agent breaks differently than a normal outage.