The Runbook Nobody Can Find During an Actual Incident
Most companies have incident runbooks somewhere, and most of those runbooks go unopened during the incident they were written for, because nobody could find them fast enough, or because what's written doesn't match what's actually happening.
This is how to write and maintain runbooks that survive contact with a real 2 a.m. page.
Link the runbook from the alert itself
A runbook buried three clicks deep in a wiki, findable only if you already remember its title, is a runbook that won't get read during a real incident when every minute matters. Link directly from the alert that fires, in the page or Slack message itself, to the specific runbook for that specific failure.
This single change, closing the gap between 'an alert fired' and 'here's what to do about it,' does more for actual incident response than almost any amount of additional runbook content. An unlinked runbook, however well written, is functionally invisible during the moment it's needed.
Write for someone who's never seen this failure before
The person on call during a specific incident often isn't the engineer who wrote the service or the runbook. Writing steps that assume deep familiarity, 'check the usual suspect,' 'restart it the normal way,' fails exactly the person most likely to need the document: someone unfamiliar with this specific failure mode, paged at an inconvenient hour.
Spell out the specific command, the specific dashboard, the specific log query. A runbook that only works for the person who wrote it isn't really a runbook, it's a note to future-them.
Test the runbook the same way you'd test a rollback
A runbook that's never been followed by someone other than its author, under conditions resembling a real incident, has an unknown number of gaps: a command that no longer works, a dashboard link that's moved, a step that assumes access the on-call engineer doesn't actually have. Game days, deliberately simulating a specific failure and having someone unfamiliar with the fix follow the runbook cold, surface these gaps safely.
This is worth treating as seriously as testing a rollback path. An incident is the wrong time to discover the runbook has a broken link or references a tool that was decommissioned last quarter.
A worked example: a runbook that assumed access nobody had
Say a database failover runbook, written by a senior engineer with admin access to the production console, tells the on-call engineer to run a specific command from that console. The engineer paged for a real incident at 3 a.m. doesn't have that access level, and the incident stretches on while access gets escalated through an on-call manager.
A game day exercise, run by someone with standard on-call permissions rather than the runbook's author, would have caught this gap during a calm afternoon instead of during a real outage. Access requirements are one of the most common, and most avoidable, runbook gaps.
Where runbooks fail when it actually matters
- Not linked from the alert, so nobody finds it fast enough during the incident
- Written assuming access or context only the original author actually has
- Never tested by anyone other than the person who wrote it
- Referencing a tool, dashboard, or command that's since changed or been decommissioned
Update the runbook as part of closing out the incident, not later
The gap between what the runbook said and what actually happened is freshest and most useful immediately after the incident, in the retrospective, not weeks later when the details have blurred. Make updating the relevant runbook a required step in incident closeout, not an optional follow-up that competes with the next fire for attention.
A runbook that only gets updated when someone happens to remember drifts from reality within a few incidents. Tying the update to the closeout process, every time, is what keeps it accurate for the next person who needs it.
Keep the runbook short enough to actually read under pressure
A comprehensive runbook that tries to cover every possible variation of a failure often ends up too long to actually use in the moment, forcing a stressed on-call engineer to skim for the relevant part instead of following clear steps. A short runbook covering the common case well, with a clear escalation path for anything that deviates from it, beats an exhaustive one that takes ten minutes just to find the right section.
If a runbook keeps growing to cover every edge case ever encountered, consider splitting it: a lean primary path for the common failure, with links to separate, shorter documents for known variations, rather than one document trying to be everything at once.
What Good Looks Like
A runbook that actually gets used during an incident is linked directly from the alert, written for someone unfamiliar with the specific failure, and tested by someone other than its author before being trusted.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should runbooks be tested with a game day exercise?
Quarterly for your highest-severity, most likely failure modes is a reasonable cadence for most teams. Less frequently for rare failure modes, but every runbook should be tested at least once by someone other than its author before you trust it during a real incident.
Who should own keeping a runbook current?
The team that owns the service it covers, with an explicit step in incident closeout to update it. Ownership without that trigger tends to mean the runbook gets written once and drifts, since updating it competes with every other priority once the incident is over.
What's the minimum a runbook needs to include to be useful?
The specific symptoms that indicate this failure, the specific diagnostic steps with exact commands or dashboard links, the specific fix, and who to escalate to if the fix doesn't work. Vague guidance is worse than a short, specific runbook covering fewer scenarios well.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Writing an Incident Response Runbook People Actually Follow at 3 A.M.
A worksheet approach to writing incident runbooks that hold up under real pressure, when the person on call is tired, stressed, and reading fast.
A Production Deployment Checklist That Actually Catches Problems
A stage-by-stage deployment checklist for distributed systems, covering rollback readiness, dependency ordering, and the checks teams skip under pressure.
Writing an Incident Runbook People Actually Follow at 2 A.M.
How to write an incident response runbook that a half-awake, stressed engineer can actually follow, instead of one that only reads well in review.
Verifying Devices Before They Touch Production, Not After
How to build device verification into a zero-trust rollout, what actually counts as a trust signal, and where teams stop checking too early.
Finding the Real Source of Latency in a Distributed System
A decision guide for narrowing down whether a slow request is a network problem, a database problem, a queue problem, or your own code.
What Actually Belongs in an Incident Response Runbook
What a useful incident response runbook actually contains: the first five minutes, a named commander, a communication cadence, and a scheduled postmortem.