Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

The Runbook Nobody Can Find During an Actual Incident

Most companies have incident runbooks somewhere, and most of those runbooks go unopened during the incident they were written for, because nobody could find them fast enough, or because what's written doesn't match what's actually happening.

This is how to write and maintain runbooks that survive contact with a real 2 a.m. page.

Link the runbook from the alert itself

A runbook buried three clicks deep in a wiki, findable only if you already remember its title, is a runbook that won't get read during a real incident when every minute matters. Link directly from the alert that fires, in the page or Slack message itself, to the specific runbook for that specific failure.

This single change, closing the gap between 'an alert fired' and 'here's what to do about it,' does more for actual incident response than almost any amount of additional runbook content. An unlinked runbook, however well written, is functionally invisible during the moment it's needed.

Write for someone who's never seen this failure before

The person on call during a specific incident often isn't the engineer who wrote the service or the runbook. Writing steps that assume deep familiarity, 'check the usual suspect,' 'restart it the normal way,' fails exactly the person most likely to need the document: someone unfamiliar with this specific failure mode, paged at an inconvenient hour.

Spell out the specific command, the specific dashboard, the specific log query. A runbook that only works for the person who wrote it isn't really a runbook, it's a note to future-them.

Test the runbook the same way you'd test a rollback

A runbook that's never been followed by someone other than its author, under conditions resembling a real incident, has an unknown number of gaps: a command that no longer works, a dashboard link that's moved, a step that assumes access the on-call engineer doesn't actually have. Game days, deliberately simulating a specific failure and having someone unfamiliar with the fix follow the runbook cold, surface these gaps safely.

This is worth treating as seriously as testing a rollback path. An incident is the wrong time to discover the runbook has a broken link or references a tool that was decommissioned last quarter.

A worked example: a runbook that assumed access nobody had

Say a database failover runbook, written by a senior engineer with admin access to the production console, tells the on-call engineer to run a specific command from that console. The engineer paged for a real incident at 3 a.m. doesn't have that access level, and the incident stretches on while access gets escalated through an on-call manager.

A game day exercise, run by someone with standard on-call permissions rather than the runbook's author, would have caught this gap during a calm afternoon instead of during a real outage. Access requirements are one of the most common, and most avoidable, runbook gaps.

Where runbooks fail when it actually matters

  • Not linked from the alert, so nobody finds it fast enough during the incident
  • Written assuming access or context only the original author actually has
  • Never tested by anyone other than the person who wrote it
  • Referencing a tool, dashboard, or command that's since changed or been decommissioned

Update the runbook as part of closing out the incident, not later

The gap between what the runbook said and what actually happened is freshest and most useful immediately after the incident, in the retrospective, not weeks later when the details have blurred. Make updating the relevant runbook a required step in incident closeout, not an optional follow-up that competes with the next fire for attention.

A runbook that only gets updated when someone happens to remember drifts from reality within a few incidents. Tying the update to the closeout process, every time, is what keeps it accurate for the next person who needs it.

Keep the runbook short enough to actually read under pressure

A comprehensive runbook that tries to cover every possible variation of a failure often ends up too long to actually use in the moment, forcing a stressed on-call engineer to skim for the relevant part instead of following clear steps. A short runbook covering the common case well, with a clear escalation path for anything that deviates from it, beats an exhaustive one that takes ten minutes just to find the right section.

If a runbook keeps growing to cover every edge case ever encountered, consider splitting it: a lean primary path for the common failure, with links to separate, shorter documents for known variations, rather than one document trying to be everything at once.

Executive Capability Standard

What Good Looks Like

A runbook that actually gets used during an incident is linked directly from the alert, written for someone unfamiliar with the specific failure, and tested by someone other than its author before being trusted.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Pick your three most frequent alert types and check whether each one links directly to a current, specific runbook.
2. Do Manually:Walk through one runbook by hand as if you'd never seen the failure before, and note every place it assumes context you don't actually have.
3. Delegate:Give each service owner responsibility for their runbooks' accuracy, with an update required as part of every incident closeout.
4. Automate:Wire runbook links directly into your alerting configuration so every page includes a link to its specific response steps.
5. Buy:Bring in an incident response specialist to run structured game days once your team has grown past what informal testing can reliably cover.

How to Get Started

Frequently Asked Questions

How often should runbooks be tested with a game day exercise?

Quarterly for your highest-severity, most likely failure modes is a reasonable cadence for most teams. Less frequently for rare failure modes, but every runbook should be tested at least once by someone other than its author before you trust it during a real incident.

Who should own keeping a runbook current?

The team that owns the service it covers, with an explicit step in incident closeout to update it. Ownership without that trigger tends to mean the runbook gets written once and drifts, since updating it competes with every other priority once the incident is over.

What's the minimum a runbook needs to include to be useful?

The specific symptoms that indicate this failure, the specific diagnostic steps with exact commands or dashboard links, the specific fix, and who to escalate to if the fix doesn't work. Vague guidance is worse than a short, specific runbook covering fewer scenarios well.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides