Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

Writing an Incident Runbook People Will Actually Follow

Most incident runbooks are written once, during a calm afternoon, and never opened again until the next compliance audit asks for one. During a real incident, people default to whatever they remember instead, which means the runbook exists but isn't actually doing its job.

Here's how to write one people reach for when it matters.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Start with roles, not just steps

Before listing what to do, define who does it: who declares an incident, who leads the response, who communicates externally, and who has the authority to make a risky call under pressure, like deciding to take a system offline. Without defined roles, the first few minutes of a real incident get spent figuring out who's in charge instead of actually responding.

Write these roles as functions, not names, so the runbook still works when the usual person is asleep, on vacation, or already handling something else. Revisit the assignments whenever the team changes, since a role sheet built around people who've since left or moved teams is a gap you don't want to discover mid-incident.

Write steps specific enough to follow under pressure

A step like "investigate the database" is useless to someone whose hands are shaking a little at 2am. A step like "check the replication lag dashboard at this link, and if it's over the expected threshold, follow the failover procedure" can actually be followed by someone who's stressed and not thinking as clearly as usual.

Write every step assuming the reader is tired, stressed, and unfamiliar with the specific system, even if that's rarely true. That's the condition the runbook actually needs to work under, not the condition it was written in.

For example, a vague step such as "check the queue" becomes "open the queue dashboard at the linked address, and if the backlog is growing faster than it drains, follow the scaling procedure below." The rewrite names the tool, the condition and the next action, so a tired responder never has to decide what "check" means. A useful test is to hand a step to someone from a different team and ask what they would do first. If they hesitate or ask a clarifying question, the step needs more detail before it stays in the runbook.

Keep it current by testing it, not just reviewing it

A runbook reviewed on paper can look complete while referencing a dashboard that was renamed months ago or a team member who's since left the company. Actually running through it, even as a tabletop exercise where the team walks through the steps out loud without a real incident happening, surfaces those gaps in a way a read-through alone doesn't.

Schedule this as a recurring exercise, not a one-time setup task, since systems and teams both keep changing after the runbook is written.

Where runbooks fail during a real incident

A handful of gaps show up repeatedly when a runbook is put to a real test:

  • A linked dashboard or tool that's moved or been renamed since the runbook was written
  • No clear criteria for when to escalate versus handle it within the current team
  • Communication templates that don't exist yet, so someone has to write customer-facing language during the incident itself
  • A runbook that assumes deep familiarity with the system, written by the one person who has it and unreadable to anyone else

Each of these turns a runbook that looks complete into one that quietly fails exactly when it's needed.

The moment the runbook needs a downtime number, not a guess

Deciding how aggressively to respond (paging more people, considering a full failover, deciding whether a partial degradation is acceptable to ride out) goes faster with a real number in hand rather than a debate in the moment. If your availability target implies a downtime budget of roughly 8.76 hours a year, only about 0.365 days1, knowing how much of that budget is already used this year, even a rough running total, tells the team how much urgency this specific incident actually deserves.

After the incident: the review that actually updates the runbook

The post-incident review is where most of the runbook's real improvement should come from, but only if it results in actual edits, not just a written summary that gets filed away. Assign a specific person to update the runbook with what was learned before the review is considered closed, not as an optional follow-up that competes with the next week's other priorities. A review that produces no edit to the runbook at all is usually a sign the review itself stayed too general to be actionable.

Executive Capability Standard

What Good Looks Like

Good here means an engineer unfamiliar with a specific system can follow the runbook during a real incident and reach the right escalation and communication decisions, not just find generic troubleshooting steps.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your current runbook as if you'd never seen the system before, and note every step that assumes context you don't actually have written down.
2. Do Manually:Run a tabletop exercise where the team walks through the runbook out loud for a plausible incident scenario.
3. Delegate:Assign one person ownership of keeping the runbook current, including updating it after every real incident review.
4. Automate:Link the runbook directly to live dashboards and alerting so referenced tools can't silently go stale without someone noticing.
5. Buy:Bring in an incident response consultant to help design the process once you're coordinating across enough teams that a single-team runbook no longer covers it.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

ClickUp

ClickUp is a reasonable place to track post-incident action items and confirm the runbook actually gets updated afterward, rather than the update becoming an open item nobody circles back to.

Visit ClickUp→

Frequently Asked Questions

How often should we test our incident runbook?

A tabletop walkthrough at least twice a year, plus an update after any real incident that revealed a gap. Waiting for a real incident to be your only test means you're finding gaps during the exact moment they're most costly to discover.

Who should own writing and maintaining the incident runbook?

A senior engineer with broad system knowledge should own it, with reviews from people at different experience levels. Include someone relatively new, since they are the most likely to notice a step that assumes context they don't have. Ownership means keeping the runbook current, not writing every step alone.

What's the most common reason runbooks don't get used during a real incident?

People don't trust that it's current, often because it's failed them before in a small way, like linking to a dashboard that no longer exists. Once a runbook loses that trust, people default to memory and improvisation instead, even when a better path was written down.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides