What Actually Belongs in an Incident Response Runbook
An incident response runbook that lists every possible failure and its root cause reads well in a calm afternoon review and is nearly useless at two in the morning during an actual outage. A useful runbook isn't a reference document, it's a script for the first confusing minutes, written by someone who wasn't panicking, for someone who will be.
This is what actually belongs in that document, and what tends to get left out.
Is an incident runbook the same as an architecture diagram?
It's tempting to fold an incident runbook into your general engineering documentation: here's the system, here's how it's wired, here's who owns what. During an actual incident, nobody has time to read background material, they need the next three steps stated plainly. Keep the runbook itself short and action-oriented, and link out to the deeper architecture docs for anyone who has time to read them once the immediate fire is out.
A runbook that takes more than a minute to find the right section in is a runbook that will get abandoned in favor of just guessing, which is exactly the outcome a runbook exists to prevent.
Write the first five minutes down explicitly, not just the eventual fix
Most runbooks focus on the fix: restart this service, roll back this deploy, fail over to this region. The first five minutes of a real incident are rarely about the fix, they're about figuring out what's actually broken, who else needs to know, and whether this is even the right runbook for what's happening.
Write that triage step out explicitly: which dashboard to check first, which log to search, what confirms this incident matches this runbook rather than a different one. An engineer who gets paged awake at three in the morning benefits far more from a clear first step than from a comprehensive list of fixes for problems they haven't yet confirmed they're facing.
A useful first-five-minutes section answers these questions in plain steps:
- Which dashboard to open first, so the responder starts from the same evidence every time.
- Which log to search, and what to look for in it.
- What confirms the incident matches this runbook rather than a different one.
- Who else needs to know right now, and how to reach them.
Who should be incident commander on a two-person team?
During an incident, someone needs to own the decision of what happens next, separate from whoever is heads-down fixing the actual problem. Without that split, the person best positioned to fix the issue also ends up fielding status questions, coordinating with other teams, and deciding when to escalate, and all three of those pull attention away from the fix itself.
On a small team, this can be as simple as naming whoever isn't currently debugging as the person who owns communication and escalation decisions for that incident. The role matters more than the org chart: even a two-person on-call rotation benefits from one person explicitly not touching the keyboard.
Decide your communication cadence before the incident, not during it
Stakeholders who aren't getting updates during an outage tend to interrupt the people fixing it to ask for one, which is the opposite of helpful. Decide in advance how often updates go out during an active incident, say every fifteen or thirty minutes depending on severity, and where they get posted, so nobody has to make that call while also trying to diagnose a database connection pool that's exhausted.
A short, honest update, still investigating, no new information, sent on schedule does more to keep people from interrupting than a detailed update sent only once something concrete is known.
For example, a team might agree that a severe outage gets an update every fifteen minutes in a single shared channel, and a minor degradation gets one every thirty. The commander posts on schedule even when the message is only that the team is still investigating and has nothing new. Write those two rules, plus the name of the channel, into the runbook itself. When the page goes off, nobody has to negotiate cadence, and the engineers fixing the problem are not interrupted by people who simply want to know whether anyone is working on it.
A common mistake: skipping the postmortem once the fix ships
Once a fix ships and the alerts clear, the temptation is to move on to whatever was interrupted, and the postmortem quietly never happens. The information about what actually went wrong, and what almost went wrong but got lucky, is freshest in the hours right after an incident and decays fast as people move on to other work.
Schedule the postmortem before the incident even ends, not as a maybe-later item, and make it about the process and the system, not about who made which call under pressure. A runbook that never gets updated based on what the last incident actually revealed just repeats the same gaps the next time something breaks.
What Good Looks Like
Good here means every active runbook states an explicit first step, names who owns incident commander decisions, and a postmortem gets scheduled before the incident is even fully resolved.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How long should an incident runbook actually be?
Short enough to scan in under a minute during an active incident. If a runbook takes longer than that to find the relevant section, most on-call engineers will start guessing instead of reading it, which defeats the reason it exists in the first place.
Who should be the incident commander on a very small team?
Whoever isn't currently debugging the problem. The role is about owning communication and escalation decisions, not about seniority, and even a two-person on-call rotation benefits from explicitly separating who fixes the issue from who manages everything else around it.
What's the biggest reason postmortems get skipped?
Once the fix ships and alerts clear, attention moves on to whatever was interrupted, and a postmortem that isn't scheduled before the incident ends tends to just never happen. The details fade fast, so scheduling it immediately matters more than most teams expect.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Writing an Incident Response Runbook People Actually Follow at 3 A.M.
A worksheet approach to writing incident runbooks that hold up under real pressure, when the person on call is tired, stressed, and reading fast.
Writing an Incident Runbook People Actually Follow at 2 A.M.
How to write an incident response runbook that a half-awake, stressed engineer can actually follow, instead of one that only reads well in review.
Writing an Incident Runbook Someone Can Actually Follow at 3 AM
A step by step way to write a streaming pipeline incident runbook that a half-awake on-call engineer can actually follow, not just a policy document.
Writing an Incident Runbook People Will Actually Follow
How to write an incident response runbook engineers actually reach for during a real outage, instead of one that sits unread until the next audit.
An On-Call Runbook for When Retrieval Quality Drops
A concrete triage order for a RAG on-call incident: outage versus quality drop, the three most common causes to check first, and a scoped kill switch.
The Runbook Nobody Can Find During an Actual Incident
Why most incident runbooks go unused during a real outage, and how to write ones that actually get followed under pressure.