An Incident Runbook Your On-Call Engineer Can Actually Use
A generic incident runbook, check the dashboard, escalate if needed, is nearly useless overnight for a model-serving outage specifically, because the diagnostic steps are different from a normal service outage: is the model down, is the provider down, is it a model quality issue masquerading as an outage, or is it actually infrastructure.
Write the runbook around those specific branches, not a generic incident template with the word model inserted.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
The first five minutes: which branch are you in
- Check the infrastructure first: is the endpoint reachable at all, are GPU nodes healthy, is the load balancer routing correctly.
- Check the provider, if you depend on a third-party model API: their status page, and whether your fallback has already engaged.
- Check whether it's a quality issue, not an outage: requests are succeeding but a spike in refusals or a change in output pattern is what triggered the alert.
- Check whether it started right after a deploy: model swap, prompt change, or config change in the last hour is the first suspect if the timing lines up.
Each branch has a different next step; guessing which one you're in wastes the minutes that matter most.
What the runbook needs for each branch
For an infrastructure outage: the rollback command or traffic-shift toggle, and who to page if it doesn't resolve within a defined window. For a provider outage: confirmation that fallback logic has engaged, and the provider's status page link, not a promise to check email for an update. For a quality issue: how to pull a sample of recent bad outputs quickly, and the model-swap rollback command, since a quality issue is usually a deploy problem in disguise.
Each of these needs the actual command or link in the runbook itself, not a description of where someone could probably find it under pressure.
A useful test of the runbook is a tabletop exercise. Pick a plausible scenario, such as a provider slowdown that starts right after a deploy, and have someone who did not write the runbook walk through it out loud using only the document. Every point where they hesitate, guess, or ask a question is a gap to fix, whether it is a missing command, an unclear owner, or an ambiguous threshold. Repeat the exercise after major architecture changes, and rotate who plays the on-call role so the document works for people other than its author.
How much downtime you can actually afford before this becomes a crisis
Knowing your error budget in the moment matters more than it sounds. At a 99.9 percent availability target, you're working with a downtime budget of roughly 8.76 hours across the entire year1, which for most teams means a single incident lasting more than an hour or two is already consuming a meaningful chunk of the whole year's allowance.
Put that number in the runbook itself, not just in a dashboard nobody checks mid-incident, so the on-call engineer has a concrete sense of urgency beyond customers are affected.
After the incident: what actually needs to happen
Write the post-incident review while details are still fresh, within a day or two, not weeks later once memory has smoothed over the actual sequence of events. Capture what alerted first, how long each diagnostic branch took, and whether the runbook itself pointed the right direction or needed a workaround.
Feed every workaround back into the runbook immediately. A runbook that doesn't get updated after every incident it was used for slowly drifts out of date exactly where it matters most.
Who owns the decision to declare an incident
Ambiguity about who can officially call something an incident, versus a rough patch that will probably resolve itself, wastes time at exactly the moment speed matters most. Name the role, not a specific person, since the on-call rotation changes, and give that role explicit authority to declare without waiting for approval from someone more senior.
A runbook that's clear on diagnostic steps but silent on who gets to pull the alarm still leaves room for a confused first ten minutes while people wait for someone else to make the call instead of just making it.
Runbook mistakes that cost time during a real incident
- Writing the runbook once at launch and never updating it as the architecture changes underneath it.
- Describing where to find a command instead of including the actual command.
- No clear branch for a quality issue that isn't technically an outage, leaving the on-call engineer improvising exactly when a script would help most.
What Good Looks Like
A working incident runbook branches explicitly for infrastructure outages, provider outages, and quality issues, includes the actual commands and links needed rather than descriptions of where to find them, and gets updated after every real incident it's used for.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
What should a model-serving incident runbook cover that a generic one doesn't?
Explicit branches for the failure modes specific to model serving: an infrastructure outage, a third-party provider outage, and a quality issue that isn't technically a service outage at all. A generic runbook built for a normal web service usually only covers the first of these three.
How often should we update our incident runbook?
After every incident it's used for, immediately, while the gaps are fresh. A runbook written once at launch and never revisited drifts out of date as your architecture changes, and it tends to be wrong exactly where a real incident needs it to be right.
Why does it matter to know our downtime budget during an active incident?
It gives the on-call engineer a concrete sense of urgency beyond a general sense that customers are affected. Put the actual number in the runbook itself, since a dashboard nobody checks mid-incident doesn't help someone make a fast escalation decision overnight.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
A Rollout Checklist for Swapping Models in Production
A rollout checklist for swapping AI models in production: evaluation gates, canary and shadow traffic, and fast rollback paths.
Why Inference Latency Creeps Up After You Ship
Where AI model-serving latency actually hides: tokenization, queueing, batching windows, and network hops, plus a worked example fix.
Zero Trust for Machines Calling Your Model Endpoints
Why internal services calling your inference endpoints still need identity checks, and where CrowdStrike-style posture checks and Tenable-style scanning fit.
Writing an Incident Response Runbook People Actually Follow at 3 A.M.
A worksheet approach to writing incident runbooks that hold up under real pressure, when the person on call is tired, stressed, and reading fast.
What SOC 2 Actually Expects From a Model-Serving Team
What SOC 2 expects from a team serving AI models: how change, access, patch, and vendor controls apply, and the evidence to have ready.
What a Real Security Audit of Model Serving Should Cover
A practical checklist for auditing AI model serving and inference: endpoint access, weight security, prompt logging, and patch timelines.