Writing an Incident Runbook for When Agents Misbehave
A standard incident runbook assumes a failure that's visible: an error rate spike, a service that's down. An agent incident often looks different at first: the system is technically up and responding, it's just giving wrong or unsafe answers, which means your usual on-call signals may not fire until customers have already noticed something is off.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
How do you define what counts as an agent incident?
Write down the specific signals that should trigger an agent incident response, separate from your standard uptime and error-rate alerts: a spike in fallback responses, a spike in a specific tool's error rate, or a report that the agent took a destructive action incorrectly. Without this list, an agent behaving badly can go unescalated because nothing about it trips a traditional outage alert.
How do you stop a misbehaving agent quickly?
The first response to a misbehaving agent is usually not a code fix, it's disabling the specific capability that's causing harm, a single tool, or the agent entirely, while the team investigates. Build this kill switch before you need it, and make sure whoever is on call actually knows where it is and has practiced using it.
A kill switch only helps if the person holding the pager can find it and trust it. For example, run a short tabletop exercise once in a while: describe a misbehaving tool, then ask the on-call engineer to show, without help, how they would disable just that tool and how they would confirm it stopped. Any hesitation shows you where the documentation is missing. A common mistake is building a switch that turns off the entire agent only, which makes people reluctant to use it. Prefer switches at the level of a single tool or capability so the response stays proportionate to the problem.
Use these first-response steps when an agent misbehaves:
- Match the report against your written list of agent incident signals and declare an incident, rather than waiting for a standard uptime alert to fire.
- Disable the specific tool or capability causing harm through the kill switch, or the whole agent if the scope is unclear.
- Pull the full decision traces for the affected sessions to see which tools ran, with which arguments and results.
- Find the root cause from those traces instead of trying to recreate the conversation from scratch.
- Add the failing case to the evaluation suite before you turn the capability back on.
Investigate using the decision trace, not just the outcome
Once the immediate behavior is stopped, pull the full decision trace for the affected sessions: which tools were called, what arguments were passed, what results came back. This is usually the fastest path to root cause, faster than trying to reproduce the exact conversation from scratch, and it's only possible if you were already logging full traces before the incident happened.
Close the loop with the evaluation suite
Before turning the affected capability back on, add the incident's specific case to your evaluation suite so it's tested automatically on every future change, not just fixed once and hoped never to recur. High-availability practice budgets a fixed amount of downtime per year depending on the target, for example roughly 8.76 hours annually at a 99.9 percent target1; treat an agent incident's resolution time against a similar discipline, moving fast, but not so fast that the fix goes out untested.
Write a short summary of what happened and why, in plain language, even for a minor incident. The habit of writing these down consistently, not just for the dramatic ones, is what makes a pattern across several small incidents visible before it becomes a large one.
A worked example: a quiet incident that needed the runbook anyway
Say an agent starts approving a small category of refund requests it shouldn't, not because of any single obvious bug, but because a recent tool change subtly altered how it interprets a discount eligibility field. Nothing about this trips a standard uptime alert, the system is fast and responsive throughout. It surfaces because a finance team member happens to notice an unusual pattern in a weekly report, three days after the underlying change shipped.
Having a runbook ready meant the response was immediate once someone flagged it: disable the refund-approval capability through the existing kill switch within minutes, pull the decision traces for the affected sessions to confirm the scope, and identify the root cause in the tool change rather than guessing at several possible explanations. Without a runbook already in place, each of those steps would have needed to be improvised in the middle of an incident that was already three days old by the time anyone noticed, with no clear owner and no agreed first move while the affected capability kept running. The finance team member who first spotted it also became part of the postmortem, since a downstream team noticing an anomaly before engineering does is itself a signal worth building a channel for. That specific gap, no automated alert existed for this pattern at all, became its own follow-up item, separate from the fix to the discount eligibility logic itself, and it's now one of the standard signals the team watches after any change touching pricing or eligibility rules.
What Good Looks Like
A working incident process for agentic systems defines specific alert signals beyond standard uptime monitoring, gives on-call a fast kill switch for a misbehaving capability, and closes every incident by adding the case to the evaluation suite.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How is an agent incident different from a normal service outage?
A normal outage is usually visible immediately through standard uptime and error monitoring. An agent incident can involve a system that's technically responding normally but giving wrong or unsafe answers, so it needs its own specific alert signals, like fallback rate or a spike in a particular tool's errors, to be caught early.
What's the first thing on-call should do when an agent misbehaves?
Stop the specific behavior, usually by disabling the tool or capability involved through a kill switch, before trying to diagnose the root cause. Investigating first while the agent keeps taking the same action live extends the damage unnecessarily.
Should every agent incident update the evaluation suite?
Yes, as a standing part of closing the incident. Adding the specific case that caused the incident to your evaluation set means it's checked automatically on every future prompt or tool change, rather than relying on someone remembering the one-off fix indefinitely.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
Writing an Incident Response Runbook People Actually Follow at 3 A.M.
A worksheet approach to writing incident runbooks that hold up under real pressure, when the person on call is tired, stressed, and reading fast.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
An Incident Runbook Your On-Call Engineer Can Actually Use
How to write an AI model-serving incident runbook with real branches for infrastructure, provider, and quality-issue outages.
Why Your Agent Loop Feels Slow, and How to Fix It
A diagnostic guide to finding where latency actually comes from in an agentic system, and which fixes help each cause instead of masking it.
Writing an Incident Runbook People Actually Follow at 2 A.M.
How to write an incident response runbook that a half-awake, stressed engineer can actually follow, instead of one that only reads well in review.