Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Writing an Incident Runbook for When Agents Misbehave

A standard incident runbook assumes a failure that's visible: an error rate spike, a service that's down. An agent incident often looks different at first: the system is technically up and responding, it's just giving wrong or unsafe answers, which means your usual on-call signals may not fire until customers have already noticed something is off.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

How do you define what counts as an agent incident?

Write down the specific signals that should trigger an agent incident response, separate from your standard uptime and error-rate alerts: a spike in fallback responses, a spike in a specific tool's error rate, or a report that the agent took a destructive action incorrectly. Without this list, an agent behaving badly can go unescalated because nothing about it trips a traditional outage alert.

How do you stop a misbehaving agent quickly?

The first response to a misbehaving agent is usually not a code fix, it's disabling the specific capability that's causing harm, a single tool, or the agent entirely, while the team investigates. Build this kill switch before you need it, and make sure whoever is on call actually knows where it is and has practiced using it.

A kill switch only helps if the person holding the pager can find it and trust it. For example, run a short tabletop exercise once in a while: describe a misbehaving tool, then ask the on-call engineer to show, without help, how they would disable just that tool and how they would confirm it stopped. Any hesitation shows you where the documentation is missing. A common mistake is building a switch that turns off the entire agent only, which makes people reluctant to use it. Prefer switches at the level of a single tool or capability so the response stays proportionate to the problem.

Use these first-response steps when an agent misbehaves:

  1. Match the report against your written list of agent incident signals and declare an incident, rather than waiting for a standard uptime alert to fire.
  2. Disable the specific tool or capability causing harm through the kill switch, or the whole agent if the scope is unclear.
  3. Pull the full decision traces for the affected sessions to see which tools ran, with which arguments and results.
  4. Find the root cause from those traces instead of trying to recreate the conversation from scratch.
  5. Add the failing case to the evaluation suite before you turn the capability back on.

Investigate using the decision trace, not just the outcome

Once the immediate behavior is stopped, pull the full decision trace for the affected sessions: which tools were called, what arguments were passed, what results came back. This is usually the fastest path to root cause, faster than trying to reproduce the exact conversation from scratch, and it's only possible if you were already logging full traces before the incident happened.

Close the loop with the evaluation suite

Before turning the affected capability back on, add the incident's specific case to your evaluation suite so it's tested automatically on every future change, not just fixed once and hoped never to recur. High-availability practice budgets a fixed amount of downtime per year depending on the target, for example roughly 8.76 hours annually at a 99.9 percent target1; treat an agent incident's resolution time against a similar discipline, moving fast, but not so fast that the fix goes out untested.

Write a short summary of what happened and why, in plain language, even for a minor incident. The habit of writing these down consistently, not just for the dramatic ones, is what makes a pattern across several small incidents visible before it becomes a large one.

A worked example: a quiet incident that needed the runbook anyway

Say an agent starts approving a small category of refund requests it shouldn't, not because of any single obvious bug, but because a recent tool change subtly altered how it interprets a discount eligibility field. Nothing about this trips a standard uptime alert, the system is fast and responsive throughout. It surfaces because a finance team member happens to notice an unusual pattern in a weekly report, three days after the underlying change shipped.

Having a runbook ready meant the response was immediate once someone flagged it: disable the refund-approval capability through the existing kill switch within minutes, pull the decision traces for the affected sessions to confirm the scope, and identify the root cause in the tool change rather than guessing at several possible explanations. Without a runbook already in place, each of those steps would have needed to be improvised in the middle of an incident that was already three days old by the time anyone noticed, with no clear owner and no agreed first move while the affected capability kept running. The finance team member who first spotted it also became part of the postmortem, since a downstream team noticing an anomaly before engineering does is itself a signal worth building a channel for. That specific gap, no automated alert existed for this pattern at all, became its own follow-up item, separate from the fix to the discount eligibility logic itself, and it's now one of the standard signals the team watches after any change touching pricing or eligibility rules.

Executive Capability Standard

What Good Looks Like

A working incident process for agentic systems defines specific alert signals beyond standard uptime monitoring, gives on-call a fast kill switch for a misbehaving capability, and closes every incident by adding the case to the evaluation suite.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Write down what signals, beyond error rate and uptime, would actually tell you an agent is behaving badly.
2. Do Manually:Manually document the steps to disable your single most important agent capability, and confirm on-call knows where to find them.
3. Delegate:Assign an engineer to own the agent-specific incident runbook and keep it updated as new capabilities ship.
4. Automate:Build a real kill switch and automatic decision-trace retrieval into your agent framework so both are ready before they're needed.
5. Buy:Bring in security and incident-response advisory support such as CrowdStrike if you need help designing detection for this class of incident.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

CrowdStrike

CrowdStrike fits for the detection and response side of an incident that turns out to involve unauthorized access to the systems your agent's tools depend on.

Visit CrowdStrike→

Frequently Asked Questions

How is an agent incident different from a normal service outage?

A normal outage is usually visible immediately through standard uptime and error monitoring. An agent incident can involve a system that's technically responding normally but giving wrong or unsafe answers, so it needs its own specific alert signals, like fallback rate or a spike in a particular tool's errors, to be caught early.

What's the first thing on-call should do when an agent misbehaves?

Stop the specific behavior, usually by disabling the tool or capability involved through a kill switch, before trying to diagnose the root cause. Investigating first while the agent keeps taking the same action live extends the damage unnecessarily.

Should every agent incident update the evaluation suite?

Yes, as a standing part of closing the incident. Adding the specific case that caused the incident to your evaluation set means it's checked automatically on every future prompt or tool change, rather than relying on someone remembering the one-off fix indefinitely.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides