Keeping Agent Workflows Running When a Region Goes Down
An agent loop typically depends on more moving parts than the rest of your application: the model provider's API, several MCP servers, and whatever databases those servers call. High availability planning for this stack means asking, for each piece, what happens to the agent when it's unreachable, not just whether the piece itself has a failover plan.
Build the worksheet: one row per dependency
For every MCP server and every external API your agents call, write down three things: what the agent should do if that dependency times out, what it should do if it returns an error, and whether the task can complete at all without it. A tool that looks up a customer's order history might be skippable, with the agent noting it couldn't check, while a tool that processes a payment cannot be skipped without the task simply failing.
This worksheet turns a vague goal like "the agent should be resilient" into a specific, reviewable list of decisions.
How much availability does your agent workflow actually need?
Higher availability costs more in engineering time and infrastructure, so match the target to what the workflow actually needs. A target of 99.9 percent leaves roughly 8.76 hours of downtime a year to work with, while 99.99 percent leaves around 52.6 minutes1. An internal agent that helps engineers triage tickets can usually live with the looser budget; an agent handling live customer conversations usually can't.
How should an agent degrade gracefully when a tool is down?
A model provider outage or a downstream MCP server outage doesn't have to take the whole agent down. Build the agent to recognize which of its tools are unavailable and either complete a reduced version of the task or hand off to a human with a clear explanation of what it couldn't check, rather than returning a raw error or, worse, guessing at an answer it can't actually support.
For example, an agent that summarizes support tickets can respond to a model provider outage by queuing the request and telling the user the summary will arrive shortly. An agent that approves refunds should stop and route to a person instead, because a delayed answer costs little while a guessed approval costs real money. Writing the degraded behavior next to each row of the worksheet, including the message the user sees, keeps the response the same no matter which conversation hits the outage first, and lets you review it before an incident instead of improvising during one.
Test the failure, don't just plan for it
Run a scheduled drill where you deliberately disable one MCP server or simulate a model provider timeout and watch what the agent actually does. Plans that look reasonable on paper often reveal a tool with no timeout at all, one that will hang the whole agent loop rather than fail fast, and you only find that by testing it.
Write down what you expected to happen before you run the drill, then compare it against what actually happened afterward. The gap between the two is usually where the real work is, and it's a much cheaper place to find that gap than during an actual outage with real customers waiting on a response.
Run a failure drill in this sequence:
- Write down what you expect the agent to do before you run anything, so the result has something to be compared against.
- Deliberately disable one MCP server or simulate a model provider timeout.
- Watch what the agent actually does, including whether any tool hangs the whole loop instead of failing fast.
- Compare the outcome with your written expectation, since the gap between the two is usually where the real work is.
- Fix the gaps, then repeat the drill after adding any new MCP server that other tools now depend on.
A worked example: one dependency, three different failure plans
Say your agent stack has one MCP tool that looks up shipping status from a third-party carrier API. For a customer-facing tracking assistant, the fallback might be to tell the user the agent couldn't reach the carrier right now and to try again shortly, since a wrong shipping estimate is worse than an honest "I don't know." For an internal support agent helping an employee investigate a complaint, the fallback might instead be to proceed with a note that shipping status couldn't be confirmed, since the employee can judge for themselves whether that gap matters for the specific case. For a fully automated workflow that issues refunds based partly on shipping delay, the fallback has to be to stop and route to a human, because acting without that data at all is the one option that's actually unsafe.
Same dependency, same failure, three different correct answers, because the right fallback depends entirely on what happens downstream of the agent's decision, not on the tool itself. Writing all three down explicitly, rather than leaving the agent to improvise a response to the same outage in each context, is what makes the difference between a graceful degradation and an inconsistent one that varies by which conversation happened to hit it first.
What Good Looks Like
A resilient agent stack has a documented fallback behavior for every dependency, an availability target that matches the workflow's actual stakes, and a tested pattern for degrading gracefully instead of hanging or guessing when something is unavailable.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need multi-region failover for every agent workflow?
No. Match the investment to the workflow's actual availability target. A customer-facing agent handling live conversations usually justifies it; an internal reporting agent that runs a few times a day usually doesn't need the added complexity and cost.
What should an agent do when a tool call times out?
Fail fast with a short timeout, then either retry once with backoff or fall back to a degraded response that tells the user or the next system what it couldn't check. A tool with no timeout at all is the most common cause of an agent loop hanging indefinitely.
How often should we run a failure drill on the agent stack?
Quarterly is a reasonable baseline, and always after adding a new MCP server that other tools now depend on. A drill after a change catches assumptions the original design didn't account for, before a real outage does.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Rolling Out Agentic Workflows Without Breaking Production
A practical rollout checklist for shipping an AI agent to production, from a shadow-mode test run through the guardrails that catch it if it misbehaves.
What High Availability Actually Costs Beyond the Second Region
A worked-example breakdown of what running a second region for failover really costs, and how to decide whether your uptime target justifies it.
Build vs. Buy for Verifying Every Device That Connects In
What zero-trust device and identity verification actually requires, what a platform gives you over a homegrown check, and how to decide between them.
Active-Active vs Active-Passive: What Your Uptime Target Buys You
A comparison of active-active, active-passive and single-region failover, with the real infrastructure and headcount cost each uptime target requires.
What Happens When a Tool Call Fails Mid-Task
A decision guide for designing fallback logic in an agent loop, so a single failed tool call degrades gracefully instead of derailing the whole task.
How Much Redundancy Your Model-Serving Stack Actually Needs
A practical look at high availability for AI model serving: active-passive versus active-active, provider fallbacks, and real uptime costs.