What Happens When a Tool Call Fails Mid-Task
A tool call can fail for reasons that have nothing to do with the model: a downstream API timing out, a database briefly unavailable, a rate limit on a third-party service. What separates a resilient agent from a fragile one is not whether these failures happen, they will, but what the agent does the moment one does.
How should you decide a tool's fallback ahead of time?
For every tool, decide in advance: should a failure here stop the whole task, should the agent retry, or should it proceed with a note that this step couldn't be completed. A tool that checks a customer's loyalty tier before applying a discount might be safe to skip with a note; a tool that actually applies the discount is not something you want the agent proceeding without.
Retry with backoff, but cap it explicitly
A single retry after a short delay handles most transient failures. Retrying indefinitely, or retrying immediately without any delay, turns a brief downstream blip into sustained load on a system that's already struggling, which can make the underlying problem worse. Cap retries at a small, explicit number and fail toward the safe fallback after that, rather than looping.
For example, a tool that fetches a shipping estimate might be retried once after a short delay and, if it still fails, the agent continues with a note that the estimate could not be confirmed. A tool that submits an order should not be retried blindly at all, because a repeat call after an unclear failure can create a duplicate. Deciding retry behavior per tool, based on whether repeating the call is safe, is more reliable than applying a single retry rule everywhere.
How does the model learn that a tool call failed?
When a tool fails, the model needs a clear, structured signal that it failed and why, distinct from a successful result that happens to be empty. A model that can't tell "no results found" from "the tool errored" will sometimes present a search failure to the user as if it were a confident "there's nothing here," which is a worse outcome than either the failure or the empty result surfaced honestly on its own.
Give the agent an honest way to say it couldn't finish
When a critical tool fails and there's no safe fallback, the agent's response to the user should say plainly what it couldn't verify or complete, rather than guessing at an answer to avoid seeming unhelpful. This is a design decision, not something that happens automatically, and it's worth testing directly: deliberately break a tool in a test environment and check what the agent actually tells the user.
Models left without explicit instruction here will often fill the gap with something plausible-sounding rather than an honest admission of uncertainty, simply because a confident-sounding answer is closer to what most of their training examples look like. Write the fallback behavior into the prompt explicitly rather than assuming the model will default to caution on its own.
Design failure handling for each tool in this order:
- Decide in advance whether a failure should stop the task, trigger a retry, or let the agent continue with a note.
- Retry once or twice after a short delay, capped explicitly, then fail toward the safe fallback.
- Return a structured failure signal that is clearly distinct from a successful but empty result.
- Write into the prompt how the agent should tell the user what it could not verify or complete.
- Test by deliberately disabling the tool in staging and reading exactly what the agent tells the user.
A worked example: two failures, two very different outcomes
Say a pricing lookup tool fails partway through a quote request. In one version of the agent, built without explicit failure handling, the model, still holding a partial result from an earlier successful call in the same conversation, presents that stale figure as the current quote, confidently and without caveat. In a second version, the same failure produces a clear message: it couldn't confirm current pricing and the user should try again shortly, with a note that the price it saw earlier in the conversation may no longer be accurate.
The second version is a worse experience in the moment, but the first one is the kind of failure that erodes trust much more seriously once a customer discovers the number they acted on was wrong. Building and testing the honest failure path deliberately, rather than trusting the model to find it unprompted, is what separates the two outcomes, and it's a small amount of prompt work compared with the cost of repairing trust after the first version's mistake reaches a customer. The test itself is simple to run: disable the pricing tool in a staging environment, send the same quote request, and read exactly what comes back. Do this for every tool whose failure could plausibly lead to a wrong action, not only the one that happened to cause an incident, since the same gap tends to exist wherever it hasn't been checked.
What Good Looks Like
Good error recovery for an agent means every tool has a decided fallback ahead of time, retries are capped explicitly, failures are surfaced to the model as distinct from empty results, and the agent tells the user honestly when it couldn't complete a task.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should an agent always retry a failed tool call?
Only up to a small, explicit limit, usually once or twice with a short delay. Retrying indefinitely can turn a brief downstream issue into sustained load that makes the underlying problem worse, and it delays the moment the agent falls back to a safer response.
How should a tool failure be different from an empty result in what the model sees?
The two need distinct, structured signals. If a model can't tell a genuine "no data found" from a tool that simply errored, it can present a failure to the user as a confident negative answer, which is worse than either outcome reported honestly.
What should an agent do when it truly can't complete a task?
Say so directly, naming what it couldn't verify or do, rather than guessing at an answer to seem more helpful. Test this specifically by deliberately failing a tool in a staging environment and checking the agent's actual response, since it rarely behaves the way you'd assume without checking.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Designing Retry and Fallback Logic That Doesn't Make Outages Worse
How to design retries, backoff, and circuit breakers so error recovery logic protects a struggling dependency instead of overwhelming it further.
Designing Fallback Logic That Doesn't Make Things Worse
How to design retry, fallback, and fail-visibly logic for AI model serving without causing a retry storm during an outage.
The Retry Logic That Makes Outages Worse, Not Better
How naive retry and fallback logic can amplify an outage instead of recovering from it, and the specific patterns that actually help.
Designing Retry Logic That Doesn't Make Things Worse
How to decide when a retry actually helps, why fixed intervals cause outages, and how to design fallback logic that degrades instead of breaking.
Designing Retry and Fallback Logic That Doesn't Undermine Your Access Controls
A checklist for building retry, idempotency, and fallback logic for zero trust APIs, plus the specific pitfalls that quietly weaken access control.
Where Retry Logic Quietly Drains Your Infrastructure Budget
How poorly designed error handling and retry logic turns a small outage into a large bill, and the specific patterns worth fixing first.