Designing Retry Logic That Doesn't Make Things Worse
Retry logic is supposed to make a system more resilient. Done carelessly, it does the opposite: a struggling service gets hit with a wave of retries from every client at once, which is often what turns a brief blip into a real outage.
Here's how to decide when a retry actually helps, and how to build one safely.
When a retry actually helps versus when it just delays the failure
A retry helps against a transient problem: a dropped connection, a brief timeout, a momentary blip in an otherwise healthy dependency. It doesn't help against a genuine, sustained failure, like a dependency that's down or a request that's fundamentally invalid, where retrying just delays an outcome that was never going to change.
Before adding a retry anywhere, ask whether the failure you're handling is actually the transient kind. Retrying a request that will fail every time wastes resources and adds latency without improving anything.
Exponential backoff and why a fixed retry interval causes outages
If every client retries a failing request after exactly the same fixed delay, all of them hit the dependency again at the same moment, which can turn a struggling service into a fully overwhelmed one right as it was starting to recover. Exponential backoff, where each retry waits longer than the last, spreads that load out over time instead of concentrating it.
Adding a small amount of randomness, jitter, to the backoff interval spreads retries out further still, which matters more as the number of clients retrying at once grows.
Fallback logic: degraded, not broken
A good fallback gives the user something useful even when the primary path fails: cached data instead of live data, a simplified response instead of a fully personalized one, or a generic recommendation instead of no recommendation at all. A bad fallback either fails just as hard as having no fallback, or worse, fails silently in a way that looks successful but returns wrong or stale information without any indication that it did.
Design the fallback to be honest about its own degraded state where that matters, rather than quietly pretending everything is normal.
For example, a product recommendation panel can fall back to a cached list of popular items when the personalization service is down, and the customer still sees something useful. A checkout confirmation cannot fall back to a guess, so it should fail clearly and let the customer retry safely. The decision rule is to ask what the user needs from that step and whether an older or simpler answer still serves them. If it does, degrade to it and log the fallback so your team knows it happened. If it does not, surface the failure honestly.
Idempotency is what makes retries safe in the first place
Retrying a request that has a side effect, like charging a customer or sending an email, is only safe if doing it twice has the same effect as doing it once. Without idempotency, a retry can double-charge a customer or send a duplicate notification, which is often a worse outcome than the original failure would have been.
Build idempotency in deliberately, usually with a unique request identifier the downstream system can use to recognize and ignore a duplicate, before you add retries to anything with a real side effect.
A worked example: a retry storm that took down a healthy service
Say a downstream service has a brief thirty-second blip. Without backoff or jitter, every client retrying at a fixed interval hits it again at the same moment, and the resulting spike is larger than the service's normal peak load, so it goes down too, even though it had actually recovered from the original blip already. Exponential backoff with jitter would have spread that same retry traffic out enough for the service to absorb it without ever tipping into a second outage.
Deciding retry policy per call, not one policy for everything
A single global retry policy applied everywhere tends to be wrong for at least some of your calls, since a quick internal lookup and a slow external payment call have very different tolerances for delay and very different consequences for a duplicate side effect. Set retry count, backoff, and idempotency requirements deliberately for each type of call based on what it actually does, rather than copying one default setting across your whole codebase. Write the reasoning down next to the policy, so the next engineer who touches that call understands why it retries three times with backoff instead of assuming it was an arbitrary choice they're free to change.
Before adding a retry to a call, check these points:
- Is the failure likely to be transient, such as a dropped connection or brief timeout, rather than a permanent problem that a retry would only delay?
- Does the retry use exponential backoff with jitter, so clients do not all hit the struggling dependency again at the same moment?
- Is the operation idempotent, meaning a second attempt has the same effect as the first, especially if it writes data or charges a customer?
- Is the number of attempts capped, with a sustained failure surfaced for handling instead of retried indefinitely?
- Does this specific call have its own policy, rather than inheriting one global setting that suits a quick lookup but not a slow call?
What Good Looks Like
Good here means every retry uses backoff and jitter instead of a fixed interval, every retried operation is genuinely idempotent, and fallback behavior is honest about being degraded rather than silently wrong.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many times should a request be retried before giving up?
There's no universal number, but two or three attempts with exponential backoff is a common starting point for most calls. Beyond that, the added latency from more attempts usually costs more than the small extra chance of success is worth, and a real, sustained failure should surface to be handled, not retried indefinitely.
Should retries be handled by the client or the server?
Usually the client, since it's the one that knows whether the request actually failed and can decide whether retrying still makes sense. A server can help by returning clear signals, like a specific status code for a transient failure versus a permanent one, so the client's retry logic has good information to act on.
Is it safe to retry a request that writes data?
Only if the operation is idempotent, meaning doing it twice has the same effect as doing it once. Without that guarantee, a retry on a write can create a duplicate record or a duplicate side effect, which is often worse than simply surfacing the original failure to be handled directly.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Designing Retry and Fallback Logic That Doesn't Make Outages Worse
How to design retries, backoff, and circuit breakers so error recovery logic protects a struggling dependency instead of overwhelming it further.
The Retry Logic That Makes Outages Worse, Not Better
How naive retry and fallback logic can amplify an outage instead of recovering from it, and the specific patterns that actually help.
Where Retry Logic Quietly Drains Your Infrastructure Budget
How poorly designed error handling and retry logic turns a small outage into a large bill, and the specific patterns worth fixing first.
Designing Retry and Fallback Logic That Doesn't Undermine Your Access Controls
A checklist for building retry, idempotency, and fallback logic for zero trust APIs, plus the specific pitfalls that quietly weaken access control.
Designing Fallback Logic That Doesn't Make Things Worse
How to design retry, fallback, and fail-visibly logic for AI model serving without causing a retry storm during an outage.
What Should Happen When Your Vector Search Call Fails
Cached results, keyword fallback, or an honest error message: decide a RAG pipeline's failure behavior in advance, per feature, not during the outage.