The Retry Logic That Makes Outages Worse, Not Better
Retry logic is supposed to make a system more resilient, but the naive version, retry immediately, retry forever, retry without limit, often makes an outage worse instead of better. When a downstream dependency is already struggling under load, a wave of immediate retries from every caller is exactly the kind of additional load that turns a brief blip into a sustained outage.
Here's what naive retry logic actually does under real failure conditions, and the patterns that avoid making things worse.
The retry storm: how a small failure becomes a big one
Picture a downstream service that briefly returns errors due to a deploy or a momentary resource spike. Every caller's naive retry logic fires at roughly the same time, often within milliseconds of each other since they were all triggered by the same failure. That synchronized wave of retries hits the already-struggling service exactly when it's least able to handle extra load, which can turn a two-second blip into a multi-minute outage as the retries themselves become the dominant source of traffic.
This is exactly the failure mode that makes an incident hard to diagnose in the moment. The on-call engineer sees a downstream dependency's error rate climbing and reasonably assumes the dependency itself is the sole problem, when a meaningful share of that load is actually their own system's retries piling on top of an already-recovering service.
How does jitter keep retries from arriving synchronized?
Instead of every caller retrying at a fixed interval, add randomized jitter to the retry delay so requests spread out over a window rather than arriving in a synchronized burst. Combined with exponential backoff, doubling the wait time on each successive attempt, this gives a struggling downstream service breathing room to recover instead of facing a wall of retries every time a fixed interval elapses.
Why cap retries and fail fast past the limit?
Unlimited retries turn a downstream outage into an upstream one: a caller stuck retrying forever holds a connection, a thread, or a queue slot the whole time, and enough stuck callers exhaust your own service's capacity even though the actual problem is somewhere else entirely. Set a hard cap on retry attempts and a clear, fast failure past that cap, so a struggling dependency degrades your service gracefully instead of taking it down entirely alongside it.
Use a circuit breaker to stop calling a service that's clearly down
A circuit breaker tracks failure rate to a specific dependency and, once it crosses a threshold, stops sending new requests to that dependency for a cooldown period, failing fast instead. This protects both sides: your own service stops wasting resources on calls likely to fail, and the struggling downstream dependency gets a reprieve from load while it recovers, rather than facing a continued stream of retries throughout its entire outage.
- Add jitter to retry delays so requests don't arrive in a synchronized wave
- Use exponential backoff, not a fixed interval, between successive attempts
- Cap total retry attempts and fail fast, clearly, once the cap is hit
- Add a circuit breaker for dependencies where repeated failures are common enough to warrant one
Review these four items together whenever a new external dependency is added, not just after an incident forces the conversation. It's far cheaper to design retry behavior correctly the first time a call site is written than to retrofit jitter, backoff and a circuit breaker into a service after a retry storm has already caused real customer-facing downtime.
Decide what a fallback actually returns, deliberately
When a retry ultimately fails, what does the calling service actually do? A generic error passed straight to the end user is honest but often not the best available option. For some flows, a cached or default value is a better fallback than a hard failure, showing a slightly stale price instead of no price at all, for example. Decide this per call site rather than defaulting to the same generic error handling everywhere; the right fallback for a checkout flow is rarely the right fallback for a non-critical recommendation widget.
What Good Looks Like
Every retry uses exponential backoff with jitter and a hard cap, dependencies with correlated failure modes have a circuit breaker, and each call site has a deliberate, tested fallback rather than a generic default.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How many retry attempts is reasonable before giving up?
There's no universal number, but two or three attempts with exponential backoff and jitter is a common, reasonable starting point for most internal calls. More than that usually adds latency without meaningfully improving success rates, especially once a dependency is genuinely struggling rather than experiencing a single transient blip.
Do we need a circuit breaker for every downstream call?
No. Reserve circuit breakers for dependencies where failures are correlated and where continuing to call a struggling service actively makes things worse for both sides. A rarely-failing, low-volume internal call may not need the added complexity.
How do we test that our retry and fallback logic actually works as designed?
Deliberately fail a dependency in a staging environment, or use a chaos-testing tool to inject failures, and watch what actually happens: do retries spread out with jitter, does the circuit breaker trip as expected, does the fallback return something reasonable. Assuming the logic works because it compiles is how these gaps go unnoticed until a real outage.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Designing Retry and Fallback Logic That Doesn't Make Outages Worse
How to design retries, backoff, and circuit breakers so error recovery logic protects a struggling dependency instead of overwhelming it further.
Designing Retry Logic That Doesn't Make Things Worse
How to decide when a retry actually helps, why fixed intervals cause outages, and how to design fallback logic that degrades instead of breaking.
Where Retry Logic Quietly Drains Your Infrastructure Budget
How poorly designed error handling and retry logic turns a small outage into a large bill, and the specific patterns worth fixing first.
Designing Retry and Fallback Logic That Doesn't Undermine Your Access Controls
A checklist for building retry, idempotency, and fallback logic for zero trust APIs, plus the specific pitfalls that quietly weaken access control.
Designing Fallback Logic That Doesn't Make Things Worse
How to design retry, fallback, and fail-visibly logic for AI model serving without causing a retry storm during an outage.
The Retry Loop That Took Down the Service It Was Trying to Save
How naive retry logic turns a small hiccup into an outage, and the specific patterns that make error recovery actually safe.