Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Designing Retry and Fallback Logic That Doesn't Make Outages Worse

Retry logic is meant to make a system more resilient, and badly designed retry logic often does the opposite: a dependency starts struggling, every caller retries immediately, and the retry storm finishes off a service that might have recovered on its own if it had gotten a moment of relief instead of a flood of repeated requests.

Good error recovery design treats a struggling dependency the way you'd want to be treated if you were the one struggling: give it room, don't pile on, and have a real answer for what to do while it recovers.

How should retries back off after a failure?

An immediate retry against a service that's failing because it's overloaded adds more load at exactly the wrong moment. Exponential backoff, waiting longer between each successive retry, gives a struggling dependency room to recover instead of compounding the problem. Add jitter, a small random variation in the wait time, so a fleet of callers doesn't retry in synchronized waves that create their own load spikes.

And always cap the total number of retries. Unlimited retries on a permanently failing request just delays the inevitable failure while consuming resources the whole time, which helps nobody.

When should a circuit breaker stop calls to a dependency?

A circuit breaker tracks recent failure rates to a dependency and, once failures cross a threshold, stops sending requests entirely for a cooldown period instead of continuing to retry against something that's clearly not going to succeed. This protects the struggling dependency from being kept down by retry traffic, and it protects your own system's resources from being tied up waiting on calls that are very unlikely to succeed.

After the cooldown, let a small number of test requests through to check whether the dependency has recovered before resuming full traffic, rather than an abrupt full resumption that could immediately overwhelm it again.

Design an actual fallback, not just a retry

Retrying assumes the request will eventually succeed. Sometimes the better answer is a fallback: cached data that's slightly stale but good enough, a degraded response that skips a non-essential enrichment step, a clear error to the user instead of a spinner that eventually times out anyway. Decide, per dependency, whether retry or fallback is the right first response, since treating everything as retryable extends outages that a good fallback would have quietly absorbed.

A 99.95% target leaves you about 4.38 hours of downtime a year to spend1, and a retry storm that keeps hammering an already struggling dependency can burn through a meaningful slice of that in one bad afternoon that a well-designed fallback would have avoided entirely.

Make idempotency a precondition for any retry

Retrying a request that isn't idempotent, one that has a side effect like charging a card or sending an email, risks duplicating that side effect every time a retry fires after a response was lost but the underlying action actually succeeded. Before adding retry logic to any write path, confirm it's genuinely safe to execute more than once, or add an idempotency key so a retry returns the original result instead of repeating the action.

This is a common source of embarrassing duplicate-charge incidents, and it's entirely avoidable by treating idempotency as a precondition for retryability, not an afterthought.

A common mistake: retrying at every layer independently

When each layer of a call chain retries independently, a client retrying three times against a service that itself retries three times against its own dependency can multiply into nine actual attempts at the bottom of the chain from a single original request. Coordinate retry behavior across layers, often by having only the outermost layer retry and having inner layers fail fast, so the effective load multiplication is understood and intentional rather than an accident of independently reasonable-looking decisions.

Trace a request through your actual call chain and count how many total attempts a single failure could generate; the number is often larger and more surprising than anyone expected going in.

A simple fix that works well in practice: pass a retry budget or attempt count along with the request itself, so an inner layer can see it's already been retried upstream and skip its own retry rather than compounding it. That one small piece of shared context prevents most of the accidental multiplication without requiring every team to coordinate their retry configuration manually.

Check each dependency call against this list:

  • Retry with exponential backoff and jitter, and always cap the number of attempts.
  • Add a circuit breaker so a dependency that is clearly down stops receiving requests during a cooldown.
  • Decide per dependency whether retry or a fallback, such as slightly stale cached data, is the first response.
  • Confirm a write path is idempotent, or add an idempotency key, before adding retries to it.
  • Let only the outermost layer retry so attempts do not multiply down the call chain.
Executive Capability Standard

What Good Looks Like

Error recovery is working when a struggling dependency gets backoff and eventual circuit breaking instead of a retry storm, and every retried write path is confirmed idempotent first.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Trace your critical call chains and count how many total attempts a single failure could generate across every layer's independent retry logic.
2. Do Manually:Manually add exponential backoff with jitter and a retry cap to your highest-traffic external calls before building anything more sophisticated.
3. Delegate:Assign an engineer to audit write paths for idempotency and add keys wherever a retry could currently duplicate a side effect.
4. Automate:Implement circuit breakers around your critical external dependencies so sustained failures stop generating retry traffic automatically.
5. Buy:Bring in a platform engineer to design resilience patterns across your service mesh once manual retry logic is scattered and inconsistent across services.

How to Get Started

Frequently Asked Questions

How many times should a client retry a failed request?

Two to three retries with exponential backoff and jitter is a reasonable default for most cases. Beyond that, you're usually just delaying an inevitable failure while consuming resources; a circuit breaker or fallback is a better response to sustained failure than more retries.

Is a circuit breaker overkill for a small application?

Not necessarily, even a simple version, tracking recent failure counts and pausing calls past a threshold, is worth building once you have any external dependency whose failure could otherwise cascade. It's a small amount of code relative to the outage duration it can prevent.

How do we know if a request is safe to retry?

Ask whether executing it twice produces a different or duplicated outcome than executing it once. If yes, it needs an idempotency key or a check for prior completion before you add retry logic to it, not just a blanket assumption that retrying is harmless.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides