Designing Fallback Logic That Doesn't Make Things Worse
Retry and fallback logic gets added to a model-serving integration reactively, usually right after an outage, and it often makes the next incident worse instead of better: a retry storm that amplifies load on an already struggling model server, or a fallback that silently returns a degraded answer with no signal that anything went wrong.
Good error recovery design happens before the incident, with explicit rules for when to retry, when to fall back, and when to just fail visibly.
Retry, fallback, or fail: three responses to three different failures
- Retry for genuinely transient failures: a timeout, a rate limit, a brief network blip. Use backoff, not immediate retries, or you turn a small problem into a bigger one.
- Fallback for a failure likely to persist briefly: switch to a secondary model or provider, or a cached or templated response, when the primary is degraded.
- Fail visibly for a failure that a retry or fallback would only mask: a malformed request, an authentication error, anything where returning a fake success hides a real bug.
Mixing these up, retrying something that should fail fast, or failing loudly on something transient, is the most common design mistake.
Why retry storms happen, and how to avoid causing one
When a model server starts struggling under load and every client retries immediately on failure, the retry traffic adds to the load that caused the problem in the first place, which can turn a slowdown into a full outage. This is the most damaging failure mode in error recovery design, and it's entirely self-inflicted.
Use exponential backoff with jitter, not a fixed retry interval, so retries from many clients spread out rather than arriving in synchronized waves. Cap the total number of retries per request, and treat a request that has exhausted its retries as a real failure to surface, not something to retry forever.
A good rule for setting retry limits is to give each request a total time budget rather than only a retry count. If the caller will give up after a fixed number of seconds anyway, retries that would finish after that point only add load. For example, if your gateway stops waiting after its timeout, fit two or three attempts inside that window with growing pauses, then return a clear error. That keeps the total work per request bounded even when many clients fail at once, which is exactly when bounded work matters most.
Fallbacks need their own signal, or nobody knows they're happening
A fallback response that looks identical to a normal one hides a real problem: your primary model or provider might be down for hours before anyone notices, because every request still technically succeeded. Tag fallback responses internally, even if the caller-facing output looks the same, and alert when the fallback rate rises above its normal baseline.
A tracked fallback is a feature. An untracked one is a blind spot that happens to look like uptime.
A worked example: a provider outage handled two different ways
Say your primary model provider goes down for twenty minutes. With a tracked fallback to a secondary provider and an alert on the fallback rate, your team knows within minutes, can decide whether the degraded quality is acceptable, and gets a clear signal when the primary recovers.
Without tracking, the same twenty minutes passes with users getting slightly worse answers and nobody on your team aware anything happened, until a support ticket or a metric review days later surfaces it. The infrastructure difference between these two outcomes is a few log fields and one alert.
Testing your fallback logic before you need it
A fallback path that's never been exercised outside of code review is a fallback you don't actually know works. Periodically force a failure in a lower environment, kill the primary model connection deliberately, and confirm the fallback fires, gets tagged correctly, and the alert triggers.
The worst time to discover a bug in your fallback logic is during the actual outage it was built for. A short, scheduled drill catches that gap while the stakes are low, instead of a few weeks after it stopped working.
Error recovery mistakes that backfire
- Retrying indefinitely with no cap, turning a transient failure into a request that never resolves either way.
- Building a fallback with no monitoring on how often it fires, so a persistent outage looks like normal operation.
- Treating every error the same way, retry everything, which wastes time on failures a retry will never fix.
What Good Looks Like
Solid error recovery design means retries are capped and backed off with jitter, fallbacks are tagged and monitored rather than silent, and failures that won't resolve on their own fail visibly instead of getting masked by logic meant for transient problems.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's the safest default for retrying a failed model request?
Exponential backoff with jitter and a hard cap on total retry attempts, reserved for genuinely transient failures like timeouts or rate limits. Immediate, uncapped retries from many clients at once are how a temporary slowdown turns into a full outage, since the retry traffic adds directly to the load causing the problem.
Should a fallback response look identical to a normal one?
To the caller, usually yes, since a degraded answer beats an error. Internally, no: tag fallback responses and alert when the fallback rate rises, or a persistent outage can run for hours looking like normal operation because every request still technically succeeded.
When should we fail a request instead of retrying or falling back?
When the failure isn't going to resolve itself: a malformed request, an authentication error, or anything where a retry or fallback would just mask a real bug. Failing visibly in those cases surfaces the actual problem instead of hiding it behind logic that was designed for transient failures, not permanent ones.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Designing Retry and Fallback Logic That Doesn't Make Outages Worse
How to design retries, backoff, and circuit breakers so error recovery logic protects a struggling dependency instead of overwhelming it further.
What Happens When a Tool Call Fails Mid-Task
A decision guide for designing fallback logic in an agent loop, so a single failed tool call degrades gracefully instead of derailing the whole task.
The Retry Logic That Makes Outages Worse, Not Better
How naive retry and fallback logic can amplify an outage instead of recovering from it, and the specific patterns that actually help.
Designing Retry Logic That Doesn't Make Things Worse
How to decide when a retry actually helps, why fixed intervals cause outages, and how to design fallback logic that degrades instead of breaking.
Designing Retry and Fallback Logic That Doesn't Undermine Your Access Controls
A checklist for building retry, idempotency, and fallback logic for zero trust APIs, plus the specific pitfalls that quietly weaken access control.
Where Retry Logic Quietly Drains Your Infrastructure Budget
How poorly designed error handling and retry logic turns a small outage into a large bill, and the specific patterns worth fixing first.