Designing Retry and Fallback Logic That Doesn't Undermine Your Access Controls
Retry and fallback logic exists to make a system resilient, but built carelessly it can quietly undermine the access control you've built everywhere else: a retry that doesn't check permissions again, a fallback path that skips the auth layer entirely because it's "just for emergencies." Resilience and access control need to be designed together, not layered on separately.
This checklist covers the specific places that pattern breaks down.
Why must a retried request pass authorization again?
A request that failed and gets automatically retried should go through the exact same authorization path as the original attempt, not a shortcut that assumes "we already checked this." Permissions can change between the first attempt and the retry, especially for a retry that's delayed by a backoff window, and skipping the recheck means an action that should now be denied might succeed anyway.
Test this specifically: revoke a permission mid-retry-sequence in a test environment and confirm the retry is denied, not just that the retry eventually succeeds when everything's healthy. This is easy to get wrong in frameworks that cache the result of a middleware check across an internal retry loop for performance reasons, so trace the actual code path rather than assuming the framework does the right thing by default.
Build idempotency keys in from the start, not after a duplicate-action incident
A network timeout doesn't tell the caller whether the request actually succeeded server-side before the connection dropped, so a naive retry can execute the same action twice, charging a customer twice, sending a notification twice. An idempotency key, generated once by the caller and reused across retries, lets your API recognize a retried request as the same operation rather than a new one.
This is far easier to design in from the start than to retrofit once a real duplicate-action incident has already happened and every downstream system has to be checked for whether it handled the duplicate gracefully. Store the idempotency key alongside the original response, not just a flag that the request happened, so a retry can return the exact original result instead of a generic "already processed" error that leaves the caller unsure what actually happened the first time.
For example, a payment endpoint times out after the charge has been created but before the response reaches the client. Without an idempotency key, the client's retry creates a second charge. With one, the server recognizes the retry, finds the stored result of the first attempt and returns it unchanged, so the client learns what actually happened. The remaining design question is how long to keep stored keys, which should be at least as long as the longest retry window any client is allowed to use.
What should a fallback path do when authorization is unavailable?
When the primary authorization service is slow or unavailable, the tempting fallback is "allow the request through" so the system stays available, which is exactly backwards from a zero trust standpoint: unavailable authorization should fail closed, denying the request, not fail open and grant access nobody actually verified.
Build your actual availability strategy around keeping the authorization service itself highly available, redundant instances, a fast local cache with a short TTL, rather than around a bypass path that undermines the entire model the moment it's needed. Audit any existing fallback code specifically for this pattern, since it tends to get added quietly during an incident, under pressure, by someone trying to restore availability fast, and then never gets removed once the immediate crisis passes.
Give circuit breakers a security-aware trip condition, not just a latency one
A circuit breaker that trips only on latency or error rate will happily keep routing traffic to a downstream service that's returning successful-looking but incorrect authorization decisions, since that failure mode doesn't show up as errors or slowness. Where practical, add a sanity check, a known test identity that should always be denied a specific action, and trip the breaker if that check itself starts passing.
An allowed downtime budget of roughly nine hours a year at 99.9% availability leaves very little room for a circuit breaker configuration that's slow to react to a real degradation1, so tune the trip conditions deliberately rather than leaving default settings in place. Test the breaker's own behavior periodically too, deliberately degrading the downstream service in staging to confirm it trips and recovers the way you expect, rather than assuming the configuration works because it hasn't been triggered by a real incident yet.
Check your retry and fallback design against these points:
- Retried requests go through the same authorization path as the original attempt, with permissions rechecked after any backoff delay.
- Callers generate an idempotency key once and reuse it across retries, and you store it alongside the original response.
- Fallback code fails closed when the authorization service is unreachable, and any emergency bypass added during an incident is removed afterward.
- Circuit breakers trip on security signals as well as latency, for example when a known test identity that should be denied starts passing.
What Good Looks Like
Good practice means every retry runs through the same authorization check as the original request, idempotency keys prevent duplicate execution, and any fallback for an unavailable authorization service fails closed rather than granting unchecked access.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Should a failed authorization check ever be retried automatically?
A failed network call to the authorization service is worth retrying with backoff. A failed authorization decision, meaning the check succeeded and returned denied, should not be retried automatically, since retrying a denial repeatedly looks like exactly the kind of probing behavior you want to detect, not paper over.
What should happen if the authorization service itself is down?
Fail closed: deny the request rather than allowing it through. Availability of the underlying feature matters less than not granting access nobody actually verified, and the better fix is making the authorization service itself more resilient, not building a bypass around it.
How long should an idempotency key stay valid?
Long enough to cover your realistic retry window plus some margin, commonly 24 hours, after which a repeated request with the same key can reasonably be treated as a new operation. Shorter windows risk a legitimate delayed retry being treated as a fresh, potentially duplicate action.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Designing Retry and Fallback Logic That Doesn't Make Outages Worse
How to design retries, backoff, and circuit breakers so error recovery logic protects a struggling dependency instead of overwhelming it further.
The Retry Logic That Makes Outages Worse, Not Better
How naive retry and fallback logic can amplify an outage instead of recovering from it, and the specific patterns that actually help.
Designing Retry Logic That Doesn't Make Things Worse
How to decide when a retry actually helps, why fixed intervals cause outages, and how to design fallback logic that degrades instead of breaking.
Where Retry Logic Quietly Drains Your Infrastructure Budget
How poorly designed error handling and retry logic turns a small outage into a large bill, and the specific patterns worth fixing first.
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Designing Fallback Logic That Doesn't Make Things Worse
How to design retry, fallback, and fail-visibly logic for AI model serving without causing a retry storm during an outage.