Why DNS Failover Alone Won't Save You During a Regional Outage
DNS failover looks like a clean solution to regional outages: monitor your primary region's health, and if it goes down, update DNS to point traffic at a healthy region instead. In practice, DNS was never designed to fail over instantly, and treating it as your only safeguard leaves a gap that shows up at exactly the wrong moment.
Here's what DNS failover actually does, where it falls short, and what closes the gap.
What DNS Failover Actually Does
A DNS failover setup runs health checks against your primary endpoint, and when those checks fail, it updates the DNS record to return a different IP address, one pointing at your backup region. Any client that does a fresh DNS lookup after the change gets routed correctly. This works well for new connections and clients with short-lived DNS caches, and it's a legitimate, low-cost first layer of protection for a lot of teams.
Where TTL Quietly Undermines the Whole Plan
The failover only helps a client once its cached DNS record expires, and that expiry is controlled by your record's time-to-live setting, which you control, and by every resolver and client along the way, which you don't. A TTL of an hour means some meaningful share of your traffic keeps hitting the dead region for up to an hour after you've already updated DNS. Worse, some ISP resolvers and corporate networks cache more aggressively than your TTL suggests, ignoring it outright in some cases. Lowering your TTL to something short, a minute or less, reduces this window but doesn't eliminate it, and very low TTLs come with their own cost in extra DNS query volume.
The Detection Delay Nobody Accounts For
Before DNS even updates, your health check has to actually notice the primary region is down, which takes some number of consecutive failed checks to avoid false positives from a single blip. Add that detection window to the TTL-driven propagation delay, and a DNS-only failover plan commonly takes several minutes from actual outage to full traffic recovery, sometimes longer. For a team with a tight availability target, that gap alone can consume a meaningful share of a whole year's downtime budget in one incident1.
What Actually Closes the Gap
Anycast routing, where the same IP address is announced from multiple regions at the network layer and traffic is routed to the nearest healthy one automatically, sidesteps the DNS propagation delay entirely, since there's no record to wait for clients to re-fetch. This is genuinely more complex to operate than DNS failover and usually requires either your own BGP setup or a CDN and networking provider that offers it as a managed feature. For most teams, a practical middle ground is a low TTL combined with a load balancer or global traffic manager that can shift traffic faster than raw DNS propagation alone, reserving full anycast for the small set of services where the extra minutes genuinely matter.
Testing This Before You Need It for Real
A failover setup that's never actually been triggered outside a diagram is not a tested failover setup. Run a real regional failover drill, actually take the primary region offline in a controlled window, and time the full recovery from detection to the last client resolving correctly, not just from when DNS updates. This is the only way to know whether your actual TTL, your actual resolver behavior, and your actual health check timing add up to a recovery time you can live with, rather than one you're hoping is fine.
A realistic failover drill follows these steps:
- Schedule a controlled window and take the primary region offline on purpose, rather than trusting a diagram.
- Time how long your health checks take to detect the failure, since several consecutive failed checks are needed to avoid false positives.
- Time DNS propagation against your real TTL, and keep timing until the last client resolves to the healthy region.
- Check clients that cache DNS longer than your TTL, such as mobile apps, and compare the total time to your recovery target.
- If the total is too slow, add a layer such as anycast routing, often through a CDN or DNS provider that offers it as a managed feature.
A Worked Example: The Mobile App That Never Failed Over
Say your web traffic recovers cleanly within a couple of minutes during a drill, but your mobile app's traffic doesn't move at all for over an hour. The likely cause: the app's HTTP client, or an underlying OS networking layer, cached the DNS resolution far longer than your record's TTL specified, something browsers and mobile networking stacks are both known to do more aggressively than a raw TTL suggests. This is exactly the kind of gap a drill catches and a diagram never will, and it usually means the mobile client needs its own retry-and-reresolve logic rather than relying on DNS alone to notice the change.
What Good Looks Like
A real multi-region failover setup accounts for detection time and DNS propagation delay together, has been tested with an actual regional outage drill, and uses anycast or a traffic manager for any service where minutes of delay would blow the downtime budget.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What's a reasonable DNS TTL for a failover-capable record?
Sixty seconds is a common choice, balancing faster failover against the extra query volume a very low TTL generates. Going lower, into single-digit seconds, rarely buys much more real-world improvement, since resolver caching behavior outside your control becomes the bigger factor at that point.
Do all DNS resolvers actually respect the TTL we set?
Most do, but not all, and you have no way to enforce it. Some ISP and corporate resolvers cache longer than the TTL specifies for their own performance reasons, which is exactly why DNS failover should be treated as a first layer, not a guarantee, for anything with a strict recovery time requirement.
Is anycast routing worth the added complexity for a smaller company?
Usually not on your own, since running your own anycast network requires BGP expertise most small teams don't have in-house. A CDN or DNS provider that offers anycast as a managed feature can get you most of the benefit without the operational burden, and it's worth it specifically for the services where a few minutes of DNS propagation delay would be a real problem.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
DNS Failover: Why It's Slower Than It Looks
A short TTL doesn't guarantee fast failover; some resolvers ignore it. What DNS failover actually controls, and when anycast is worth the jump.
What Happens to Your Traffic During a DNS Failover, Exactly?
A Q&A walkthrough of what actually happens during a DNS-based failover: TTL behavior, health checks, and why some clients don't fail over at all.
DNS Failover: What Actually Happens When a Region Dies
Health-check failover vs. anycast routing, why TTL is the hidden variable, and how to actually test failover instead of trusting the runbook.
Anycast DNS Failover: What It Actually Buys You
How anycast DNS routing differs from a simple health-check failover, what it protects against, and where it can't replace application redundancy.
DNS Failover Versus a Load Balancer: What Each One Actually Fixes
A comparison of DNS-based failover and load balancer failover for surviving a regional or full-service outage, and where each approach falls short on its own.
Setting Up DNS Failover That Actually Fails Over When It Matters
How DNS based failover actually works, where TTLs and caching quietly undermine it, and how to test a failover policy before you need it during an outage.