DNS Failover: What Actually Happens When a Region Dies
A DNS failover plan that's never been tested is a hope, not a plan. The gap between "our DNS provider supports health-check-based failover" and "our service actually recovers within an acceptable window when a region goes dark" is usually TTL caching behavior nobody accounted for, and client-side DNS caching that ignores the TTL entirely.
Here's how to decide between health-check failover and anycast, and why testing matters more than either choice.
How Does Health-Check Failover Work, and What Limits It?
A DNS provider that monitors an endpoint's health and swaps the returned IP when it fails is straightforward to set up and understood by most teams. Its ceiling is TTL: a low TTL means faster failover propagation but more DNS query volume and less benefit from caching; a high TTL means resolvers and clients hold onto the old, dead IP longer after the health check has already failed over. Many clients and some resolvers also ignore TTL entirely and cache longer than instructed, which means even a well-tuned TTL doesn't guarantee every client fails over on your expected timeline.
Anycast: Failover Without Waiting for DNS to Propagate At All
Anycast routes traffic to the topologically nearest healthy instance advertising the same IP address at the network layer, which means failover doesn't depend on DNS propagation or client caching behavior at all, a dead region simply stops being routed to by the network. This is a heavier infrastructure lift, it typically requires your own IP address space and BGP relationships or a provider that offers anycast as a managed service, and it's usually the right investment once DNS TTL-bound failover has proven too slow or unreliable for your actual recovery time requirements.
Multi-Provider Redundancy Protects Against the DNS Provider Itself
Health-check failover and anycast both assume your DNS provider itself stays up; a DNS provider outage, which has happened to every major provider at some point, takes down failover along with everything else if you're single-sourced. Running authoritative DNS across two providers, with your zone kept in sync between them, protects against this specific failure mode, though it adds operational complexity, particularly around keeping records consistent and coordinating any changes across both providers without drift.
Why Isn't an Untested Failover a Real Plan?
Run an actual failover drill in a lower environment: take the primary region offline the way a real outage would, and time how long it actually takes for traffic to shift, including the client-side caching behavior you can't fully control from your DNS configuration alone. This regularly surfaces gaps a written runbook doesn't: a health check with too generous a failure threshold, a CDN layer caching DNS resolution longer than the origin's TTL, a client library that caches DNS results for its own process lifetime regardless of TTL.
A useful failover drill covers these points:
- Take the primary region offline in a lower environment the way a real outage would.
- Time how long traffic actually takes to shift, including client and CDN caching you can't fully control.
- Test internal resolution too, since split-horizon DNS can still point at the dead region.
- Compare the measured recovery time against your availability target, then tune TTL and health checks.
- Update the status page and incident channels even when the failover is clean.
Availability Targets Should Drive the TTL and Health-Check Tuning
The gap between a 99.9% and a 99.99% availability target isn't a rounding error, it's the difference between roughly nine hours and roughly one hour of allowed downtime across a full year1. A health-check interval and failure threshold tuned for a looser target will burn through a tighter target's entire annual downtime budget in a single slow-to-detect incident, so pick your TTL and health-check cadence starting from the actual availability target you've committed to, not a default your DNS provider ships with.
Split-Horizon DNS Adds a Layer Failover Drills Often Miss
If internal services resolve your domain differently than external clients do, through split-horizon DNS or an internal resolver override, a failover drill that only tests the external path can pass while an internal dependency still points at the dead region. Map out every place your domain resolves differently for internal versus external traffic before you trust a single drill's result as proof the whole system fails over correctly, since it's common for an internal service's DNS configuration to drift out of sync with the external-facing failover setup without anyone noticing until it matters.
Communicating a Failover Event Beyond the Immediate Fix
A successful automated failover still means your primary region is down, and status page updates, internal incident channels, and any customer-facing communication should reflect that even if end users never noticed a service interruption. Treat a clean failover as a real incident worth a lightweight postmortem, not just a success story, since the same health check and threshold tuning that made the failover fast this time is the thing most likely to have drifted by the next time it's actually needed.
What Good Looks Like
Your failover should survive losing an entire region without anyone touching a DNS record by hand in the middle of an incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How low should our TTL be for a service that needs fast failover?
Short enough that most resolvers respect it within your recovery time target, commonly sixty seconds or less for services with strict failover requirements, balanced against the extra query volume a low TTL generates. Test actual propagation time in a drill rather than assuming the configured TTL is what clients will honor.
Is anycast worth it for a company our size?
It depends on your actual recovery time requirement and whether DNS TTL-bound failover has already proven too slow in a real drill. Anycast is a meaningful infrastructure investment; many teams are well served by aggressive health-check failover with multi-provider DNS redundancy before reaching for anycast specifically.
What's the most common thing a failover drill uncovers that the runbook missed?
Client or CDN-layer DNS caching that ignores or outlasts the configured TTL. Teams often discover their actual failover time is several times longer than the TTL alone would suggest, because something in the request path is caching DNS results independently.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
DNS Failover: Why It's Slower Than It Looks
A short TTL doesn't guarantee fast failover; some resolvers ignore it. What DNS failover actually controls, and when anycast is worth the jump.
What Happens to Your Traffic During a DNS Failover, Exactly?
A Q&A walkthrough of what actually happens during a DNS-based failover: TTL behavior, health checks, and why some clients don't fail over at all.
Setting Up DNS Failover That Actually Fails Over When It Matters
How DNS based failover actually works, where TTLs and caching quietly undermine it, and how to test a failover policy before you need it during an outage.
Why DNS Failover Alone Won't Save You During a Regional Outage
What DNS-based failover actually does and doesn't protect against, including TTL and caching pitfalls, and what to pair it with for a real multi-region setup.
Anycast DNS Failover: What It Actually Buys You
How anycast DNS routing differs from a simple health-check failover, what it protects against, and where it can't replace application redundancy.
DNS Failover Versus a Load Balancer: What Each One Actually Fixes
A comparison of DNS-based failover and load balancer failover for surviving a regional or full-service outage, and where each approach falls short on its own.