What Happens to Your Traffic During a DNS Failover, Exactly?
DNS-based failover looks simple in a diagram: a health check fails, a record updates, traffic moves to the healthy region. What actually happens is messier, governed by TTLs, resolver caching behavior outside your control, and health check logic that can flap or lag in ways that make a real failover feel much slower than the diagram suggested.
Here are the questions engineering teams actually ask the first time they design or rely on a DNS failover setup, answered directly.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
If our TTL is 60 seconds, does failover really take 60 seconds?
Sixty seconds is the best case, not a guarantee. A resolver isn't obligated to honor your TTL exactly, some ISP and corporate resolvers cache more aggressively than the record specifies, and any client or application with its own DNS cache, a connection pool that resolved once and reused the connection, won't fail over until that cache expires or the connection drops. Plan for a real-world failover window meaningfully longer than your TTL, and test it under realistic conditions rather than trusting the TTL number alone.
How does the health check actually decide a region is down?
Most DNS failover services run their own health checks from multiple external locations against an endpoint you specify, and mark a region unhealthy after a configurable number of consecutive failures. A single check location is fragile, since a network issue between the checker and your region can look identical to your region actually being down. Use checks from multiple independent locations and require agreement across a majority of them before failing over, or a normal transient blip between one checker and one region can trigger an unnecessary failover.
What's the difference between failover routing and weighted or latency-based routing?
Failover routing is binary: primary is used exclusively until it's marked unhealthy, then traffic moves entirely to secondary. Weighted routing splits traffic by a fixed percentage regardless of health, useful for canary testing but not a failover mechanism on its own. Latency-based routing sends each client to whichever healthy region responds fastest for them, which is genuinely useful for multi-region active-active setups but adds complexity that a simple primary-secondary failover setup doesn't need. Most teams starting out should use plain failover routing and add complexity only once they have a specific reason to.
Why did some of your customers not fail over even though the health check clearly caught the outage?
This is almost always a caching resolver or client-side DNS cache holding onto the old record past its TTL, or an application that resolved the hostname once at startup and never re-resolves it during the process lifetime. Test failover with a long-running client process specifically, not just a fresh DNS lookup from a terminal, since a fresh lookup will correctly show the new record while a live, already-connected client may not notice for a much longer window.
How do we actually test this before we need it in a real incident?
Run a scheduled game day where you intentionally fail the health check for the primary region, not just simulate it in a test environment, and measure real client-observed failover time end to end, including from a long-running connection. An untested failover plan that turns out to take far longer than the team believed is exactly the kind of gap worth finding on a scheduled game day rather than during a real outage. Weigh what you find against the actual downtime budget the business has signed up for: a commitment to three-nines availability leaves an annual allowance well under half a day, which a single slow failover can eat into meaningfully on its own1.
Run the failover game day like this:
- Schedule a game day and intentionally fail the primary region's health check, instead of only simulating the failure in a test environment.
- Measure client-observed failover time end to end, including from a long-running client process that resolved the hostname once at startup.
- Confirm health checks run from multiple independent locations, so one checker's network problem can't trigger a false failover.
- Look for connection pools and application-level DNS caches that keep sending traffic to the old region after the record has updated.
- Compare the measured failover window with your TTL, and plan around the longer real-world number.
A worked example: the connection pool that never noticed the outage
Say your primary region goes down and the DNS record updates within seconds, well inside the TTL you configured. A backend service with a long-lived database connection pool, opened at process startup and never re-resolved, keeps sending traffic to the now-unreachable primary anyway, since the pool holds an already-established connection rather than performing a fresh lookup on every request. The service layer looks correctly failed-over from a DNS perspective while a specific internal client quietly keeps erroring until it happens to restart. This is why testing failover from a fresh terminal lookup alone gives a false sense of confidence.
What Good Looks Like
A tested DNS failover setup uses multi-location health checks requiring majority agreement, accounts for real-world caching behavior beyond the stated TTL, and gets validated on a scheduled game day rather than assumed to work.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Should we set our DNS TTL as low as possible to speed up failover?
There's a real tradeoff. A very low TTL, under 30 seconds, does reduce failover time for well-behaved clients but increases the query load on your DNS provider and doesn't help against resolvers or clients that cache more aggressively than the TTL specifies regardless.
Do we need health checks from more than one location?
Strongly recommend it. A single check location can't distinguish your region actually being down from a network problem specific to that one checker, which is a common cause of an unnecessary, disruptive failover triggered by a false positive.
Is DNS failover enough on its own, or do we need a load balancer too?
DNS failover works well for routing between regions or entirely separate infrastructure stacks. Within a single region, a load balancer's health checks typically react faster and more precisely, since they aren't subject to DNS caching delays at all. Most resilient architectures use both, at different layers.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
DNS Failover: Why It's Slower Than It Looks
A short TTL doesn't guarantee fast failover; some resolvers ignore it. What DNS failover actually controls, and when anycast is worth the jump.
Why DNS Failover Alone Won't Save You During a Regional Outage
What DNS-based failover actually does and doesn't protect against, including TTL and caching pitfalls, and what to pair it with for a real multi-region setup.
DNS Failover: What Actually Happens When a Region Dies
Health-check failover vs. anycast routing, why TTL is the hidden variable, and how to actually test failover instead of trusting the runbook.
Anycast DNS Failover: What It Actually Buys You
How anycast DNS routing differs from a simple health-check failover, what it protects against, and where it can't replace application redundancy.
DNS Failover Versus a Load Balancer: What Each One Actually Fixes
A comparison of DNS-based failover and load balancer failover for surviving a regional or full-service outage, and where each approach falls short on its own.
Setting Up DNS Failover That Actually Fails Over When It Matters
How DNS based failover actually works, where TTLs and caching quietly undermine it, and how to test a failover policy before you need it during an outage.