Enterprise DevSecOps & Automated CompliancePlaybook3 min readUpdated September 2026

What Happens to Your Traffic During a DNS Failover, Exactly?

DNS-based failover looks simple in a diagram: a health check fails, a record updates, traffic moves to the healthy region. What actually happens is messier, governed by TTLs, resolver caching behavior outside your control, and health check logic that can flap or lag in ways that make a real failover feel much slower than the diagram suggested.

Here are the questions engineering teams actually ask the first time they design or rely on a DNS failover setup, answered directly.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

If our TTL is 60 seconds, does failover really take 60 seconds?

Sixty seconds is the best case, not a guarantee. A resolver isn't obligated to honor your TTL exactly, some ISP and corporate resolvers cache more aggressively than the record specifies, and any client or application with its own DNS cache, a connection pool that resolved once and reused the connection, won't fail over until that cache expires or the connection drops. Plan for a real-world failover window meaningfully longer than your TTL, and test it under realistic conditions rather than trusting the TTL number alone.

How does the health check actually decide a region is down?

Most DNS failover services run their own health checks from multiple external locations against an endpoint you specify, and mark a region unhealthy after a configurable number of consecutive failures. A single check location is fragile, since a network issue between the checker and your region can look identical to your region actually being down. Use checks from multiple independent locations and require agreement across a majority of them before failing over, or a normal transient blip between one checker and one region can trigger an unnecessary failover.

What's the difference between failover routing and weighted or latency-based routing?

Failover routing is binary: primary is used exclusively until it's marked unhealthy, then traffic moves entirely to secondary. Weighted routing splits traffic by a fixed percentage regardless of health, useful for canary testing but not a failover mechanism on its own. Latency-based routing sends each client to whichever healthy region responds fastest for them, which is genuinely useful for multi-region active-active setups but adds complexity that a simple primary-secondary failover setup doesn't need. Most teams starting out should use plain failover routing and add complexity only once they have a specific reason to.

Why did some of your customers not fail over even though the health check clearly caught the outage?

This is almost always a caching resolver or client-side DNS cache holding onto the old record past its TTL, or an application that resolved the hostname once at startup and never re-resolves it during the process lifetime. Test failover with a long-running client process specifically, not just a fresh DNS lookup from a terminal, since a fresh lookup will correctly show the new record while a live, already-connected client may not notice for a much longer window.

How do we actually test this before we need it in a real incident?

Run a scheduled game day where you intentionally fail the health check for the primary region, not just simulate it in a test environment, and measure real client-observed failover time end to end, including from a long-running connection. An untested failover plan that turns out to take far longer than the team believed is exactly the kind of gap worth finding on a scheduled game day rather than during a real outage. Weigh what you find against the actual downtime budget the business has signed up for: a commitment to three-nines availability leaves an annual allowance well under half a day, which a single slow failover can eat into meaningfully on its own1.

Run the failover game day like this:

  1. Schedule a game day and intentionally fail the primary region's health check, instead of only simulating the failure in a test environment.
  2. Measure client-observed failover time end to end, including from a long-running client process that resolved the hostname once at startup.
  3. Confirm health checks run from multiple independent locations, so one checker's network problem can't trigger a false failover.
  4. Look for connection pools and application-level DNS caches that keep sending traffic to the old region after the record has updated.
  5. Compare the measured failover window with your TTL, and plan around the longer real-world number.

A worked example: the connection pool that never noticed the outage

Say your primary region goes down and the DNS record updates within seconds, well inside the TTL you configured. A backend service with a long-lived database connection pool, opened at process startup and never re-resolved, keeps sending traffic to the now-unreachable primary anyway, since the pool holds an already-established connection rather than performing a fresh lookup on every request. The service layer looks correctly failed-over from a DNS perspective while a specific internal client quietly keeps erroring until it happens to restart. This is why testing failover from a fresh terminal lookup alone gives a false sense of confidence.

Executive Capability Standard

What Good Looks Like

A tested DNS failover setup uses multi-location health checks requiring majority agreement, accounts for real-world caching behavior beyond the stated TTL, and gets validated on a scheduled game day rather than assumed to work.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your DNS provider's documentation on exactly how their health checks and failover routing decide to switch records.
2. Do Manually:Manually test failover once by intentionally failing a health check and measuring real client-observed recovery time.
3. Delegate:Assign an engineer to own DNS failover configuration and the game-day testing schedule as a standing responsibility.
4. Automate:Automate a recurring game day that fails the primary region's health check on a schedule and reports measured failover time.
5. Buy:Use a managed DNS provider with built-in multi-location health checking rather than building your own health check infrastructure from scratch.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Deel

Global contractor and EOR employment platform across 150+ countries.

Visit Deel→
Rippling

Unified global HR, payroll, and IT equipment deployment.

Visit Rippling→

Frequently Asked Questions

Should we set our DNS TTL as low as possible to speed up failover?

There's a real tradeoff. A very low TTL, under 30 seconds, does reduce failover time for well-behaved clients but increases the query load on your DNS provider and doesn't help against resolvers or clients that cache more aggressively than the TTL specifies regardless.

Do we need health checks from more than one location?

Strongly recommend it. A single check location can't distinguish your region actually being down from a network problem specific to that one checker, which is a common cause of an unnecessary, disruptive failover triggered by a false positive.

Is DNS failover enough on its own, or do we need a load balancer too?

DNS failover works well for routing between regions or entirely separate infrastructure stacks. Within a single region, a load balancer's health checks typically react faster and more precisely, since they aren't subject to DNS caching delays at all. Most resilient architectures use both, at different layers.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides