Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

DNS Failover: Why It's Slower Than It Looks

A primary region goes down, DNS failover is supposedly configured, and customers are still hitting the dead region twenty minutes later, because a resolver somewhere cached the old answer well past the TTL that was set, or never honored it in the first place.

DNS failover is a real tool, but it's a request to resolvers to behave a certain way, not a guarantee. Knowing where its limits are keeps you from planning around a recovery time it can't actually deliver.

What DNS Failover Actually Controls, and What It Doesn't

DNS controls which IP address a new lookup resolves to. It does nothing about a connection or session a client already established before the failure, and it can't force a resolver to honor a short TTL; some ISP and corporate resolvers cache well past whatever value you configured, regardless of what the record says.

TTL Is a Request, Not a Guarantee

Lowering TTL reduces propagation time for most resolvers that respect it, but some don't, and you generally have no visibility into which of your users are sitting behind one that ignores it. Plan around your worst observed resolver behavior, not your configured TTL value; the two can differ by a wide margin during a real incident.

Health Check Design Determines How Fast You Actually Fail Over

A health check that only pings a load balancer, not the actual application behind it, can report healthy while the application itself is failing every real request. Point health checks at something that reflects genuine application health, and set a sensible interval and failure threshold so a single transient blip doesn't trigger a failover nobody actually needed.

Run the check from more than one vantage point if your provider supports it. A health check running from a single location can itself lose connectivity to a perfectly healthy region, triggering a failover based on a problem with the checker rather than the service being checked, which is its own kind of false positive worth designing around.

Anycast as the Faster Alternative

Anycast routing announces the same IP address from multiple locations and lets network-layer routing send each client to the nearest healthy one, sidestepping the DNS caching problem entirely because no new lookup is required to shift traffic. The tradeoff is that it needs a network provider that actually supports it and represents a bigger infrastructure commitment than DNS-based failover.

Most small and mid-sized teams start with DNS failover, since it's simpler to set up and reason about, and only move to anycast once DNS's propagation delay has caused a real incident that a lower TTL alone couldn't fix.

The two approaches aren't mutually exclusive. A common middle ground is anycast for the entry point that serves traffic day to day, with DNS-based failover kept as a fallback path for the rare case that even the anycast network itself needs to route around a whole provider's regional outage.

A Realistic Failover Checklist

  • Set TTL as low as your DNS provider allows for records you might need to fail over quickly, while accepting that some resolvers still won't respect it.
  • Point health checks at something that reflects real application health, not just infrastructure being reachable.
  • Test an actual failover on a schedule, not only once at initial setup, since DNS and health check behavior can drift as infrastructure changes.
  • Know your real observed failover time from an actual test, not the number implied by your TTL setting.

What Clients Experience During the Gap

Even a fast, well-tested failover leaves a window where some share of users are still resolving the old address, and their requests either time out or hit a region that's actively being drained. Design the client side for that gap: reasonable connection timeouts instead of ones long enough to leave a user staring at a spinner, and retry logic on the client that doesn't just hammer the same failed address in a tight loop.

Communicate the realistic gap internally too, to whoever owns customer-facing status updates. Telling customers a failover completed the moment DNS changed, when a meaningful share of them are still resolving the old address for several more minutes, sets an expectation the actual system can't meet.

Executive Capability Standard

What Good Looks Like

Good DNS failover means you know your real observed failover time from an actual tested failover, not your configured TTL, and your health checks reflect genuine application health rather than just infrastructure reachability.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Run a real test failover and measure how long it actually takes for traffic to shift, rather than trusting your configured TTL as the answer.
2. Do Manually:Manually review your health check configuration to confirm it hits something that reflects real application health, not just a load balancer ping.
3. Delegate:Assign one engineer to own a recurring failover test schedule so drift gets caught before an actual incident.
4. Automate:Automate a scheduled failover drill with alerting if the observed recovery time exceeds your target.
5. Buy:Bring in a network specialist to evaluate anycast if DNS propagation delay has already caused a real incident.

How to Get Started

Frequently Asked Questions

Does lowering TTL guarantee fast failover?

No. It reduces propagation time for resolvers that respect it, but some ISP and corporate resolvers cache past the configured value regardless. Plan your real recovery time estimate around the worst resolver behavior you've actually observed in testing, not around the TTL number itself.

What's the practical difference between DNS failover and anycast?

DNS failover requires a new lookup to point clients at a healthy location, which is subject to caching delays outside your control. Anycast announces the same IP from multiple locations and lets network routing send traffic to the nearest healthy one, avoiding the DNS caching problem entirely, at the cost of needing a provider that supports it.

How often should I actually test failover?

On a regular schedule, not just once at initial setup. DNS behavior, health check configuration, and the application itself all drift as infrastructure changes, so a failover that worked cleanly six months ago isn't a guarantee it still works today without a fresh test.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides