Why DNS Failover Isn't as Fast as You Think It Is
DNS failover sounds like it should be instant: a health check fails, the record changes, traffic moves. In practice it's one of the slower failover mechanisms available, and treating it like a fast one is how a five-minute outage turns into a much longer one.
Understanding where the delay actually comes from, and when anycast routing sidesteps it, changes how much you can rely on DNS alone during an incident.
What DNS failover can and can't do
A health-checked DNS record changes which IP address a resolver hands back the next time it looks your domain up. It does nothing for a connection that's already open, and it does nothing for a client or resolver that's still holding an older, cached answer. DNS failover moves future lookups, not current traffic.
That's a meaningful limitation for anything long-lived, a persistent connection, a mobile client that cached a resolution hours ago, a corporate resolver that ignores your stated cache time. For short-lived HTTP requests from clients that respect DNS caching properly, it works closer to as advertised.
The TTL problem nobody tests until an outage
The TTL you set on a record is a request, not a guarantee. Some resolvers and ISPs cache longer than you asked for, and some clients cache the result inside the application layer on top of whatever the OS resolver does. Setting a low TTL helps, but it doesn't make every resolver in the path honor it.
Anycast sidesteps the whole problem, because failover happens at the network routing layer instead of through DNS propagation. The same IP address is announced from multiple locations, and the network routes a client to whichever healthy location is closest, so there's no cached answer to wait out.
Health checks that catch a real failure
A health check that hits a shallow endpoint returning a fixed 200 response tells you the process is running, not that the service works. If that endpoint doesn't touch the database, the queue, or whatever dependency actually breaks user requests, your failover will sit there passing checks while real traffic fails.
Build the check to exercise something close to what a real request needs. It should fail when the thing that actually matters is down, and it should not flap on a single slow response, add a short window of consecutive failures before triggering failover, so one bad health check doesn't send traffic somewhere and back for no reason.
Anycast versus latency-based routing
Anycast announces one IP address from many network locations and lets routing decide which one answers, giving you failover measured in seconds because it happens below DNS entirely. It's what CDNs and some DNS providers run their own edge networks on.
Latency-based or geo DNS routing instead hands out different IP addresses depending on where the resolver is, still bound by whatever TTL the resolver actually respects. It's simpler to reason about and doesn't require owning network infrastructure, but a real failover through it moves at DNS speed, not network speed. Pick anycast when failover time matters most, and geo DNS when you mainly want to route by region and can tolerate a slower switch.
Testing the failover before a real outage does it for you
A failover setup that's never been exercised is a hypothesis, not a plan. Run a deliberate drill: fail the health check on purpose, or manually flip the record, and time how long it actually takes for traffic to move, not how long the runbook says it should take.
Watch what happens from a few different vantage points and resolvers, not just your own laptop, since that's the only way to catch a resolver in the wild that's ignoring your TTL. A drill that surfaces a slow or partial failover is worth far more before an incident than during one, and it's the only way to know whether the secondary region you're failing over to can actually absorb full production traffic instead of falling over itself the moment it starts receiving it.
Write down what the drill actually showed, the real time to full cutover, which vantage points lagged, and treat that number as your honest failover time until the next drill changes it. A runbook with an untested time-to-failover figure in it is a guess dressed up as a fact, and it's the kind of guess that gets relied on at exactly the wrong moment.
A useful failover drill follows these steps:
- Fail the health check on purpose, or flip the record by hand, so the failover happens when you choose instead of during a real incident.
- Time how long traffic actually takes to move, and record that measured number rather than trusting the figure written in the runbook.
- Check the result from several vantage points and resolvers, not just your own laptop, since caching behavior differs across ISPs and clients.
- Note which long-lived connections and application-layer caches ignored the change, because DNS failover only moves future lookups, not current traffic.
- Compare the measured time with what your outage tolerance allows, and consider anycast if the gap is too wide.
What Good Looks Like
Good failover means a real outage in one region routes traffic to a healthy one inside the time your users will tolerate, and you know that because you tested it, not because the runbook says it should work.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is anycast the same thing as a CDN?
No, though CDNs typically run on anycast networks. Anycast is a routing technique, announcing one IP from many locations. A CDN is a broader service built on top of that, adding caching, edge compute, and content delivery.
How low should my DNS TTL be?
Low enough to fail over in a time you can live with, but every lookup has a cost, so going extremely low mainly adds resolver query load without a proportional benefit once you're below what most resolvers will actually honor.
Does DNS failover replace a load balancer?
No. A load balancer distributes and reroutes traffic within a region at the network layer, often in milliseconds. DNS failover operates a layer above that, deciding which region or provider traffic goes to in the first place.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
DNS Failover: Why It's Slower Than It Looks
A short TTL doesn't guarantee fast failover; some resolvers ignore it. What DNS failover actually controls, and when anycast is worth the jump.
Active-Active or Active-Passive: Choosing a DNS Failover Setup
A decision guide for choosing between active-active and active-passive DNS failover, and the health check design that makes either one work.
What Happens to Your Traffic During a DNS Failover, Exactly?
A Q&A walkthrough of what actually happens during a DNS-based failover: TTL behavior, health checks, and why some clients don't fail over at all.
DNS Failover: What Actually Happens When a Region Dies
Health-check failover vs. anycast routing, why TTL is the hidden variable, and how to actually test failover instead of trusting the runbook.
Why DNS Failover Alone Won't Save You During a Regional Outage
What DNS-based failover actually does and doesn't protect against, including TTL and caching pitfalls, and what to pair it with for a real multi-region setup.
Anycast DNS Failover: What It Actually Buys You
How anycast DNS routing differs from a simple health-check failover, what it protects against, and where it can't replace application redundancy.