DNS Failover Versus a Load Balancer: What Each One Actually Fixes
DNS failover and load balancer failover fix different problems: DNS failover reroutes traffic across regions, while a load balancer reroutes it between instances inside one region. They work at different layers with different speeds, and treating them as interchangeable produces failover plans that look solid on paper but stall during a real outage.
What DNS failover actually does
DNS failover changes which IP address a domain name resolves to, based on health checks against each candidate endpoint, so new connections get routed to a healthy region once the unhealthy one is detected. Its core limitation is caching: resolvers and clients cache a DNS answer for a duration you set, called the time to live, and any client holding an old cached answer will keep connecting to the dead endpoint until that cache expires, regardless of how fast your health check detected the failure. This is why a failover strategy that only changes DNS, without accounting for caching, tends to look instant on a status dashboard while still failing for a meaningful share of real users.
What load balancer failover actually does
A load balancer sitting in front of multiple backend instances or regions can detect an unhealthy target and stop routing new traffic to it within seconds, without the client needing to do anything or re-resolve anything, since the client is only ever talking to the load balancer's own stable address. Its limitation is scope: a single load balancer typically operates within one region or one cloud provider's network, so it doesn't help when the failure is the load balancer's own region going dark entirely. Confirm which failure modes your specific load balancer setup actually covers before assuming it protects against a full regional event, since that assumption is exactly where the gap tends to hide.
Where each one falls short alone
DNS failover alone means every client is exposed to its own cache's stale answer for up to the full time-to-live window, which can be minutes, during exactly the period when fast failover matters most. Load balancer failover alone means you have no answer for a full regional outage that takes the load balancer itself down along with everything behind it. Neither one, used by itself, actually covers the full range of failures a production system needs to survive, and a failover plan built around only one of them will look complete right up until the specific failure it doesn't cover actually happens.
A quick decision rule helps here. For each failure you want to survive, ask whether it takes out a single instance or an entire region. A single instance is a load balancer problem, and its health checks should handle it within seconds. An entire region, including the load balancer itself, is a DNS problem, and recovery will be bounded by how long clients cache the old answer. For example, if a plan says a regional outage recovers in about a minute, check whether that number comes from the health check or from the time to live clients actually honor. Write the answer next to each failure mode in your runbook, so the expected recovery time reflects what users will experience rather than what a dashboard reports.
The pattern that covers both layers
Use a load balancer within each region for fast, sub-second failover between instances in that region, and DNS failover across regions for the case where an entire region, including its load balancer, goes down. Keep the DNS time-to-live short enough that a real cross-region failover completes in a reasonable window, while accepting that it will never be as fast as the load balancer's own instant, in-region failover, because that's a fundamentally different mechanism with a fundamentally different floor on speed. Document which layer is expected to handle which class of failure, so an on-call engineer isn't left guessing mid-incident about why traffic hasn't shifted yet.
A failover plan that covers both layers usually includes these pieces:
- Run a load balancer inside each region, so unhealthy instances stop receiving new traffic within seconds and no client has to re-resolve DNS.
- Configure DNS failover across regions, backed by health checks, for the case where an entire region and its load balancer go dark.
- Set the DNS time to live low enough that a cross-region failover finishes in a window your business can tolerate, and lower it before an outage rather than during one.
- Write down which layer handles which class of failure, so an on-call engineer knows why traffic has or hasn't shifted yet.
- Run a failover drill and measure the recovery time clients actually experience, since cached DNS answers can outlast the configured value.
A worked example: the failover that took twelve minutes longer than planned
Say a company configures DNS failover with a lengthy default time-to-live left over from before anyone thought carefully about failover speed, then tests a regional outage and finds clients still hitting the dead region well after the health check correctly detected the failure and updated the DNS record. The health check worked. The DNS record updated. Cached resolvers around the internet simply hadn't caught up yet, because nothing about a health check can force a client's own DNS cache to expire early. Shortening the time-to-live ahead of time, before an outage, is the only real fix for this specific gap.
What Good Looks Like
A resilient failover setup combines fast, in-region load balancer failover for individual instance failures with cross-region DNS failover, tuned to a deliberately short time-to-live, for the case of a full regional outage.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How short should our DNS time-to-live be for failover purposes?
Short enough that a cross-region failover completes within an acceptable window for your business, commonly under a few minutes, while still being long enough that you're not creating excessive DNS query load under normal operation. Test the actual failover time in a drill rather than assuming the configured value is what clients will experience.
Can we skip DNS failover entirely if we only operate in one region?
You can, but that also means a full regional outage takes your entire service down with no automated path to recovery beyond manually standing up capacity elsewhere. Whether that's an acceptable risk depends on how costly downtime actually is for your specific business, not a general best practice that applies to everyone equally.
Do all clients actually respect a DNS record's time-to-live correctly?
No, and that's an important caveat: some resolvers and client libraries cache DNS answers longer than the configured time-to-live, or ignore it in some circumstances. Plan for failover to take somewhat longer in practice than your configured time-to-live suggests, rather than treating that number as a hard guarantee.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
DNS Failover: Why It's Slower Than It Looks
A short TTL doesn't guarantee fast failover; some resolvers ignore it. What DNS failover actually controls, and when anycast is worth the jump.
What Happens to Your Traffic During a DNS Failover, Exactly?
A Q&A walkthrough of what actually happens during a DNS-based failover: TTL behavior, health checks, and why some clients don't fail over at all.
DNS Failover: What Actually Happens When a Region Dies
Health-check failover vs. anycast routing, why TTL is the hidden variable, and how to actually test failover instead of trusting the runbook.
Why DNS Failover Alone Won't Save You During a Regional Outage
What DNS-based failover actually does and doesn't protect against, including TTL and caching pitfalls, and what to pair it with for a real multi-region setup.
Anycast DNS Failover: What It Actually Buys You
How anycast DNS routing differs from a simple health-check failover, what it protects against, and where it can't replace application redundancy.
Setting Up DNS Failover That Actually Fails Over When It Matters
How DNS based failover actually works, where TTLs and caching quietly undermine it, and how to test a failover policy before you need it during an outage.