Anycast DNS Failover: What It Actually Buys You
DNS failover and anycast routing get talked about as the same thing, and they're not. A simple failover swaps which address a DNS record points to when a health check fails. Anycast routes a single address to whichever of several physical locations is closest, or healthiest, at the network layer, before DNS even enters the picture for that request.
Knowing which one you actually have, and which one you actually need, matters more than either term alone.
Simple DNS failover: fast to set up, slow to actually fail over
A health check monitors your primary endpoint, and when it fails, updates a DNS record to point at a backup. The catch is DNS caching: every resolver and client that's already cached the old record keeps using it until that cache's time-to-live expires, regardless of how fast your health check reacted.
A short time-to-live helps but has its own cost, more DNS query volume hitting your nameservers, and some resolvers ignore short values and cache longer than you asked for anyway. Simple DNS failover is a reasonable fit for a lower-traffic service where a few minutes of continued requests to a dead endpoint is an acceptable cost.
Anycast: routing at the network layer, not the DNS layer
With anycast, the same address is announced from multiple physical locations, and network routing sends each client's traffic to whichever announcement is closest or healthiest at that moment, without any DNS change or cache expiry involved. That's what makes it fast to react to a regional outage: routing convergence happens at the network layer in seconds, not whenever every client's DNS cache happens to expire.
The cost is operational complexity: you need infrastructure actually running and healthy at multiple points of presence, and that routing behavior isn't something most application teams have deep experience debugging when it doesn't do what they expected.
Protecting the uptime budget that actually matters to your SLA
Anycast failover exists to protect that same budget: at four nines you're working with about 52.6 minutes of downtime a year to spend, not 52.6 minutes per incident1.
A regional outage that takes minutes to fail over at the network layer, instead of however long a cached DNS record takes to expire across every client, is the difference between staying inside that budget for the year and blowing through it in a single incident.
Where neither one saves you
Both DNS failover and anycast route traffic away from a dead endpoint; neither one fixes an application that's up but returning wrong data, a database that's healthy but out of sync with its replica, or a dependency failure inside the app that a network-level health check never sees.
Routing-layer redundancy is necessary and it's not sufficient: pair it with application-level health checks that actually exercise a real code path, not just confirm a port is open.
A worked example: a regional outage during a peak traffic window
Say a single cloud region hosting part of your infrastructure has a network-level outage during your highest-traffic hour of the week. With anycast, the routing layer detects the unhealthy announcement and shifts traffic to a healthy region within seconds, and most users never notice anything beyond a handful of retried requests. With simple DNS failover alone, every client that cached the primary record within the last several minutes keeps sending requests into the outage until that cache expires, turning a brief regional blip into an extended, visible incident for a meaningful share of your traffic.
The gap between those two outcomes is exactly the operational complexity anycast costs you, and whether that gap is worth paying for depends on how much traffic realistically hits you during your worst-case outage window.
Testing failover before you need it
A failover mechanism that's never actually been triggered outside of an outage is untested infrastructure, no different from a disaster recovery plan nobody's rehearsed. Schedule a deliberate failover test, taking a healthy endpoint offline on purpose during a low-traffic window, and confirm both that traffic actually reroutes and that your monitoring correctly identifies what happened rather than paging on a mystery. Teams that skip this step tend to discover their failover doesn't work correctly during the first real outage, which is the worst possible time to find out.
A useful failover test follows these steps:
- Schedule the test for a low-traffic window instead of waiting for a real outage to find out whether failover works.
- Take a healthy endpoint offline on purpose, rather than reviewing the failover design on paper.
- Confirm traffic actually reroutes to the backup, and note how long clients keep hitting the dead endpoint while cached DNS records expire.
- Check that monitoring correctly identifies what failed, so the on-call engineer sees the cause and not just a symptom.
- Verify the application behind the backup returns correct data, since routing-layer failover does not catch an app that is up but wrong.
What Good Looks Like
A good failover setup matches the mechanism, simple DNS or anycast, to how fast a regional outage actually needs to be routed around given the SLA's real error budget, and pairs either one with application-level health checks.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Do we need anycast, or is simple DNS failover enough?
For a lower-traffic service where a few minutes of continued requests to a dead endpoint during DNS cache expiry is tolerable, simple failover is enough and far simpler to operate. Anycast is worth the added operational complexity once regional outages need to fail over in seconds, not minutes, and your SLA's error budget doesn't have room for the delay.
Why didn't our DNS failover work immediately when the primary went down?
Almost certainly DNS caching. Every resolver and client that already cached the old record keeps using it until that cache's time-to-live expires, regardless of how fast your health check detected the failure and updated the record. A shorter value helps but doesn't eliminate the delay, since some resolvers cache longer than requested anyway.
Does anycast or DNS failover protect against an application-level bug?
No. Both route traffic away from a dead network endpoint, but neither one checks whether the application behind a healthy-looking endpoint is actually returning correct data. Pair either one with application-level health checks that exercise a real code path, not just confirm a port is open.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
DNS Failover: Why It's Slower Than It Looks
A short TTL doesn't guarantee fast failover; some resolvers ignore it. What DNS failover actually controls, and when anycast is worth the jump.
What Happens to Your Traffic During a DNS Failover, Exactly?
A Q&A walkthrough of what actually happens during a DNS-based failover: TTL behavior, health checks, and why some clients don't fail over at all.
DNS Failover: What Actually Happens When a Region Dies
Health-check failover vs. anycast routing, why TTL is the hidden variable, and how to actually test failover instead of trusting the runbook.
Why DNS Failover Alone Won't Save You During a Regional Outage
What DNS-based failover actually does and doesn't protect against, including TTL and caching pitfalls, and what to pair it with for a real multi-region setup.
DNS Failover Versus a Load Balancer: What Each One Actually Fixes
A comparison of DNS-based failover and load balancer failover for surviving a regional or full-service outage, and where each approach falls short on its own.
Setting Up DNS Failover That Actually Fails Over When It Matters
How DNS based failover actually works, where TTLs and caching quietly undermine it, and how to test a failover policy before you need it during an outage.