AI Model Serving & Inference OptimizationPlaybook3 min readUpdated September 2026

Why DNS Failover Isn't as Fast as You Think It Is

DNS failover sounds like it should be instant: a health check fails, the record changes, traffic moves. In practice it's one of the slower failover mechanisms available, and treating it like a fast one is how a five-minute outage turns into a much longer one.

Understanding where the delay actually comes from, and when anycast routing sidesteps it, changes how much you can rely on DNS alone during an incident.

What DNS failover can and can't do

A health-checked DNS record changes which IP address a resolver hands back the next time it looks your domain up. It does nothing for a connection that's already open, and it does nothing for a client or resolver that's still holding an older, cached answer. DNS failover moves future lookups, not current traffic.

That's a meaningful limitation for anything long-lived, a persistent connection, a mobile client that cached a resolution hours ago, a corporate resolver that ignores your stated cache time. For short-lived HTTP requests from clients that respect DNS caching properly, it works closer to as advertised.

The TTL problem nobody tests until an outage

The TTL you set on a record is a request, not a guarantee. Some resolvers and ISPs cache longer than you asked for, and some clients cache the result inside the application layer on top of whatever the OS resolver does. Setting a low TTL helps, but it doesn't make every resolver in the path honor it.

Anycast sidesteps the whole problem, because failover happens at the network routing layer instead of through DNS propagation. The same IP address is announced from multiple locations, and the network routes a client to whichever healthy location is closest, so there's no cached answer to wait out.

Health checks that catch a real failure

A health check that hits a shallow endpoint returning a fixed 200 response tells you the process is running, not that the service works. If that endpoint doesn't touch the database, the queue, or whatever dependency actually breaks user requests, your failover will sit there passing checks while real traffic fails.

Build the check to exercise something close to what a real request needs. It should fail when the thing that actually matters is down, and it should not flap on a single slow response, add a short window of consecutive failures before triggering failover, so one bad health check doesn't send traffic somewhere and back for no reason.

Anycast versus latency-based routing

Anycast announces one IP address from many network locations and lets routing decide which one answers, giving you failover measured in seconds because it happens below DNS entirely. It's what CDNs and some DNS providers run their own edge networks on.

Latency-based or geo DNS routing instead hands out different IP addresses depending on where the resolver is, still bound by whatever TTL the resolver actually respects. It's simpler to reason about and doesn't require owning network infrastructure, but a real failover through it moves at DNS speed, not network speed. Pick anycast when failover time matters most, and geo DNS when you mainly want to route by region and can tolerate a slower switch.

Testing the failover before a real outage does it for you

A failover setup that's never been exercised is a hypothesis, not a plan. Run a deliberate drill: fail the health check on purpose, or manually flip the record, and time how long it actually takes for traffic to move, not how long the runbook says it should take.

Watch what happens from a few different vantage points and resolvers, not just your own laptop, since that's the only way to catch a resolver in the wild that's ignoring your TTL. A drill that surfaces a slow or partial failover is worth far more before an incident than during one, and it's the only way to know whether the secondary region you're failing over to can actually absorb full production traffic instead of falling over itself the moment it starts receiving it.

Write down what the drill actually showed, the real time to full cutover, which vantage points lagged, and treat that number as your honest failover time until the next drill changes it. A runbook with an untested time-to-failover figure in it is a guess dressed up as a fact, and it's the kind of guess that gets relied on at exactly the wrong moment.

A useful failover drill follows these steps:

  1. Fail the health check on purpose, or flip the record by hand, so the failover happens when you choose instead of during a real incident.
  2. Time how long traffic actually takes to move, and record that measured number rather than trusting the figure written in the runbook.
  3. Check the result from several vantage points and resolvers, not just your own laptop, since caching behavior differs across ISPs and clients.
  4. Note which long-lived connections and application-layer caches ignored the change, because DNS failover only moves future lookups, not current traffic.
  5. Compare the measured time with what your outage tolerance allows, and consider anycast if the gap is too wide.
Executive Capability Standard

What Good Looks Like

Good failover means a real outage in one region routes traffic to a healthy one inside the time your users will tolerate, and you know that because you tested it, not because the runbook says it should work.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map every place your infrastructure depends on DNS to route traffic, and check the actual TTL set on each of those records today.
2. Do Manually:Run a manual failover drill: point traffic at your secondary region by hand and time how long it actually takes clients to pick up the change.
3. Delegate:Have someone own the health check logic specifically, checking the dependency that actually breaks requests, not just an endpoint that always returns success.
4. Automate:Configure automated, health-checked DNS failover so the switch happens without someone paging in first.
5. Buy:Move latency-sensitive traffic onto an anycast network from a CDN or DNS provider that operates its own edge, instead of building failover on ordinary authoritative DNS.

How to Get Started

Frequently Asked Questions

Is anycast the same thing as a CDN?

No, though CDNs typically run on anycast networks. Anycast is a routing technique, announcing one IP from many locations. A CDN is a broader service built on top of that, adding caching, edge compute, and content delivery.

How low should my DNS TTL be?

Low enough to fail over in a time you can live with, but every lookup has a cost, so going extremely low mainly adds resolver query load without a proportional benefit once you're below what most resolvers will actually honor.

Does DNS failover replace a load balancer?

No. A load balancer distributes and reroutes traffic within a region at the network layer, often in milliseconds. DNS failover operates a layer above that, deciding which region or provider traffic goes to in the first place.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides