Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

DNS Failover, Answered: TTLs, Health Checks and What Actually Fails Over

DNS failover gets configured once, tested with a single manual cutover, and then trusted for years without anyone revisiting whether the health check actually verifies what the team assumes it does. Here are straight answers to the questions that come up once it's time to actually rely on it during a real incident, not just during the original setup.

Why doesn't a low TTL make failover instant?

A TTL is a request to resolvers to stop caching a record after a certain time, not a guarantee. Some ISPs and corporate resolvers cache longer than the TTL specifies, and some clients or applications cache a resolved IP independently of DNS entirely, holding a connection open to the old address well past when DNS itself would have updated. A low TTL shortens the tail of stale traffic after a failover; it doesn't eliminate it, and treating it as instant is the gap that shows up mid-incident when a meaningful slice of traffic is still hitting the dead endpoint minutes after the DNS record changed.

For example, imagine a team sets a low TTL, triggers failover and sees the dashboard turn green within a minute. Meanwhile, request logs at the old address show traffic still arriving from resolvers and long-lived client connections that never re-resolved. Build that tail into your incident runbook. Tell responders that failover completes gradually, and don't declare recovery until traffic to the dead endpoint has actually dropped. Watching the old endpoint's request volume is a more reliable signal than checking that the DNS record changed.

What does a DNS health check actually verify?

Typically, a TCP connection or an HTTP request against a specific endpoint, checked for a response code or a simple string match, nothing more. That confirms the process is up and the port is listening; it doesn't confirm the actual request path behind that endpoint is healthy. A health check hitting a lightweight status page can report green while the real application, several dependencies deep, is failing every genuine request. Point the health check at something that exercises a real, representative path, not the cheapest endpoint to keep green.

When does anycast make manual failover configuration pointless?

Anycast routes a request to the nearest healthy instance at the network layer itself, so many regional outages resolve automatically without any DNS record ever changing, which is a meaningfully faster response than DNS propagation allows. It only helps if you're actually running instances in more than one location for it to route between; anycast in front of a single region doesn't add any failover capability, it just adds a routing layer with nothing to fail over to.

What's the failure mode nobody plans for?

The primary region looking healthy to the health checker while it's actually serving errors to real users, because the health-check endpoint doesn't touch the same dependency that's actually broken. A database connection pool exhausted for real user traffic but not for the lightweight health-check query, or a degraded third-party dependency that only a subset of real endpoints call, both produce this exact gap. The health check needs to fail when the thing users experience fails, which usually means it has to be less lightweight than the endpoint that's cheapest to keep monitoring, even though that makes the check itself a little more expensive to run on every interval.

How do we actually test this before we need it in a real incident?

Run a scheduled, deliberate failover drill, not just a one-time test at setup. Fail over to the secondary on a known schedule, confirm real traffic actually shifts within the time you expect, and fail back. A failover path that hasn't been exercised in months is effectively untested, because DNS provider configuration, health check endpoints, and the secondary region's actual capacity to handle full production load can all drift silently in the meantime, and the first time anyone would notice is during the incident the failover was built for.

Run a failover drill using these steps:

  1. Fail over to the secondary on a known schedule, instead of testing once at setup and trusting it for years.
  2. Confirm real traffic actually shifts within the time you expect, allowing for resolvers and clients that cache longer than the TTL.
  3. Push enough real traffic through the secondary to see how it behaves at production load, since standby sizing is a common failure.
  4. Fail back, and note any drift in DNS provider configuration or health check endpoints before the next drill.

What breaks in the secondary that a drill usually catches

The most common finding from a real drill isn't in the DNS layer at all: it's a secondary region sized for standby, not full load, that starts failing the moment real traffic actually lands on it. A drill that only confirms DNS resolved correctly, without pushing enough real traffic through the secondary to see how it behaves under production load, misses exactly the failure a real incident would expose. Size the secondary for what it would actually need to carry, and confirm that with traffic, not with a capacity spreadsheet.

Executive Capability Standard

What Good Looks Like

A working DNS failover setup has a health check that exercises real dependencies rather than a lightweight status page, a TTL tuned to actually shorten the stale-traffic tail, and a failover path tested on a real schedule, not just once at setup.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Review what your current health check endpoint actually verifies, and whether it touches the same dependencies a real user request does.
2. Do Manually:Manually run a failover drill: shift real traffic to the secondary, measure the actual cutover time, and fail back.
3. Delegate:Have an engineer rework the health check to exercise a representative real path, if the current one is a lightweight status endpoint.
4. Automate:Schedule recurring, automated failover drills so drift in configuration or secondary-region capacity gets caught before a real incident.
5. Buy:Bring in infrastructure or networking advisory if you're evaluating anycast for the first time and want the routing and failover design reviewed before committing to it.

How to Get Started

Frequently Asked Questions

How low should our DNS TTL be for failover purposes?

Low enough to meaningfully shorten the stale-traffic tail, often in the range of thirty to sixty seconds for records you might need to fail over, but expect some resolvers and clients to hold on longer regardless. Pair a low TTL with anycast where you can, since anycast's network-layer routing isn't subject to the same caching behavior.

Should our health check hit the same endpoint real users hit?

Not necessarily the same endpoint, but it must exercise the same critical dependencies a real request does. A lightweight status page can stay green while the real path fails, so the check should touch the primary database and any required third-party call.

How often should we actually test failover?

On a regular schedule, not just once at setup, since DNS configuration, health check endpoints, and secondary-region capacity can all drift over time. A quarterly drill that actually shifts real traffic and measures the real cutover time catches drift before an actual incident does.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides