Active-Active or Active-Passive: Choosing a DNS Failover Setup
DNS failover routes traffic away from a failed region or data center by changing what a DNS query resolves to. It sounds like a simple switch, but two decisions determine whether it actually works when you need it: whether you run active-active or active-passive, and whether your health check actually detects the failures that matter instead of just confirming a server is technically reachable.
This walks through both decisions, plus the DNS time-to-live setting most teams configure once and never revisit, which quietly determines how long failover actually takes regardless of how fast your health check reacts.
What DNS Failover Can and Can't Fix
DNS failover solves the case where an entire region, data center, or point of presence becomes unreachable: a cloud provider outage, a network partition, a catastrophic failure at one location. It does not solve a partial degradation where the service is technically reachable but slow or returning errors for some requests, since a basic health check often still reports that endpoint as healthy.
It also doesn't fail over instantly. DNS resolvers and clients cache responses for the duration of your TTL, and some resolvers or clients cache longer than they should regardless of what you set. Plan around a failover window measured in minutes at best, not the instant cutover a load balancer in front of a single region can offer within that region.
Active-Passive: Simpler, With a Real Failover Delay
Active-passive keeps one region serving all traffic while a second stands ready but idle. It's simpler to reason about: only one region is actively serving requests at any time, so there's no need to worry about data consistency or session handling across two live regions simultaneously. The cost is failover time: detecting the failure, updating DNS, and waiting for caches to expire all happen after the primary region is already down, during which real users are affected.
This fits services where a failover delay of several minutes is acceptable and where running a second region hot, actively serving traffic, adds cost or complexity, data replication, session affinity, that isn't worth it for your actual availability requirements.
Active-Active: Faster Failover, More Complexity
Active-active serves traffic from multiple regions simultaneously, so when one fails, DNS failover just stops sending new traffic there while existing capacity in the other region absorbs the load, rather than needing to spin up or warm a passive region from a cold start. This is meaningfully faster in practice, since the failover mechanism doesn't have to wait for a standby region to actually be ready to serve traffic.
The cost is real operational complexity: your data layer needs a strategy for consistency across regions, whether that's a globally distributed database, regional read replicas with careful write routing, or eventual consistency your application can tolerate. Don't take on active-active for the failover speed benefit alone if your data architecture isn't already built for multi-region operation, since retrofitting that afterward is a much larger project than the DNS configuration itself.
Designing a Health Check That Fails for the Right Reasons
A health check that only confirms the server responds to a ping or returns a healthy status on a static endpoint will report healthy right up until a downstream dependency, a database, a critical internal API, fails and takes the actual user-facing functionality down with it. Build the health check to exercise a real, representative path: a lightweight version of the actual critical transaction, not just process liveness.
Balance this against false positives: a health check that's too strict, failing on a slow but not actually broken downstream dependency, triggers unnecessary failovers that add their own risk and confusion. Tune the check to fail on things that genuinely indicate the region can't serve real traffic, and let genuinely transient slowness recover on its own without triggering a full regional failover.
The TTL Tradeoff Nobody Configures Correctly
Your DNS record's TTL directly trades off two things, and most teams pick one value and never revisit it:
- A short TTL means failover propagates faster once triggered, since clients and resolvers stop trusting the old answer sooner.
- A short TTL also means more DNS query volume and, for a paid DNS provider, higher cost, along with slightly higher latency on every fresh resolution.
- A long TTL reduces query volume and cost but extends how long some fraction of clients keep hitting the failed endpoint after failover has already been triggered.
- Some resolvers and clients ignore or override your TTL entirely, which is a real limitation of DNS failover, not a configuration mistake you can fully engineer around.
Set the TTL short enough to matter for your actual failover time requirements, and treat DNS failover as one layer of a broader resilience strategy, not the only mechanism standing between an outage and your users.
What Good Looks Like
A good DNS failover setup means the health check fails specifically when real user transactions would fail, and the TTL is set deliberately based on your actual failover time requirement, not left at a default.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is active-active always better since failover is faster?
Not if your data layer isn't ready for multi-region operation. Active-active requires a real strategy for data consistency across regions, and retrofitting that after choosing active-active for the failover speed alone is a much bigger project than the DNS setup itself.
How fast does DNS failover actually happen?
Plan for minutes, not seconds, even with a short TTL. Some resolvers and clients cache longer than your configured TTL regardless of what you set, so DNS failover should be one layer of resilience, not the only mechanism you rely on for fast recovery.
What's wrong with a health check that just pings the server?
It reports healthy even when a downstream dependency the actual user-facing feature depends on has failed. Build the check around a lightweight version of a real critical transaction instead, so it fails for the reasons that actually mean the region can't serve users.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
DNS Failover: Why It's Slower Than It Looks
A short TTL doesn't guarantee fast failover; some resolvers ignore it. What DNS failover actually controls, and when anycast is worth the jump.
What Happens to Your Traffic During a DNS Failover, Exactly?
A Q&A walkthrough of what actually happens during a DNS-based failover: TTL behavior, health checks, and why some clients don't fail over at all.
Why DNS Failover Isn't as Fast as You Think It Is
How DNS-based failover and anycast routing actually work, why TTLs slow failover down, and how to build a setup you've tested before you need it.
Anycast DNS Failover: What It Actually Buys You
How anycast DNS routing differs from a simple health-check failover, what it protects against, and where it can't replace application redundancy.
DNS Failover: What Actually Happens When a Region Dies
Health-check failover vs. anycast routing, why TTL is the hidden variable, and how to actually test failover instead of trusting the runbook.
Why DNS Failover Alone Won't Save You During a Regional Outage
What DNS-based failover actually does and doesn't protect against, including TTL and caching pitfalls, and what to pair it with for a real multi-region setup.