Model Context Protocol & Agentic ArchitecturePlaybook3 min readUpdated September 2026

Active-Active or Active-Passive: Choosing a DNS Failover Setup

DNS failover routes traffic away from a failed region or data center by changing what a DNS query resolves to. It sounds like a simple switch, but two decisions determine whether it actually works when you need it: whether you run active-active or active-passive, and whether your health check actually detects the failures that matter instead of just confirming a server is technically reachable.

This walks through both decisions, plus the DNS time-to-live setting most teams configure once and never revisit, which quietly determines how long failover actually takes regardless of how fast your health check reacts.

What DNS Failover Can and Can't Fix

DNS failover solves the case where an entire region, data center, or point of presence becomes unreachable: a cloud provider outage, a network partition, a catastrophic failure at one location. It does not solve a partial degradation where the service is technically reachable but slow or returning errors for some requests, since a basic health check often still reports that endpoint as healthy.

It also doesn't fail over instantly. DNS resolvers and clients cache responses for the duration of your TTL, and some resolvers or clients cache longer than they should regardless of what you set. Plan around a failover window measured in minutes at best, not the instant cutover a load balancer in front of a single region can offer within that region.

Active-Passive: Simpler, With a Real Failover Delay

Active-passive keeps one region serving all traffic while a second stands ready but idle. It's simpler to reason about: only one region is actively serving requests at any time, so there's no need to worry about data consistency or session handling across two live regions simultaneously. The cost is failover time: detecting the failure, updating DNS, and waiting for caches to expire all happen after the primary region is already down, during which real users are affected.

This fits services where a failover delay of several minutes is acceptable and where running a second region hot, actively serving traffic, adds cost or complexity, data replication, session affinity, that isn't worth it for your actual availability requirements.

Active-Active: Faster Failover, More Complexity

Active-active serves traffic from multiple regions simultaneously, so when one fails, DNS failover just stops sending new traffic there while existing capacity in the other region absorbs the load, rather than needing to spin up or warm a passive region from a cold start. This is meaningfully faster in practice, since the failover mechanism doesn't have to wait for a standby region to actually be ready to serve traffic.

The cost is real operational complexity: your data layer needs a strategy for consistency across regions, whether that's a globally distributed database, regional read replicas with careful write routing, or eventual consistency your application can tolerate. Don't take on active-active for the failover speed benefit alone if your data architecture isn't already built for multi-region operation, since retrofitting that afterward is a much larger project than the DNS configuration itself.

Designing a Health Check That Fails for the Right Reasons

A health check that only confirms the server responds to a ping or returns a healthy status on a static endpoint will report healthy right up until a downstream dependency, a database, a critical internal API, fails and takes the actual user-facing functionality down with it. Build the health check to exercise a real, representative path: a lightweight version of the actual critical transaction, not just process liveness.

Balance this against false positives: a health check that's too strict, failing on a slow but not actually broken downstream dependency, triggers unnecessary failovers that add their own risk and confusion. Tune the check to fail on things that genuinely indicate the region can't serve real traffic, and let genuinely transient slowness recover on its own without triggering a full regional failover.

The TTL Tradeoff Nobody Configures Correctly

Your DNS record's TTL directly trades off two things, and most teams pick one value and never revisit it:

  • A short TTL means failover propagates faster once triggered, since clients and resolvers stop trusting the old answer sooner.
  • A short TTL also means more DNS query volume and, for a paid DNS provider, higher cost, along with slightly higher latency on every fresh resolution.
  • A long TTL reduces query volume and cost but extends how long some fraction of clients keep hitting the failed endpoint after failover has already been triggered.
  • Some resolvers and clients ignore or override your TTL entirely, which is a real limitation of DNS failover, not a configuration mistake you can fully engineer around.

Set the TTL short enough to matter for your actual failover time requirements, and treat DNS failover as one layer of a broader resilience strategy, not the only mechanism standing between an outage and your users.

Executive Capability Standard

What Good Looks Like

A good DNS failover setup means the health check fails specifically when real user transactions would fail, and the TTL is set deliberately based on your actual failover time requirement, not left at a default.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Study your actual failure history: how often has a full region gone down versus a partial degradation that a basic health check wouldn't have caught.
2. Do Manually:Manually trigger a failover in a staging environment and time how long it actually takes for traffic to fully shift, including client-side caching delays.
3. Delegate:Have whoever owns your infrastructure own the health check design specifically, since a health check that's too strict or too lenient undermines the whole failover setup regardless of which topology you chose.
4. Automate:Automate health checks against a real representative transaction path, not just server liveness, so failover triggers for the failures that actually matter to users.
5. Buy:Use a managed DNS provider with built-in health-check-based failover rather than building your own health check and DNS update pipeline, unless your failover requirements are unusual enough to need custom logic.

How to Get Started

Frequently Asked Questions

Is active-active always better since failover is faster?

Not if your data layer isn't ready for multi-region operation. Active-active requires a real strategy for data consistency across regions, and retrofitting that after choosing active-active for the failover speed alone is a much bigger project than the DNS setup itself.

How fast does DNS failover actually happen?

Plan for minutes, not seconds, even with a short TTL. Some resolvers and clients cache longer than your configured TTL regardless of what you set, so DNS failover should be one layer of resilience, not the only mechanism you rely on for fast recovery.

What's wrong with a health check that just pings the server?

It reports healthy even when a downstream dependency the actual user-facing feature depends on has failed. Build the check around a lightweight version of a real critical transaction instead, so it fails for the reasons that actually mean the region can't serve users.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides