Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

When Multi-Region Routing Sends Traffic to the Wrong Place

Multi-region routing looks solved the day you turn it on. Users land in the closest healthy region, latency drops, and the dashboard goes green. The real test comes later, during an actual regional incident, when the system behaves nothing like it did in the demo.

Most of that gap traces back to three decisions made early and rarely revisited: how region health gets measured, how fast the router reacts to a bad signal, and what happens to requests and sessions that were mid-flight when a region went dark. Get those right and failover is uneventful. Get them wrong and a single availability zone blip turns into an outage that touches every region.

Why Your Health Checks Are Lying to the Router

Most routing layers check the thing that's easiest to check, not the thing that matters. A TCP handshake or a shallow health endpoint can return green while the database connection pool behind it is exhausted, or a downstream dependency in that region is timing out. The router keeps sending traffic to a region that's technically up and practically unusable.

Deep health checks fix this, but they carry their own risk: a check that calls three downstream services to verify itself, and then times out slowly, can make the health signal the bottleneck during an incident. Aim for a check that verifies the two or three dependencies that actually gate a successful request, with a strict timeout, and nothing beyond that. Watch your DNS caching too: a health check that flips in seconds does nothing if resolvers and clients are still holding the old answer for minutes.

Split-Brain: When Two Regions Both Think They're Primary

Active-active architectures promise better utilization and faster failover, and they deliver both, right up until a network partition lets two regions each believe they own the same write. Whichever conflict-resolution rule you picked, last-write-wins, vector clocks, a designated tiebreaker region, gets exercised for real for the first time during an outage, which is the worst possible moment to discover it's wrong.

If your data layer can't genuinely support multi-writer semantics for a given table, don't pretend it can with routing tricks. Pin writes for that data to a single region and make failover to a backup writer an explicit, monitored, and deliberately slower operation than read failover. A brief pause in write availability during a real regional failure beats silently corrupted data that nobody notices until reconciliation.

What Happens to Sessions That Were Mid-Flight During a Cutover

Sticky sessions and in-region caches are the part of multi-region design nobody wants to think about, because reasoning about routing tables is more interesting than reasoning about a shopping cart that lived only in a regional cache. When a cutover happens, those sessions either get dropped, silently reconstructed with stale data, or, worst case, partially replayed against the new region's state.

Decide, deliberately, what a user experiences when their session's home region disappears mid-request: a forced re-login, a rebuilt cart pulled from durable storage, or a short, visible retry. Whatever you pick, make it consistent and make it something support can explain in one sentence, instead of something engineering discovers three tickets into an incident.

Running a Failover Test Without Waking Up On-Call

If you've committed to 99.9 percent uptime, you're working with well under a full day of downtime for the entire year, about 8.76 hours total1, and a botched regional cutover can burn through that budget in minutes. The safest way to learn how your routing actually behaves under failure is to fail a region over on purpose, on a schedule, with someone watching.

Start by shifting a single-digit percentage of read traffic to the secondary region during low-traffic hours and confirm the health signal, the dashboards, and the alerting all agree with reality. Widen the traffic share, then include writes, then finally schedule a drill that takes a region fully out of rotation for a defined window. Each step should leave a short written record of what broke, because the value of a drill is the list of assumptions it corrects, not the successful run.

Keep Your Rollback Path Outside the Region That's Down

A surprising number of multi-region setups keep their routing configuration, feature flags, or DNS management console dependent on the very region that's most likely to fail. When that region goes down, the team can see the problem but can't route around it, because the control plane for fixing it lives inside the blast radius.

Host your routing control plane, your DNS provider access, and your incident runbooks somewhere that survives the failure you're planning for. That usually means a third-party DNS or traffic-management service, credentials that don't depend on regional single sign-on being up, and a runbook stored somewhere reachable when your primary tooling isn't.

Before you trust a failover, check each of these:

  • Health checks test the dependencies that matter, such as the database connection pool, instead of a shallow handshake that stays green.
  • Only one region can own a given write during a network partition, so two regions never both act as primary.
  • Sessions and in-region caches have a plan for what happens to requests that were mid-flight during the cutover.
  • DNS caching is accounted for, since real users and third-party resolvers keep a cached answer long after your dashboards update.
  • Routing configuration, feature flags and the DNS console do not depend on the region most likely to fail.
Executive Capability Standard

What Good Looks Like

Good multi-region routing means a region can fail, traffic reroutes automatically within your target time, and no engineer has to manually flip DNS or reassign a database writer during the incident.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Map every place a request depends on region-specific state: sessions, caches, database writers, feature flag services, so you know exactly what a cutover has to handle.
2. Do Manually:Run a scheduled failover drill where an engineer manually shifts traffic and watches dashboards, sessions, and error rates in real time.
3. Delegate:Hand ownership of the routing health checks and cutover runbook to a senior engineer who updates them every time the dependency map changes.
4. Automate:Build automated health checks with strict timeouts and let the router shift read traffic without a human in the loop during a regional failure.
5. Buy:Bring in a traffic-management or DNS provider built for regional failover instead of hand-rolling health-check logic inside your own load balancer.

How to Get Started

Frequently Asked Questions

How do I know if I actually need multi-region routing instead of a CDN and multiple availability zones?

Multi-AZ handles most single-datacenter failures and is far simpler to run correctly. Reach for true multi-region when a regulator, a major customer contract, or your own downtime budget requires surviving the loss of an entire cloud region, not just a rack or a zone. If you can't name that specific requirement, multi-AZ with a CDN in front usually gets you more reliability per engineering hour.

Should new regions be active-active or active-passive?

Active-passive is easier to reason about and to test, and it's the right default unless you have a clear latency or capacity reason to serve live traffic from both regions. Active-active adds real value for global user bases with tight latency needs, but it also adds conflict-resolution complexity that has to be designed and drilled, not assumed to work when a partition actually happens.

What's the single most common cause of a failed regional cutover?

Stale DNS caching, by a wide margin. Teams test failover by watching their own dashboards, which pick up the new health state instantly, while real users and third-party resolvers are still holding a cached answer that points at the dead region for minutes. Test with a client that respects your actual cache lifetimes, not just your monitoring.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides