Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Failover and High Availability: The Questions Worth Asking First

Failover is one of those topics where teams reach for the most sophisticated architecture before answering the more basic question: how much downtime can this business actually tolerate, and what is it worth paying to avoid it.

These are the questions worth working through in order, before committing to a specific failover design, since the answer to the first one usually changes what the right answer is to the rest.

How many nines does a small company actually need?

Fewer than most teams assume. Going from 99 percent to 99.9 percent availability cuts allowed downtime from 3.65 days a year to about 8.76 hours, and going from 99.9 to 99.99 cuts it again to roughly 52.6 minutes1. Each step up costs meaningfully more in engineering complexity, and for most products that first jump matters far more to customers than any of the jumps after it.

Pick the target based on what your contracts and your customers' patience actually require, not what sounds impressive in a pitch deck. A target you can actually hit and explain to a customer beats a more ambitious one that quietly slips every quarter.

Active-active or active-passive: which fits your team?

Active-active, where multiple regions or instances all handle live traffic, gives you the fastest failover because there's nothing to promote, but it demands that your data layer handle writes from multiple locations without conflicts, which is genuinely hard and easy to get subtly wrong.

Active-passive, where a standby only takes over on failure, is simpler to reason about and cheaper to run, at the cost of a failover delay while the standby spins up and takes over. For most small and mid-sized teams, active-passive with a well-tested failover path beats an active-active setup nobody fully understands, since an architecture your team can actually operate under pressure is worth more than one that's theoretically faster but poorly understood.

What actually causes failover to fail when you need it?

Almost always, it's the parts nobody tested. DNS records with a TTL too long to fail over quickly. A standby database that's been silently falling behind on replication for weeks. Application code with a hardcoded reference to the primary region's endpoint that nobody remembered was there.

The failover mechanism itself usually works in a demo. What breaks it in production is a dependency nobody accounted for, discovered for the first time during the actual outage.

How do you test failover without causing the outage you're avoiding?

Start small and scheduled. Fail over a low-traffic service during a planned window, with the team watching, before ever trying it on anything customer-critical. Treat the drill like a real incident: time how long it actually takes, and write down every step that required a human to remember something instead of following a runbook.

A failover plan that's never been exercised is a hypothesis, not a plan. Run the drill quarterly at minimum, and after any significant architecture change to the services it covers.

A first failover drill can follow these steps:

  1. Pick a low-traffic service for the first drill, not anything customer-critical.
  2. Schedule the failover for a planned window and have the team watching while it runs.
  3. Run it like a real incident, with the same roles and communication you would use in an actual outage.
  4. Note what did not work, such as slow DNS changes or a standby that fell behind on replication, and fix it before the next drill.
  5. Move on to more critical services only after the small drills go cleanly.

When is redundancy not worth the cost?

For internal tools, batch jobs that can retry later, or anything where a short outage costs inconvenience rather than revenue, full multi-region redundancy is usually overbuilt. Save the investment for the paths that are actually customer-facing and time-sensitive, and accept a simpler recovery process, restore from backup, restart from the last checkpoint, for everything else.

Matching the redundancy investment to what each service actually needs, rather than applying one standard everywhere, is usually the difference between a high-availability strategy that's affordable and one that isn't.

What role does replication lag play in a failover?

A standby that looks ready can still lose data on failover if it was behind the primary when the primary went down. Any write that hadn't replicated yet is gone, or worse, conflicts with what the standby thinks happened next once it takes over.

Monitor replication lag as its own signal, separate from whether the standby is technically online. A standby that's up but hours behind gives you a false sense of protection, and finding that out during a real failover is a much worse time to learn it than during a routine check.

Executive Capability Standard

What Good Looks Like

Good failover practice means the availability target matches what the business actually needs, the failover path has been tested under realistic conditions, and someone can execute it correctly under pressure without waiting on the one person who built it.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Write down your current recovery time and recovery point for your most critical service, based on what would actually happen today, not what the architecture diagram implies.
2. Do Manually:Run a scheduled failover drill on a low-traffic service and time every step by hand.
3. Delegate:Give an engineer ownership of the failover runbook, with a standing requirement to re-test it after any change to the services it covers.
4. Automate:Build automated health checks and failover triggers for the paths where a manual response would be too slow to matter.
5. Buy:Bring in a fractional CTO or infrastructure specialist to design multi-region architecture once a single failover path stops being enough.

How to Get Started

Frequently Asked Questions

Is multi-region always better than multi-availability-zone?

Not for most failure modes. Multi-AZ protects against a data center failure within one region and is far simpler to operate. Multi-region protects against a regional outage or a compliance requirement to keep data in a specific geography, but it adds real complexity that isn't worth it unless you actually need that protection.

How long should a failover drill take to run?

Budget half a day the first few times, including setup and a full debrief on what didn't go cleanly. As the process matures and the runbook gets more reliable, a routine drill should take less than an hour, which is itself a useful sign of how ready the system actually is.

Do we need a fully automated failover, or is a manual trigger okay?

A manual trigger is fine, and often safer, as long as someone is on call to pull it and the runbook is specific enough that any engineer on the rotation can follow it under pressure. Automate the failover mechanism itself once the manual process is proven reliable, not before.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides