Failover and High Availability: The Questions Worth Asking First
Failover is one of those topics where teams reach for the most sophisticated architecture before answering the more basic question: how much downtime can this business actually tolerate, and what is it worth paying to avoid it.
These are the questions worth working through in order, before committing to a specific failover design, since the answer to the first one usually changes what the right answer is to the rest.
How many nines does a small company actually need?
Fewer than most teams assume. Going from 99 percent to 99.9 percent availability cuts allowed downtime from 3.65 days a year to about 8.76 hours, and going from 99.9 to 99.99 cuts it again to roughly 52.6 minutes1. Each step up costs meaningfully more in engineering complexity, and for most products that first jump matters far more to customers than any of the jumps after it.
Pick the target based on what your contracts and your customers' patience actually require, not what sounds impressive in a pitch deck. A target you can actually hit and explain to a customer beats a more ambitious one that quietly slips every quarter.
Active-active or active-passive: which fits your team?
Active-active, where multiple regions or instances all handle live traffic, gives you the fastest failover because there's nothing to promote, but it demands that your data layer handle writes from multiple locations without conflicts, which is genuinely hard and easy to get subtly wrong.
Active-passive, where a standby only takes over on failure, is simpler to reason about and cheaper to run, at the cost of a failover delay while the standby spins up and takes over. For most small and mid-sized teams, active-passive with a well-tested failover path beats an active-active setup nobody fully understands, since an architecture your team can actually operate under pressure is worth more than one that's theoretically faster but poorly understood.
What actually causes failover to fail when you need it?
Almost always, it's the parts nobody tested. DNS records with a TTL too long to fail over quickly. A standby database that's been silently falling behind on replication for weeks. Application code with a hardcoded reference to the primary region's endpoint that nobody remembered was there.
The failover mechanism itself usually works in a demo. What breaks it in production is a dependency nobody accounted for, discovered for the first time during the actual outage.
How do you test failover without causing the outage you're avoiding?
Start small and scheduled. Fail over a low-traffic service during a planned window, with the team watching, before ever trying it on anything customer-critical. Treat the drill like a real incident: time how long it actually takes, and write down every step that required a human to remember something instead of following a runbook.
A failover plan that's never been exercised is a hypothesis, not a plan. Run the drill quarterly at minimum, and after any significant architecture change to the services it covers.
A first failover drill can follow these steps:
- Pick a low-traffic service for the first drill, not anything customer-critical.
- Schedule the failover for a planned window and have the team watching while it runs.
- Run it like a real incident, with the same roles and communication you would use in an actual outage.
- Note what did not work, such as slow DNS changes or a standby that fell behind on replication, and fix it before the next drill.
- Move on to more critical services only after the small drills go cleanly.
When is redundancy not worth the cost?
For internal tools, batch jobs that can retry later, or anything where a short outage costs inconvenience rather than revenue, full multi-region redundancy is usually overbuilt. Save the investment for the paths that are actually customer-facing and time-sensitive, and accept a simpler recovery process, restore from backup, restart from the last checkpoint, for everything else.
Matching the redundancy investment to what each service actually needs, rather than applying one standard everywhere, is usually the difference between a high-availability strategy that's affordable and one that isn't.
What role does replication lag play in a failover?
A standby that looks ready can still lose data on failover if it was behind the primary when the primary went down. Any write that hadn't replicated yet is gone, or worse, conflicts with what the standby thinks happened next once it takes over.
Monitor replication lag as its own signal, separate from whether the standby is technically online. A standby that's up but hours behind gives you a false sense of protection, and finding that out during a real failover is a much worse time to learn it than during a routine check.
What Good Looks Like
Good failover practice means the availability target matches what the business actually needs, the failover path has been tested under realistic conditions, and someone can execute it correctly under pressure without waiting on the one person who built it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Is multi-region always better than multi-availability-zone?
Not for most failure modes. Multi-AZ protects against a data center failure within one region and is far simpler to operate. Multi-region protects against a regional outage or a compliance requirement to keep data in a specific geography, but it adds real complexity that isn't worth it unless you actually need that protection.
How long should a failover drill take to run?
Budget half a day the first few times, including setup and a full debrief on what didn't go cleanly. As the process matures and the runbook gets more reliable, a routine drill should take less than an hour, which is itself a useful sign of how ready the system actually is.
Do we need a fully automated failover, or is a manual trigger okay?
A manual trigger is fine, and often safer, as long as someone is on call to pull it and the runbook is specific enough that any engineer on the rotation can follow it under pressure. Automate the failover mechanism itself once the manual process is proven reliable, not before.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
What High Availability Actually Costs Beyond the Second Region
A worked-example breakdown of what running a second region for failover really costs, and how to decide whether your uptime target justifies it.
Active-Active vs Active-Passive: What Your Uptime Target Buys You
A comparison of active-active, active-passive and single-region failover, with the real infrastructure and headcount cost each uptime target requires.
What High Availability Really Costs, and What It Buys You
A plain-language look at the real cost of failover and redundancy, matched against what different availability targets actually mean in downtime terms.
What an Hour of Downtime Actually Costs You
How to work out your real cost of downtime, match it to an availability target, and decide whether a second region is actually worth paying for yet.
How Much Redundancy Your Vector Store Actually Needs
Replicated indexes, snapshot restore, and multi-region setups each buy different recovery guarantees. Match the approach to your actual uptime target.
Working Out Your Pipeline's Actual Downtime Budget
A worked example of turning an availability target into a real downtime budget for a streaming pipeline, and what that means for failover design.