Active-Active vs Active-Passive: What Your Uptime Target Buys You
Every jump in availability target costs roughly an order of magnitude more in infrastructure and operational discipline than the one before it, which is why picking a target without pricing out the approach behind it is how teams end up either over-building for risk they don't have or under-building for an outage they can't afford.
The three common approaches, single-region, active-passive and active-active, aren't a menu you pick from once. Most companies move through them in order as their contracts and traffic grow, and knowing what each one actually buys you helps you time that move instead of guessing at it.
What are you accepting with single-region hosting?
A single application instance in a single region is the cheapest setup and the one most small teams start with. Its failure mode is total: if the instance, the availability zone, or the region has a problem, you're down until someone (or something) restarts service elsewhere. This is a reasonable choice when your users would forgive an hour of downtime a handful of times a year, and it's a poor choice once a customer contract includes an uptime commitment you haven't actually engineered for.
Even within "single-region," running multiple instances across separate availability zones in that region is a cheap, meaningful step up: it protects against a single machine or zone failing without the cost or complexity of a second region. Plenty of teams conflate "single-region" with "single point of failure" when the actual gap is just not having spread their existing instances across zones.
Active-passive: redundancy that waits its turn
Active-passive keeps a standby copy of your infrastructure, often in a second availability zone or region, that takes over when the primary fails. The standby can be running and idle (faster failover, more cost) or provisioned on demand when a failure is detected (slower failover, less cost). Failover isn't instant: detecting the failure, promoting the standby, and redirecting traffic typically takes anywhere from under a minute to several minutes depending on how much of that process is automated versus manual.
The operational discipline this requires is easy to underestimate: a standby that's never actually been failed over to in a drill is a standby you don't really know works. Teams that skip failover drills often discover gaps (an expired credential, a config that only exists on the primary) during a real incident instead of a scheduled test.
Database replication is usually the hardest part to get right here. An asynchronous standby database can lag behind the primary by seconds or more, which means a failover under load can lose the most recent writes. Decide up front, and document, how much data loss is acceptable in a failover; that number should drive whether you need synchronous replication, which costs latency on every write, or asynchronous, which is cheaper but riskier during a failover.
Active-active: no failover step, at the cost of running everything twice
Active-active runs full production capacity in two or more locations simultaneously, splitting real traffic across them, so there's no failover step at all: traffic just shifts away from a failed location. This buys the fastest recovery of the three approaches, but it requires your data layer to handle writes safely from multiple locations, which is a genuinely hard problem for most databases and often means real architectural changes, not just infrastructure changes.
Most teams that need this level of resilience are already operating at a scale, and with a contractual uptime commitment, that justifies the engineering investment. Adopting it earlier than that usually means solving a distributed-data problem you didn't need to solve yet.
For example, a team that adopts active-active before it's needed often discovers the hard requirement only once they hit a real conflict: two users in different regions edit the same record within seconds of each other, and now the application has to decide which write wins, or how to merge them. That's a product decision as much as an infrastructure one, and it's far cheaper to make deliberately, before it's forced by an incident, than to retrofit after.
How do you match a failover approach to your uptime target?
A 99.9% target allows roughly 8.76 hours of downtime a year, while 99.99% shrinks that to about 53 minutes and 99.999% to roughly 5 minutes1. Active-passive with a well-drilled runbook can realistically hit 99.9% to 99.95% uptime1. Getting past that into 99.99% availability territory is where active-active, or at minimum a heavily automated active-passive setup with sub-minute detection, becomes necessary rather than aspirational1.
Before committing to a target, check what your actual customer contracts require. Founders regularly find they've been designing for 99.99% uptime because it sounded appropriately serious, when their signed agreements only commit to 99.9%, a target that costs meaningfully less to operate1.
Match the approach to your situation like this:
- Single-region fits when users would forgive an hour of downtime a handful of times a year, ideally with instances spread across availability zones.
- Active-passive fits once a customer contract includes an uptime commitment, provided you have a well-drilled failover runbook.
- Active-active fits only when your data layer can handle writes safely from several locations and the target justifies running everything twice.
- Whatever you choose, rehearse failover on a schedule so the first real test is not an actual incident.
What Good Looks Like
Failover redundancy means your recovery approach (single-region, active-passive or active-active) is a deliberate match to your actual uptime commitment, verified by drills rather than assumed from the architecture diagram.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should we actually run a failover drill?
Quarterly is a reasonable baseline for active-passive setups, with an additional drill after any significant infrastructure change. Teams that only discover failover issues during a real incident are almost always teams that have never scheduled a drill.
Can we start with active-passive and migrate to active-active later?
Yes, and it's the common path. Active-passive validates that your failover process, monitoring and runbooks actually work before you take on the harder problem of multi-location data writes that active-active requires.
Does a multi-region CDN count as failover redundancy?
It helps for static assets and cached content, but it doesn't protect your application or database layer. A CDN outage in front of a single-region backend still leaves you down if that backend region fails.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
What High Availability Really Costs, and What It Buys You
A plain-language look at the real cost of failover and redundancy, matched against what different availability targets actually mean in downtime terms.
What High Availability Actually Costs Beyond the Second Region
A worked-example breakdown of what running a second region for failover really costs, and how to decide whether your uptime target justifies it.
Active-Active vs. Active-Passive for Your Identity and Policy Layer
Comparing active-active and active-passive failover for the identity and policy services zero trust APIs depend on, with real tradeoffs on each side.
Working Out Your Pipeline's Actual Downtime Budget
A worked example of turning an availability target into a real downtime budget for a streaming pipeline, and what that means for failover design.
What an Hour of Downtime Actually Costs You
How to work out your real cost of downtime, match it to an availability target, and decide whether a second region is actually worth paying for yet.
Failover and High Availability: The Questions Worth Asking First
Straight answers on how many nines you actually need, active-active versus active-passive, and why untested failover often fails when you need it.