Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

Active-Active vs Active-Passive: What Your Uptime Target Buys You

Every jump in availability target costs roughly an order of magnitude more in infrastructure and operational discipline than the one before it, which is why picking a target without pricing out the approach behind it is how teams end up either over-building for risk they don't have or under-building for an outage they can't afford.

The three common approaches, single-region, active-passive and active-active, aren't a menu you pick from once. Most companies move through them in order as their contracts and traffic grow, and knowing what each one actually buys you helps you time that move instead of guessing at it.

What are you accepting with single-region hosting?

A single application instance in a single region is the cheapest setup and the one most small teams start with. Its failure mode is total: if the instance, the availability zone, or the region has a problem, you're down until someone (or something) restarts service elsewhere. This is a reasonable choice when your users would forgive an hour of downtime a handful of times a year, and it's a poor choice once a customer contract includes an uptime commitment you haven't actually engineered for.

Even within "single-region," running multiple instances across separate availability zones in that region is a cheap, meaningful step up: it protects against a single machine or zone failing without the cost or complexity of a second region. Plenty of teams conflate "single-region" with "single point of failure" when the actual gap is just not having spread their existing instances across zones.

Active-passive: redundancy that waits its turn

Active-passive keeps a standby copy of your infrastructure, often in a second availability zone or region, that takes over when the primary fails. The standby can be running and idle (faster failover, more cost) or provisioned on demand when a failure is detected (slower failover, less cost). Failover isn't instant: detecting the failure, promoting the standby, and redirecting traffic typically takes anywhere from under a minute to several minutes depending on how much of that process is automated versus manual.

The operational discipline this requires is easy to underestimate: a standby that's never actually been failed over to in a drill is a standby you don't really know works. Teams that skip failover drills often discover gaps (an expired credential, a config that only exists on the primary) during a real incident instead of a scheduled test.

Database replication is usually the hardest part to get right here. An asynchronous standby database can lag behind the primary by seconds or more, which means a failover under load can lose the most recent writes. Decide up front, and document, how much data loss is acceptable in a failover; that number should drive whether you need synchronous replication, which costs latency on every write, or asynchronous, which is cheaper but riskier during a failover.

Active-active: no failover step, at the cost of running everything twice

Active-active runs full production capacity in two or more locations simultaneously, splitting real traffic across them, so there's no failover step at all: traffic just shifts away from a failed location. This buys the fastest recovery of the three approaches, but it requires your data layer to handle writes safely from multiple locations, which is a genuinely hard problem for most databases and often means real architectural changes, not just infrastructure changes.

Most teams that need this level of resilience are already operating at a scale, and with a contractual uptime commitment, that justifies the engineering investment. Adopting it earlier than that usually means solving a distributed-data problem you didn't need to solve yet.

For example, a team that adopts active-active before it's needed often discovers the hard requirement only once they hit a real conflict: two users in different regions edit the same record within seconds of each other, and now the application has to decide which write wins, or how to merge them. That's a product decision as much as an infrastructure one, and it's far cheaper to make deliberately, before it's forced by an incident, than to retrofit after.

How do you match a failover approach to your uptime target?

A 99.9% target allows roughly 8.76 hours of downtime a year, while 99.99% shrinks that to about 53 minutes and 99.999% to roughly 5 minutes1. Active-passive with a well-drilled runbook can realistically hit 99.9% to 99.95% uptime1. Getting past that into 99.99% availability territory is where active-active, or at minimum a heavily automated active-passive setup with sub-minute detection, becomes necessary rather than aspirational1.

Before committing to a target, check what your actual customer contracts require. Founders regularly find they've been designing for 99.99% uptime because it sounded appropriately serious, when their signed agreements only commit to 99.9%, a target that costs meaningfully less to operate1.

Match the approach to your situation like this:

  • Single-region fits when users would forgive an hour of downtime a handful of times a year, ideally with instances spread across availability zones.
  • Active-passive fits once a customer contract includes an uptime commitment, provided you have a well-drilled failover runbook.
  • Active-active fits only when your data layer can handle writes safely from several locations and the target justifies running everything twice.
  • Whatever you choose, rehearse failover on a schedule so the first real test is not an actual incident.
Executive Capability Standard

What Good Looks Like

Failover redundancy means your recovery approach (single-region, active-passive or active-active) is a deliberate match to your actual uptime commitment, verified by drills rather than assumed from the architecture diagram.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Check your signed customer contracts for any uptime commitments you've already made, and compare that number against your current architecture's likely real-world availability.
2. Do Manually:Set up a manual active-passive failover runbook and run it as a scheduled drill at least once a quarter.
3. Delegate:Assign an infrastructure engineer to automate failure detection and standby promotion so failover time drops from a manual, multi-minute process to an automated one.
4. Automate:Build automated, tested failover into your deployment pipeline so a region or zone failure triggers recovery without a human paging in.
5. Buy:Bring in an SRE or infrastructure consultant to design active-active architecture once your data volume and uptime commitments justify the investment.

How to Get Started

Frequently Asked Questions

How often should we actually run a failover drill?

Quarterly is a reasonable baseline for active-passive setups, with an additional drill after any significant infrastructure change. Teams that only discover failover issues during a real incident are almost always teams that have never scheduled a drill.

Can we start with active-passive and migrate to active-active later?

Yes, and it's the common path. Active-passive validates that your failover process, monitoring and runbooks actually work before you take on the harder problem of multi-location data writes that active-active requires.

Does a multi-region CDN count as failover redundancy?

It helps for static assets and cached content, but it doesn't protect your application or database layer. A CDN outage in front of a single-region backend still leaves you down if that backend region fails.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides