Active-Active vs. Active-Passive for Your Identity and Policy Layer
In a zero trust setup, your identity provider and policy engine aren't optional infrastructure, they're on the critical path of every request. If they go down, it's not a degraded experience, it's every API call failing at once, which makes their failover strategy worth as much attention as the APIs they protect.
The choice is usually between active-active (multiple regions serving traffic simultaneously) and active-passive (a standby that takes over on failure), and each has real costs the other doesn't.
What active-passive actually costs you
Active-passive is cheaper to run day to day, since the standby doesn't need to handle production load, but it shifts the cost into the failover moment itself: detection time, promotion time, and the risk that the standby has quietly drifted out of sync with the primary because nobody's watching it under real load. The failover you never test is the failover that fails when you need it.
If you go this route, schedule regular, real failover drills, not tabletop exercises, where you actually promote the standby and route production traffic to it, then measure how long it took and what broke.
What active-active actually costs you
Active-active removes the failover moment (there's no single point that has to notice failure and promote a standby) but it adds real, ongoing complexity: your policy decisions and identity state have to stay consistent across regions in something close to real time, or you risk a request being allowed in one region and denied in another for the same caller at the same moment. That consistency problem is the actual cost, and it's a permanent engineering burden, not a one-time setup cost.
This is usually worth it once you're operating at a scale where a full regional failure is a realistic scenario you plan for, not a hypothetical.
Match the choice to the availability tier you actually need
A downtime budget at 99% availability tolerates an active-passive setup with a slow, careful failover, but that downtime budget shrinks fast once you move toward 99.99% availability or tighter, where a manual or even semi-automated promotion process can burn through the whole allowance in a single incident1. Decide your real availability target first, based on what your customers actually need and what you're contractually committing to, then choose the architecture that target requires, rather than defaulting to whichever pattern is more familiar to your team.
For example, a small B2B team with two engineers on call and a modest availability commitment might start with active-passive and run a real promotion drill every quarter. If a drill shows that detection and promotion take longer than the downtime budget allows, that result is the evidence for moving to active-active, not a vague sense that redundancy feels safer. A useful decision rule: pick the simplest pattern whose measured failover time fits inside the allowance your contracts imply, and revisit the choice when the commitment tightens.
Test the failure mode that actually matters: partial degradation
Full regional outages get all the attention in failover planning, but partial degradation, your identity provider responding slowly rather than failing outright, is both more common and harder to handle well. A slow policy engine can make every request wait rather than fail cleanly, which is often worse for users than an outright outage they can at least detect and route around.
Build a specific test for this: inject artificial latency into your identity or policy service in a staging environment and watch what your APIs do. If they queue requests indefinitely instead of timing out and falling back to a cached decision or a safe default, that's a gap independent of which failover architecture you chose.
Decide your fail-open versus fail-closed default in advance
When the identity or policy service is unreachable, does the request get denied (fail closed, safer but can cause a broad outage from a single dependency failure) or allowed with a cached or default decision (fail open, keeps things running but weakens the guarantee during the outage)? This decision needs to be made deliberately and in advance, per endpoint if necessary, not discovered live during an incident when someone is under pressure to bring things back up.
Financial and destructive-action endpoints should almost always fail closed. Read-only, low-sensitivity endpoints are a more reasonable candidate for a fail-open exception, if you make that call before the outage rather than during it.
Settle these points before you pick a failover pattern:
- Your real availability target, based on what customers need and what you have committed to in contracts.
- Whether your team can run real failover drills that promote the standby and route production traffic to it.
- Whether policy decisions and identity state can stay consistent across regions in close to real time.
- How your APIs behave when the identity provider is slow instead of down, including timeouts and cached fallbacks.
- Which endpoints fail closed and which, if any, are allowed a fail-open exception, decided before an outage.
What Good Looks Like
Good practice means your identity and policy layer's failover behavior, both which pattern you use and your fail-open versus fail-closed defaults, is a deliberate decision matched to a stated availability target, and it's actually tested, not just documented.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
Which failover pattern should a small engineering team choose first?
Active-passive, because the ongoing operational complexity of active-active usually isn't worth it until you're at a scale where a full regional failure is something you plan for as a real scenario. Get comfortable running real failover drills on active-passive before taking on the harder cross-region consistency problem.
How often should we actually test failover, not just plan for it?
At minimum quarterly for a production-critical identity or policy layer, and after any significant architecture change to either service. A failover test that reveals nothing wrong every single time for years is itself worth a second look, since it may mean the test isn't actually exercising a real failure.
Should every endpoint use the same fail-open or fail-closed default?
No. The right default depends on what the endpoint does. A payment or permissions-change endpoint should fail closed by default; a low-sensitivity, read-only endpoint is a more defensible candidate for fail-open. Set this per endpoint deliberately rather than applying one blanket rule everywhere.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Continuous Device Verification for a Zero-Trust API
How continuous device and identity verification actually works in a zero-trust architecture, and where to draw the line for a small engineering team.
Active-Active vs Active-Passive: What Your Uptime Target Buys You
A comparison of active-active, active-passive and single-region failover, with the real infrastructure and headcount cost each uptime target requires.
What High Availability Actually Costs Beyond the Second Region
A worked-example breakdown of what running a second region for failover really costs, and how to decide whether your uptime target justifies it.
Rolling Out Zero Trust in Production Without a Broad Outage
A checklist for rolling out stricter API authentication and authorization in production, and the pitfalls that turn a rollout into an incident.
What High Availability Really Costs, and What It Buys You
A plain-language look at the real cost of failover and redundancy, matched against what different availability targets actually mean in downtime terms.
What an Hour of Downtime Actually Costs You
How to work out your real cost of downtime, match it to an availability target, and decide whether a second region is actually worth paying for yet.