What High Availability Really Costs, and What It Buys You
Every additional nine of availability costs more than the last one, and past a certain point the spend stops making sense for most companies outside a handful of regulated industries. The hard part is not building redundancy, it is deciding how much of it your product and your customers actually need.
This is a way to match your availability target against its real cost, so the decision to add or skip a standby database or a second region is based on a number instead of a feeling.
What each availability target actually means in downtime
The nines everyone quotes translate to very different real-world numbers. A 99% target allows about 3.65 days of downtime a year, while a 99.9% target allows a little under 8.76 hours, and a 99.99% uptime target allows only around 52.6 minutes1. The jump between those last two tiers is the one that usually forces a real architecture change, since a budget measured in minutes leaves almost no room for a slow, manual failover process.
Most B2B products with defined business hours and no per-minute contractual penalty can run comfortably at that middle target. Going further is a real engineering investment, and it should be a deliberate choice tied to a specific customer commitment, not a default picked because it sounds more serious.
What redundancy actually costs, in headcount and infrastructure
A standby database replica, a second region, and the automation to fail between them cost money every month whether or not you ever use them, plus the ongoing engineering time to keep the failover path working as the system changes. That is the real price of high availability: not the initial build, but the continuous maintenance that keeps a failover path from silently rotting.
Budget for both when deciding whether the next nine is worth it. A failover system nobody has tested in six months because the team moved on to other work is a liability dressed up as insurance.
Testing the failover you are already paying for
The return on investment for redundancy only exists if the failover actually works when triggered for real. Schedule a failover drill on a low-traffic day, treat it like a fire drill, and fix whatever breaks. Most teams find at least one thing wrong the first time they actually try it, a stale connection string, a missing environment variable, a runbook step that assumes a person who left the company.
A drill that surfaces problems on a quiet Tuesday afternoon is doing its job. The alternative is finding those same problems during a real outage, at a much higher cost.
Run the drill in this order:
- Schedule the drill on a low-traffic day and tell the team it is a fire drill, not a surprise.
- Trigger the failover the way you would in a real incident, following the written runbook step by step.
- Note everything that breaks, such as a stale connection string, a missing environment variable or a runbook step that assumes someone who left.
- Fix each problem, update the runbook, and repeat the drill after any significant architecture change.
When to skip the next nine entirely
If your product is used during business hours by a small number of accounts, a planned maintenance window and a fast recovery process may serve customers better than round-the-clock automatic failover that costs meaningfully more every month. Ask what a customer actually loses during a short maintenance window late on a weekend, and size your investment to that answer, not to a target borrowed from a much larger company's requirements.
For example, consider a product used by a small number of accounts during business hours. Those customers may lose very little during a planned maintenance window late on a weekend, so a well-rehearsed recovery process may serve them better than a costly standby region. The decision rule: write down what a customer actually loses per minute of downtime, compare that with the monthly cost and upkeep of the redundancy, and add more redundancy only when the loss is larger. Revisit the answer when a new contract commits you to a specific uptime target. A common mistake is copying a target from a much larger company without asking whether your customers would notice the difference.
A concrete example of what the next nine requires
Say your team currently recovers from a regional outage by having an engineer manually redirect traffic and restore from backup, a process that reliably takes forty-five minutes end to end. That works fine against a 99.9% uptime target, which allows a little under 8.76 hours of downtime a year1, but it would burn through an entire year's downtime budget in a single incident if the target moved to 99.99%, where the whole year's allowance is only about 52.6 minutes1.
Closing that gap means automating the failover so it takes seconds instead of minutes, which is a meaningfully bigger engineering investment than most teams expect when they first see the target written as a percentage instead of as a number of minutes.
What Good Looks Like
Good here means you can state your actual availability target in downtime minutes per year, not just as a percentage, and your last failover test happened within the last six months with the results written down.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
What availability target should a typical B2B SaaS company aim for?
A 99.9% uptime target is a reasonable default for most B2B products without per-minute contractual penalties, since it still allows a little under 8.76 hours of downtime a year1 while staying achievable without a full multi-region build. Go higher only when a specific customer contract or use case requires it.
How often should we test our failover process?
At minimum twice a year, and immediately after any significant architecture change to the systems involved. A failover path that worked cleanly at your last test can silently break after a database migration or a change to how services authenticate with each other.
Is multi-region redundancy worth it for an early-stage company?
Usually not yet. The engineering time is better spent on product until you have specific customers whose contracts require it or a genuine pattern of costly regional outages. Multi-region architecture is easier to add once you understand your real traffic and failure patterns than to guess at upfront.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Active-Active vs Active-Passive: What Your Uptime Target Buys You
A comparison of active-active, active-passive and single-region failover, with the real infrastructure and headcount cost each uptime target requires.
What High Availability Actually Costs Beyond the Second Region
A worked-example breakdown of what running a second region for failover really costs, and how to decide whether your uptime target justifies it.
Working Out Your Pipeline's Actual Downtime Budget
A worked example of turning an availability target into a real downtime budget for a streaming pipeline, and what that means for failover design.
What an Hour of Downtime Actually Costs You
How to work out your real cost of downtime, match it to an availability target, and decide whether a second region is actually worth paying for yet.
Failover and High Availability: The Questions Worth Asking First
Straight answers on how many nines you actually need, active-active versus active-passive, and why untested failover often fails when you need it.
How Much Redundancy Your Vector Store Actually Needs
Replicated indexes, snapshot restore, and multi-region setups each buy different recovery guarantees. Match the approach to your actual uptime target.