What an Hour of Downtime Actually Costs You
Work out what an hour of downtime costs you before deciding how much redundancy to buy. A second region, active failover and multi-zone everything all reduce risk, but they also add real money and real complexity every month, whether or not you ever use them.
The way to decide how much redundancy is enough is to work out what an hour of downtime actually costs you first, then buy protection against that number.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Work out your own cost of downtime before buying redundancy
Say your product generates revenue continuously through the day. Estimate what an hour of full outage costs in lost transactions, plus any support burden and customer trust cost that's harder to put a number on but still real. That figure is what you're actually protecting against, not an abstract goal of "more uptime."
A team that skips this step tends to either over-invest in redundancy they don't need yet, or under-invest and get surprised by how expensive a bad outage actually was.
Match your availability target to your downtime budget
Every availability target implies an exact downtime budget, whether you've calculated it or not. A 99.9 percent uptime target leaves you about 8.76 hours of downtime a year to work with, only about 0.365 days1. Chasing one more nine of availability shrinks that budget by roughly a factor of ten each time, and the cost of the infrastructure to get there tends to grow at least as fast.
Pick the target that matches what an outage actually costs you, not the highest number that sounds impressive in a pitch deck.
Active-active versus active-passive: the tradeoff that matters
Active-passive keeps a standby ready to take over, which is cheaper and simpler but means a real failover event, including detecting the failure and cutting traffic over, before service is restored. Active-active runs both regions serving real traffic all the time, which removes that failover gap but requires your data layer to handle writes safely from two places at once, which is a genuinely harder engineering problem.
Most teams should start with active-passive and only move to active-active once the failover gap itself, not just the redundancy, is the thing costing real money.
A worked example: is a second region worth it yet
If an hour of downtime costs you a modest amount and happens rarely, the ongoing cost of running a second full region can easily exceed what you're protecting against over a year. If an hour of downtime costs you a large amount, or happens to coincide with your busiest sales periods, the math flips quickly, and a second region pays for itself the first time it's actually needed.
Run this comparison with your own numbers before committing, rather than defaulting to whatever a well-funded competitor happens to run.
Testing failover on a schedule instead of hoping it works
Redundancy you haven't tested is a theory, not a plan. Schedule an actual failover drill, not just a review of the architecture diagram, and treat a drill that reveals a gap as a success, since that's exactly what it's for.
- Pick a recurring date for a failover drill, not "whenever there's time"
- Fail over during business hours the first time, so people are around to catch problems
- Write down what broke and fix it before the next drill, not just the next real incident
- Confirm your monitoring actually detects the failover, not just that the failover itself worked
What backups alone don't cover
A solid backup strategy protects your data. It doesn't protect your uptime, because restoring from a backup takes time, and every minute of that restore is still an outage from your users' point of view. Teams sometimes treat backups as their redundancy plan and only discover the gap when a real failure happens and the restore takes far longer than anyone expected.
Treat backups and failover as two separate problems with two separate answers: backups are your defense against losing data permanently, and failover, or a fast, tested recovery process at minimum, is your defense against being down for longer than you can tolerate. A strong answer to one doesn't cover for a weak answer to the other.
What Good Looks Like
Good here means you know what an hour of downtime actually costs you, you've picked an availability target that matches that number, and you've tested your failover mechanism recently enough to trust it.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Do we need a multi-region setup if we're a small company?
Not automatically. It depends on what an hour of downtime actually costs you and how often you expect to need it. Many small teams are better served by a solid single-region setup with tested backups and a fast recovery plan than by the ongoing cost of a second region.
How often should we actually test failover?
At least twice a year, and after any significant change to the systems involved in failing over. A failover mechanism that hasn't been tested since it was built is likely to have drifted out of sync with how the system actually runs today.
What's the biggest mistake teams make with redundancy?
Buying redundancy based on what sounds prudent rather than what an actual outage would cost them, then never testing whether the redundancy they bought actually works when triggered. Both mistakes are common, and either one on its own can waste the whole investment.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
What High Availability Actually Costs Beyond the Second Region
A worked-example breakdown of what running a second region for failover really costs, and how to decide whether your uptime target justifies it.
What High Availability Really Costs, and What It Buys You
A plain-language look at the real cost of failover and redundancy, matched against what different availability targets actually mean in downtime terms.
Active-Active vs Active-Passive: What Your Uptime Target Buys You
A comparison of active-active, active-passive and single-region failover, with the real infrastructure and headcount cost each uptime target requires.
Failover and High Availability: The Questions Worth Asking First
Straight answers on how many nines you actually need, active-active versus active-passive, and why untested failover often fails when you need it.
How Much Redundancy Your Vector Store Actually Needs
Replicated indexes, snapshot restore, and multi-region setups each buy different recovery guarantees. Match the approach to your actual uptime target.
Working Out Your Pipeline's Actual Downtime Budget
A worked example of turning an availability target into a real downtime budget for a streaming pipeline, and what that means for failover design.