Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

What an Hour of Downtime Actually Costs You

Work out what an hour of downtime costs you before deciding how much redundancy to buy. A second region, active failover and multi-zone everything all reduce risk, but they also add real money and real complexity every month, whether or not you ever use them.

The way to decide how much redundancy is enough is to work out what an hour of downtime actually costs you first, then buy protection against that number.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Work out your own cost of downtime before buying redundancy

Say your product generates revenue continuously through the day. Estimate what an hour of full outage costs in lost transactions, plus any support burden and customer trust cost that's harder to put a number on but still real. That figure is what you're actually protecting against, not an abstract goal of "more uptime."

A team that skips this step tends to either over-invest in redundancy they don't need yet, or under-invest and get surprised by how expensive a bad outage actually was.

Match your availability target to your downtime budget

Every availability target implies an exact downtime budget, whether you've calculated it or not. A 99.9 percent uptime target leaves you about 8.76 hours of downtime a year to work with, only about 0.365 days1. Chasing one more nine of availability shrinks that budget by roughly a factor of ten each time, and the cost of the infrastructure to get there tends to grow at least as fast.

Pick the target that matches what an outage actually costs you, not the highest number that sounds impressive in a pitch deck.

Active-active versus active-passive: the tradeoff that matters

Active-passive keeps a standby ready to take over, which is cheaper and simpler but means a real failover event, including detecting the failure and cutting traffic over, before service is restored. Active-active runs both regions serving real traffic all the time, which removes that failover gap but requires your data layer to handle writes safely from two places at once, which is a genuinely harder engineering problem.

Most teams should start with active-passive and only move to active-active once the failover gap itself, not just the redundancy, is the thing costing real money.

A worked example: is a second region worth it yet

If an hour of downtime costs you a modest amount and happens rarely, the ongoing cost of running a second full region can easily exceed what you're protecting against over a year. If an hour of downtime costs you a large amount, or happens to coincide with your busiest sales periods, the math flips quickly, and a second region pays for itself the first time it's actually needed.

Run this comparison with your own numbers before committing, rather than defaulting to whatever a well-funded competitor happens to run.

Testing failover on a schedule instead of hoping it works

Redundancy you haven't tested is a theory, not a plan. Schedule an actual failover drill, not just a review of the architecture diagram, and treat a drill that reveals a gap as a success, since that's exactly what it's for.

  • Pick a recurring date for a failover drill, not "whenever there's time"
  • Fail over during business hours the first time, so people are around to catch problems
  • Write down what broke and fix it before the next drill, not just the next real incident
  • Confirm your monitoring actually detects the failover, not just that the failover itself worked

What backups alone don't cover

A solid backup strategy protects your data. It doesn't protect your uptime, because restoring from a backup takes time, and every minute of that restore is still an outage from your users' point of view. Teams sometimes treat backups as their redundancy plan and only discover the gap when a real failure happens and the restore takes far longer than anyone expected.

Treat backups and failover as two separate problems with two separate answers: backups are your defense against losing data permanently, and failover, or a fast, tested recovery process at minimum, is your defense against being down for longer than you can tolerate. A strong answer to one doesn't cover for a weak answer to the other.

Executive Capability Standard

What Good Looks Like

Good here means you know what an hour of downtime actually costs you, you've picked an availability target that matches that number, and you've tested your failover mechanism recently enough to trust it.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Work out a real estimate of what an hour of downtime costs your business, in lost revenue and support burden.
2. Do Manually:Run a scheduled failover drill by hand, during business hours, and document exactly what broke.
3. Delegate:Assign one engineer ownership of the failover runbook and the drill calendar, so it doesn't quietly lapse.
4. Automate:Automate health checks and failover triggers so a real failure cuts over without someone needing to notice and act manually first.
5. Buy:Bring in outside infrastructure help to design a second region once your own numbers show it's actually worth the ongoing cost.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

ClickUp

ClickUp works well for keeping a running record of who's due to run the next failover drill and confirming it actually happened on schedule.

Visit ClickUp→

Frequently Asked Questions

Do we need a multi-region setup if we're a small company?

Not automatically. It depends on what an hour of downtime actually costs you and how often you expect to need it. Many small teams are better served by a solid single-region setup with tested backups and a fast recovery plan than by the ongoing cost of a second region.

How often should we actually test failover?

At least twice a year, and after any significant change to the systems involved in failing over. A failover mechanism that hasn't been tested since it was built is likely to have drifted out of sync with how the system actually runs today.

What's the biggest mistake teams make with redundancy?

Buying redundancy based on what sounds prudent rather than what an actual outage would cost them, then never testing whether the redundancy they bought actually works when triggered. Both mistakes are common, and either one on its own can waste the whole investment.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides