Working Out Your Pipeline's Actual Downtime Budget
Your pipeline's downtime budget is the yearly downtime your availability target allows, and it should decide how much failover infrastructure you build. Most teams pick a target because it sounds serious, without working out what it costs to hit or what it buys in return.
Here's a worked example: pick a target, translate it into a real downtime budget, and use that budget to decide what your pipeline's failover setup actually needs to do.
How do you turn an availability target into hours of downtime?
Availability targets sound abstract until you convert them into a downtime budget: at 99 percent that is about 3.65 days of allowed downtime a year, at 99.9 percent about 8.76 hours, at 99.95 percent about 4.38 hours, at 99.99 percent about 52.6 minutes, and at 99.999 percent about 5.26 minutes1.
Each step up costs meaningfully more engineering effort and infrastructure than the last, so pick the target your pipeline's downstream dependents actually need, not the most impressive-sounding number. A pipeline feeding an internal analytics dashboard doesn't need the same budget as one feeding a live payment authorization flow.
Multi-AZ brokers are the floor, not the whole answer
Running brokers across multiple availability zones protects against a single zone failure, which is the most common infrastructure failure mode, but it doesn't protect against a bad deploy, a schema change that breaks every consumer at once, or a regional outage. Multi-AZ is table stakes; it's not a complete failover strategy on its own.
If your downtime budget is tight enough to require protection against a regional failure, that means active replication to a second region, which is a substantially bigger commitment in both cost and operational complexity than multi-AZ alone. Most pipelines don't actually need this; check your real budget from the last step before building it.
What RTO and RPO does your pipeline need?
Recovery time objective (how long you can be down) and recovery point objective (how much data you can afford to lose) are different numbers, and a failover plan needs both. A stream processor that fails over instantly but loses the last few seconds of unprocessed events has a great RTO and a nonzero RPO; whether that's acceptable depends entirely on what the data is.
Say your pipeline feeds a real-time inventory count: losing a few seconds of updates during a failover is recoverable on the next sync. If it feeds a payment ledger, that same gap is not acceptable, and your failover design needs synchronous replication or an equivalent guarantee, not just a fast restart.
Test failover the way you'd test a fire drill, not a checklist
A failover plan that's never been exercised tends to fail on details nobody thought to write down: a hardcoded broker address, a consumer group that doesn't rejoin cleanly, a downstream system that can't handle a brief burst of reprocessed events. Run a real failover test on a non-production environment, on a schedule, not just after an incident forces you to.
Measure your actual RTO during that drill against the target you set, not the one you assumed. The gap between the two is usually where the next quarter's failover work should go.
Decide what happens to in-flight events during the switch
The moment a failover happens, some events are mid-processing: read from the topic but not yet written to their destination. Decide in advance whether your consumers are built to safely reprocess an event they already handled (an idempotent write, keyed on an event ID so a repeat doesn't double-count) or whether a failover risks duplicate side effects downstream.
Idempotency is usually the cheaper fix than trying to guarantee exactly-once delivery through the failover itself, and it pays off beyond disaster recovery too, since the same protection covers ordinary consumer restarts and retries. If a downstream write can't be made idempotent (a webhook that triggers an external charge, say), that's the one path that needs its own dedupe logic before you can trust failover for it.
Write down who declares a failover, not just how to run it
The technical steps matter less than most teams assume compared with a clear answer to a simpler question: who has the authority to declare a failover and start the process, at any hour, without waiting for a meeting. Ambiguity here costs more real downtime than any missing runbook step.
Name the role, not a specific person, so the answer still holds when that person is asleep or on vacation, and make sure whoever is on call actually knows they hold that authority before the night they need to use it.
To turn an availability target into a failover plan, work through these steps:
- Pick the availability target your downstream dependents actually need and convert it into a yearly downtime budget.
- Run brokers across multiple availability zones as the baseline, then judge whether your budget justifies more.
- Set both a recovery time objective and a recovery point objective for the data the pipeline carries.
- Make consumers idempotent, keyed on an event ID, so events in flight during the switch can be safely reprocessed.
- Rehearse failover in a non-production environment on a schedule, not only after an incident.
- Name the role that can declare a failover at any hour without waiting for a meeting.
What Good Looks Like
Failover is production ready when the pipeline has an explicit availability target translated into a downtime budget, defined RTO and RPO numbers, and a failover plan that's been tested on a schedule, not just written down.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How do we pick the right availability target for our pipeline?
Work backward from what actually depends on it. A pipeline feeding a customer-facing, real-time feature needs a tighter target than one feeding an internal dashboard refreshed hourly. Match the target, and the cost of hitting it, to the actual business impact of downtime, not to what sounds appropriately impressive in a planning document.
Is multi-region replication worth it for most teams?
Usually not, unless your downtime budget is tight enough that a regional cloud outage would be unacceptable, or a regulation requires it. Multi-AZ redundancy covers the far more common failure mode of a single zone going down, at a fraction of the cost and complexity of full regional replication.
What's the difference between RTO and RPO in practice?
RTO is how long you're down before service resumes; RPO is how much data you lose in the process. A fast failover with data loss and a slow failover with none are both valid designs depending on what the pipeline carries, which is why both numbers need to be set deliberately, not left implicit.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Active-Active vs Active-Passive: What Your Uptime Target Buys You
A comparison of active-active, active-passive and single-region failover, with the real infrastructure and headcount cost each uptime target requires.
What High Availability Really Costs, and What It Buys You
A plain-language look at the real cost of failover and redundancy, matched against what different availability targets actually mean in downtime terms.
What High Availability Actually Costs Beyond the Second Region
A worked-example breakdown of what running a second region for failover really costs, and how to decide whether your uptime target justifies it.
Blue-Green, Canary, or Rolling: Deploying Stream Processors
A decision guide to rolling, blue-green, and canary deploys for stateful stream processors, plus the rollback plan most teams never actually test.
Verifying Every Service That Talks to Your Pipeline
Which parts of zero-trust verification to build and which to buy, so every producer and consumer on a streaming pipeline proves its identity.
Retry, Circuit Break, or Dead-Letter: Handling a Failing Consumer
A comparison of retries, circuit breakers, and dead-letter queues for a failing stream consumer, and how to combine them without masking a real outage.