Disaster Recovery Plan: Setting RTO and RPO per Service
A disaster recovery plan states, for each service, how long you can be down (RTO) and how much data you can lose (RPO), and the exact steps to get back. Set the targets by business impact, then choose the cheapest design that meets them.
This template walks through tiering, targets, recovery patterns, a runbook outline and testing.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What do RTO and RPO actually mean?
Recovery time objective (RTO) is the longest outage a service can have before the damage is unacceptable. Recovery point objective (RPO) is the longest window of data you can afford to lose, measured backward from the failure. An RPO of one hour means losing at most the last hour of writes.
They're independent. A reporting dashboard might tolerate a day of downtime but no data loss; a real-time feed might tolerate losing a few minutes of history but not an hour of downtime. Set each per service, not for the company as a whole, since a single blanket target either overspends or under-protects.
How to tier your services
You don't need precision, just useful buckets. A four-step exercise:
- List every service and datastore, including identity, DNS, email, and the CI system you'd need to redeploy.
- For each one, ask what breaks for customers and for revenue if it's down for one hour, one day and one week.
- Group them into three tiers: revenue-critical, important, and can wait.
- Give each tier an RTO and RPO. For example, say revenue-critical is one hour and five minutes, important is one day and one hour, and can-wait is one week and one day.
Those tier numbers are examples; yours should come from what your customers and contracts require. Link them to availability promises: a 99.99% target leaves only about 52.6 minutes of downtime per year1, so a one-hour RTO can't coexist with that promise.
Which recovery pattern fits each tier?
Patterns trade cost for speed. From cheapest to most expensive:
- Backup and restore: data is copied elsewhere, and you rebuild infrastructure from code and restore data after a disaster. Slowest, cheapest. Fits the can-wait tier.
- Pilot light: the database is replicated to a second region, but application servers stay off until needed.
- Warm standby: a scaled-down copy of the whole stack runs continuously in another region and scales up on failover.
- Active-active: traffic is served from multiple regions at once. Fastest, but also the most complex and costly, and it forces hard questions about data consistency.
Most small companies land on backup-and-restore or pilot light for the database. For managed databases such as AWS RDS, cross-region copies and read replicas are the usual building blocks; check what your configuration supports. The backup half deserves its own checklist: retention, off-account copies and timed restore drills.
What goes in the runbook?
The runbook is what you follow when nobody has time to think. Each service section should have:
- Who declares a disaster and who leads the recovery, with backups named.
- The order of operations, since identity, DNS and secrets usually have to come back before applications do.
- Exact commands or console steps for restoring data and redeploying, tested by someone other than the author.
- Where credentials and infrastructure code live if your primary account or laptop is unavailable.
- How you'll communicate with customers, using the same approach as your incident response plan.
- Verification steps to confirm the service is really healthy before you announce recovery.
Don't forget dependencies outside your control, such as a payment provider or email service. Note the workaround, or accept that they're part of your RTO.
How to test the plan without causing an outage
Start with a tabletop exercise: walk through a scenario such as 'the primary region is unavailable' and follow the runbook aloud, noting every gap. Next, run a real restore into an isolated environment and time it against your RTO. Only after that consider a live failover test in staging, then a controlled window in production if the business needs it.
Record the measured recovery time and data-loss window. If they miss the targets, you have two choices: spend more to improve the design, or change the targets with the business's agreement. Both are legitimate; ignoring the gap is not. See also MongoDB Atlas vs AWS RDS vs Supabase for how datastore choice affects recovery options, and the ransomware response plan for the security variant.
What Good Looks Like
Each service has a written RTO and RPO, a recovery pattern that meets them, and a measured test result within the past year.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
What is the difference between RTO and RPO?
RTO is how long a service can be down before harm becomes unacceptable. RPO is how much data, measured in time, you can lose. A service can have a short RTO with a long RPO, or the reverse.
Do small companies need a disaster recovery plan?
Yes, a short one. Customers and security questionnaires ask for it, and a plan forces you to find single points of failure. Two pages with tiers, targets and tested restore steps is a good start.
Is a multi-region setup required for disaster recovery?
Not always. Backup and restore, or a replicated database with a rebuilt application layer, can meet modest targets at lower cost. Multi-region designs are for services whose RTO and RPO are very short.
How often should the plan be tested?
Run a tabletop exercise at least twice a year and a real restore test quarterly for critical data. Also retest after big architecture changes or when key staff leave.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
Incident Response Plan for a Startup: A Fill-In Outline
An incident response plan outline for small engineering teams: roles, the first 15 minutes, communication steps, a security branch and a review process.
MongoDB Atlas vs AWS RDS vs Supabase: Managed Database Comparison
Compare MongoDB Atlas, AWS RDS, and Supabase for managed databases: document vs relational schemas, automated backups, developer velocity, and cloud cost.
Ransomware Response Plan: Who Does What in the First 24 Hours
An outline for a ransomware response plan: roles, first-hour containment steps, the payment question, communications and how to recover safely.
Active-Active vs Active-Passive: What Your Uptime Target Buys You
A comparison of active-active, active-passive and single-region failover, with the real infrastructure and headcount cost each uptime target requires.
What an Hour of Downtime Actually Costs You
How to work out your real cost of downtime, match it to an availability target, and decide whether a second region is actually worth paying for yet.
Moving to the Cloud: A Migration Plan for a Small Business
A phased plan for a small business cloud migration: inventory, choose a strategy per system, build a landing zone, pilot, cut over and retire the old.