Cloud FinOps & Infrastructure ScalingPlaybook3 min readUpdated September 2026

The Backup You've Never Restored Isn't a Backup

A backup job with a green checkmark every night for two years feels like safety. It isn't, not on its own. A backup is only as good as the restore nobody's actually tried, and the first time most teams discover a gap, a missing table, a broken script, a restore that takes six hours instead of six minutes, is during the incident it was supposed to solve.

The fix isn't a better backup tool. It's a drill you actually run.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Is a successful backup job the same as a restorable backup?

The job succeeding means the write completed and the file landed in storage. It says nothing about whether that file is structurally valid, whether the restore script still matches your current schema, or whether whoever needs to run it during an incident knows how. Treat 'backup succeeded' and 'restore verified' as two separate facts you track, not one. Most teams only ever measure the first.

For example, a team can track two separate columns: last successful backup job and last verified restore. The first column is green every day, while the second reads never. That gap is the finding. Closing it takes one scheduled drill that restores a copy into an isolated environment, checks a few known records, and writes the outcome next to the date. Until that drill exists, the honest statement about recoverability is that it hasn't been proven, no matter how many green checkmarks the backup job has collected.

Restore into an isolated environment, never toward production

A drill that restores into a throwaway environment, a fresh database instance nobody else is using, lets you test the real mechanics without any risk to what's actually running. Automate the teardown afterward so the drill environment doesn't linger as an unmonitored, unpatched copy of production data sitting around indefinitely. If your data includes anything sensitive, treat the drill environment with the same access controls as production itself, not a lighter version.

How do you measure how long a restore really takes?

The number that matters during a real incident is how long the restore takes, not whether it's technically possible. A restore that works but takes six hours is a very different commitment than one that takes six minutes, especially against a tight downtime budget: a 99.95% uptime target leaves roughly 4.38 hours a year to spend across every incident combined1, so a single slow restore can burn most of an entire year's allowance. Write the actual timed result down after every drill, not an estimate.

Test the restore your real incident would need, not the easy one

Restoring last night's full backup is the easy case. The harder, more realistic ones are point-in-time restore to a specific minute before a bad migration ran, or restoring a single table without touching everything else. If your only tested path is 'restore everything from last night,' you don't actually know whether the narrower, more common request, get this one table back to ten minutes ago, is even possible with your current tooling. Say a bad migration corrupts a customer's order history at 2 p.m.: the request you'll actually get is 'get this one table back to before 2 p.m. without touching anything else,' not 'restore the whole database from last night.' If that narrower request has never been tested, you won't find out it's unsupported until someone's asking for it during the incident itself.

Common pitfall: a runbook only the person who wrote it can follow

A restore runbook that references institutional knowledge, the person who set this up left eight months ago, and their shorthand notes assume context nobody else has, is a runbook that fails exactly when it's needed most, during an incident when the original author might be asleep, on vacation, or gone. Have someone who didn't write the runbook run the next drill from it, cold. Every place they get stuck is a gap to fix before it's a real outage.

Put drills on a calendar, and log every one the same way

Quarterly is a reasonable default cadence for most systems, monthly for anything that would be genuinely painful to lose. A tool like ClickUp can hold the drill calendar and the result of each one, timed duration, what broke, what got fixed, so a pattern across drills is visible instead of scattered across whoever happened to run each one. Compliance platforms such as Vanta can often ingest that same drill log as evidence when a SOC 2 auditor asks whether disaster recovery is actually tested, not just documented as a policy; confirm the integration you need in the vendor's current documentation.

A useful restore drill follows these steps:

  1. Restore into a throwaway environment that nobody else uses, never toward production.
  2. Time the restore from start to finish and write the actual result down, not an estimate.
  3. Test the request a real incident would produce, such as point in time recovery or a single table.
  4. Have someone who didn't write the runbook run the drill from it, cold.
  5. Tear the drill environment down afterward and log what broke and what got fixed.
Executive Capability Standard

What Good Looks Like

Good backup and restore practice means a timed, logged restore drill run on a regular calendar, tested by someone other than whoever set the backup up, covering both a full restore and a narrower, more realistic partial one.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Check when your backups were last actually restored, not just when the last backup job succeeded, and whether anyone timed it.
2. Do Manually:Run one full restore drill by hand into an isolated environment and time it from start to finish.
3. Delegate:Assign an owner for the drill calendar and results log so drills happen on schedule regardless of who's available that quarter.
4. Automate:Automate the drill environment's provisioning and teardown so running a drill doesn't require manual setup each time.
5. Buy:Bring in outside help to design the restore process if nobody on the team has tested a point-in-time or partial restore before.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How often should we actually run a restore drill?

Quarterly is a reasonable default for most production systems, monthly for anything where a day of lost data would be genuinely painful. The right cadence depends more on how often your schema and data actually change than on a fixed calendar rule: fast-changing systems drift out of a tested state faster.

Does a green backup job mean the backup is restorable?

No. It confirms the write completed, not that the file is structurally valid, that the restore script still matches your current schema, or that anyone knows how to run it under pressure. Those are separate claims, and only an actual restore drill tests the second one.

What's the most common blind spot in restore testing?

Only ever testing the easy case: a full restore of last night's backup into a clean environment. Real incidents more often need a narrower restore, one table, a specific point in time before a bad migration, and teams that have never tested that path often discover it isn't actually supported by their current tooling.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides