The Backup You Haven't Tested Is Just a Hope
A backup job that completes successfully tells you a file was written somewhere. It tells you nothing about whether that file can rebuild a working database. The only way to know is to actually restore it, on a schedule, and treat that restore as the real test.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Step 1: Restore to an Isolated Environment, Not Production
Spin up a separate environment for the restore, never test against production and never test in a way that could touch real customer traffic. This sounds obvious, but under time pressure, teams cut this corner and it's exactly the corner that turns a routine drill into a real incident.
An isolated restore target also means you can run the drill on a schedule without coordinating around production maintenance windows, which is a big part of why these drills stop happening in the first place.
Step 2: Measure Actual Restore Time, Not Backup Completion Time
The number that matters during a real incident is how long it takes to get back to a working state, not how long the backup job took to run. Time the full restore from scratch: provisioning the target, pulling the backup, running it, and verifying the schema and row counts look right.
That's your real recovery time, and it should be measured against your allowed downtime budget1. Say your restore takes six hours end to end: against a target that allows only a few hours of downtime a year, that's a recovery plan that fails on paper before it's ever tested for real.
For example, the backup job reports success in about forty minutes every night, so the team assumes recovery is quick. A drill shows that provisioning a new database takes an hour, pulling the file takes another, and verification takes a third. Real recovery time is several hours, not forty minutes. Record that end-to-end number and compare it with what your uptime commitment allows. If it doesn't fit, the fix might be a warm standby or a faster restore path, and you found that out on a quiet afternoon rather than during an outage.
Step 3: Verify Data Integrity, Not Just That the Restore Ran
A restore that completes without an error can still be missing recent transactions, have corrupted rows, or be silently out of sync with what your application expects. Run a set of specific checks after every restore: row counts against known baselines, a handful of specific records you can verify by hand, and any foreign key constraints that would flag corruption.
Write these checks down once and reuse them every drill, so verification doesn't depend on someone remembering what "looks right" means for your particular schema, and so a different engineer running the drill next quarter gets the same thorough, repeatable check every single time, no matter who happens to be running the drill that quarter.
Step 4: Rotate Who Runs the Drill
If only one engineer has ever run a restore, you don't have a tested recovery process, you have one person's private knowledge that happens to work. Rotate the drill across the team so more than one person has actually done it under realistic conditions, with the runbook as the only guide, not that person's memory filling in the gaps.
This also surfaces gaps in the written runbook fast: if a different engineer gets stuck on step three, that's a real documentation problem worth fixing before an actual incident exposes it at 2am.
Step 5: Schedule It Like You Mean It
A restore drill that's "on the roadmap" without a specific date on the calendar doesn't happen. Put it on a recurring quarterly schedule, treat a skipped drill as a real gap worth flagging, not a minor miss, and track the last successful restore date somewhere visible so nobody's guessing how stale your recovery confidence actually is.
What to Do When a Drill Actually Fails
A failed drill is a good outcome in the sense that it found the gap before a real incident did. Treat it exactly like a production incident, with a written root cause and a specific fix, not a quiet retry until it happens to work.
Teams sometimes treat a failed drill as an embarrassment to fix quietly and move on from. That instinct is backwards: a documented, understood failure from a drill is far cheaper than the same failure discovered during a real outage, and it deserves the same rigor a real incident review would get, complete with an owner and a follow-up deadline.
After each drill, confirm the following before closing it out:
- The restore ran in an isolated environment that could not touch production or real customer traffic.
- Total restore time, from provisioning through verification, is recorded and compared with your downtime budget.
- Row counts, hand-verified records, and foreign key checks passed against your written baseline.
- Someone other than the usual engineer ran the drill using only the runbook.
- The next drill has a date on the calendar, and any failure has an owner and a deadline.
What Good Looks Like
A real backup strategy is proven by a timed, verified restore run on a fixed schedule by more than one person, not by a green checkmark on the backup job itself.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How often should we actually run a restore drill?
Quarterly is a reasonable baseline for most small teams, more often for anything storing data you genuinely couldn't afford to lose or be down for long. The exact cadence matters less than actually keeping to whatever schedule you set.
Is testing backups on a staging environment with real data safe?
Only if that staging environment already meets the same access controls and data handling rules as production. If it doesn't, restore to a fully isolated, locked-down environment instead, or scrub the data as part of the restore process before anyone else can access it.
What's the biggest mistake teams make with backup verification?
Trusting the backup job's own success status as proof the backup works. A completed job confirms a file was written, not that the file can rebuild a working, correct database. Only an actual restore confirms that.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
A Backup You Haven't Restored From Is Just a File
A backup job that succeeds every night tells you nothing about whether a restore will actually work. Here is a runbook for testing the part that matters.
A Runbook for Verifying Database Backups Actually Restore
A step-by-step runbook for proving your database backups restore cleanly, run on a schedule instead of trusted on faith until a real outage.
Proving Your Pipeline Backups Actually Restore
A runbook for actually testing that your event pipeline's backups restore cleanly, instead of trusting a green checkmark on a backup job.
The Backup You've Never Restored Isn't a Backup
A nightly backup job that succeeds every night tells you almost nothing about whether you can actually recover. The drill format that closes that gap.
Why Your Backups Might Not Actually Restore
How to verify database backups actually restore, how often to run restore drills, and what to measure besides pass or fail so you trust them.
A Runbook for Proving Your Backups Actually Restore
A step-by-step way to verify database backups actually restore, on a schedule, instead of discovering a gap the first time you need a backup for real.