A Backup You Haven't Restored From Is Just a File
Every team with a database has a backup job, and most of those teams have never actually restored from one. The backup job succeeding tells you a file was written somewhere. It tells you nothing about whether that file can rebuild a working database, how long that rebuild takes, or whether anyone on the team knows the steps well enough to run them at two in the morning.
The only way to know a backup actually works is to restore from it, on a schedule, before an outage forces you to find out for the first time under pressure.
Why a Green Backup Job Doesn't Mean a Working Restore
A backup job can succeed for weeks while writing a file that's silently corrupted, incomplete, or missing a table that got added after the backup script was last updated. Nothing about a green checkmark in your job scheduler checks any of that, because the job's own definition of success is just "the write completed without an error."
The only test that actually validates a backup is using it to rebuild a working database somewhere else, and comparing what comes out against what you expected to see. Anything short of that is trusting the file, not verifying it.
Setting a Real Recovery Time Objective, Not an Aspirational One
A recovery time objective, how long you're willing to be down while restoring, only means something once you've actually measured how long a restore takes on your real data volume, not a guess based on how big the file looks. Teams routinely set an RTO of an hour without ever having timed a full restore, and then discover during a real incident that it takes four.
If your availability target is 99.9 percent, you're defending about 8.76 hours of downtime for the entire year1, and a single restore that runs long can consume a meaningful share of that budget in one incident. Time an actual restore before you commit to a number anyone outside engineering will hold you to.
The Restore Drill: A Runbook You Can Actually Follow
Pick a recent backup and restore it into an isolated environment on a fixed schedule, monthly is reasonable for most teams. Run a specific verification query against the restored data, row counts on your most important tables, a spot check on recent records, not just "did the restore command exit with a success code."
Write down every step as you go the first time, including the parts that felt obvious in the moment, because the person running this drill during a real incident may not be the person who ran it calmly last month. A runbook written under no time pressure is far more reliable than instructions reconstructed from memory during one.
A monthly restore drill can follow these steps:
- Pick a recent backup and restore it into an isolated environment that mirrors production closely enough to be a fair test.
- Time the restore, and compare the result against the recovery time objective you have set.
- Compare row counts on your most critical tables against a known-recent snapshot.
- Spot check that a handful of recent records exist and look correct.
- Record anything that failed, such as a missing credential, network path or tool version, and fix it before the next drill.
What the Drill Usually Finds
The most common discovery isn't that the backup is missing data, it's that the restore process depends on a credential, a network path, or a tool version that quietly changed since the runbook was last tested. A restore script that worked fine six months ago can fail on a database version bump, a rotated credential, or a decommissioned staging environment it was pointed at.
The second most common discovery is time: a restore that takes far longer than anyone assumed, because nobody had measured it against current data volume rather than the volume from when the process was first built. Both findings are only valuable if the drill actually runs; a runbook nobody has executed recently is a document, not a tested procedure.
Making the Drill Someone's Actual Job, Not a Someday Task
Restore drills get skipped for the same reason tech debt gets skipped: nothing forces them onto the calendar, and there's always a feature that feels more urgent this week. Put the drill on a recurring calendar invite with a named owner, and treat a skipped drill the same way you'd treat a skipped deploy of a security patch, not the same way you'd treat a deferred nice to have.
Track the last successful drill date somewhere visible, the same dashboard where you'd track uptime, so it's obvious at a glance whether the team's confidence in its backups is current or stale.
What Good Looks Like
Good backup verification means a recent backup has actually been restored, on a schedule, into an isolated environment, with a specific check confirming the data that came out matches what you expected, not just that the restore command exited without an error.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Frequently Asked Questions
How often should we actually run a restore drill?
Monthly is a reasonable default for most teams, and always after a database version upgrade or a meaningful schema change, since either can quietly break a restore process that worked fine before. A drill that's never repeated only tells you about the system as it existed on the day you ran it.
Do we need a separate environment to restore into?
Yes, an isolated environment that mirrors production closely enough to be a fair test, not your actual production database. Restoring into an environment that's meaningfully different in scale or configuration can hide problems that would only appear at real production volume.
What's the minimum verification after a restore, if we're short on time?
Row counts on your most critical tables compared against a known-recent snapshot, plus a spot check that a handful of recent records actually exist and look correct. That's not exhaustive, but it catches the most common failure, a restore that silently drops or truncates recent data, in a few minutes of work.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
The Backup You Haven't Tested Is Just a Hope
A step-by-step way to actually verify your database backups restore cleanly, instead of trusting a green checkmark from the backup job.
A Runbook for Verifying Database Backups Actually Restore
A step-by-step runbook for proving your database backups restore cleanly, run on a schedule instead of trusted on faith until a real outage.
Proving Your Pipeline Backups Actually Restore
A runbook for actually testing that your event pipeline's backups restore cleanly, instead of trusting a green checkmark on a backup job.
The Backup You've Never Restored Isn't a Backup
A nightly backup job that succeeds every night tells you almost nothing about whether you can actually recover. The drill format that closes that gap.
Proving a Database Backup Can Actually Be Restored
A worked example of running a real database restore drill, and the specific ways backups that report success still fail to restore.
Why Your Backups Might Not Actually Restore
How to verify database backups actually restore, how often to run restore drills, and what to measure besides pass or fail so you trust them.