Why Your Backups Might Not Actually Restore
A backup job that reports success every night is not the same thing as a backup that restores cleanly. The backup process can complete without error while the file is subtly corrupted, incomplete, or simply untested against the exact restore procedure you'd need during a real incident.
The only way to know a backup actually works is to restore it, on purpose, before you need to. This is a practical way to build that habit without it becoming a burden nobody keeps up with.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
A Green Backup Job Proves Less Than You Think
Most backup tooling reports success based on the job completing, not based on the resulting file being usable. A database that locked mid-backup, a snapshot taken during an inconsistent write, or a storage target silently truncating a large file can all still produce a job that reports success while leaving you with a backup that won't actually restore.
Treat "backup completed" and "backup is restorable" as two separate claims that each need their own verification, because they fail independently of each other.
How do you build a restore drill you'll actually repeat?
Automate restoring the latest backup into an isolated environment on a fixed schedule, weekly for a fast-moving product, monthly at minimum for a stable one, rather than relying on someone remembering to test it manually. A drill that depends on a person's memory eventually stops happening the first time that person is busy.
Measure two things every time: did the restore complete without error, and does the resulting data pass a basic integrity check, like row counts matching or a known record being present and correct. A restore that completes but returns corrupted or incomplete data is arguably worse than one that fails outright, because it looks successful.
A restore drill that holds up looks like this:
- Restore the latest backup into an isolated environment automatically, so the drill never touches production and never depends on someone remembering to run it.
- Put the drill on a fixed schedule, weekly for a fast-moving product and at least monthly for a stable one.
- Record whether the restore completed, and how long it took from start to a usable database.
- Run a basic integrity check on the restored data, since a restore that finishes with incomplete data can look fine until someone relies on it.
How long should a database restore take?
How long a restore actually takes matters as much as whether it works, because that time is what you're spending against your availability commitment during a real incident. A team promising 99.9% uptime has a downtime budget of about 0.365 days a year, roughly nine hours, and if your restore drill shows the process alone takes several hours, that's a meaningful share of your entire annual budget spent on data recovery before the rest of the incident response even starts1.
Track this restore time drill over drill. A restore that gets slower as your data grows is a warning sign worth acting on well before an actual incident forces the question.
Test the Failure Case You're Actually Worried About
A drill that always restores the most recent, healthy backup doesn't tell you what happens if you need to restore from three backups ago, because the most recent one was itself corrupted or taken after a bad deploy. Periodically restore an older backup specifically to confirm your retention window actually contains a usable point to recover from.
This matters more than it sounds like it should: the scenario where you need an old backup is usually the scenario where something has already gone wrong for a while, which is exactly the situation your drills should be built to survive.
Check the restore against a point in time before whatever caused the problem, not just against the most recent snapshot before the incident. A backup taken an hour before a bad migration ran is still a bad backup, even though it's technically recent.
Write Down Who Runs This During a Real Incident
A restore drill that only one specific engineer knows how to run is a single point of failure hiding inside your disaster recovery plan. Document the exact steps, including any credentials or access needed, somewhere the whole on-call rotation can find and follow it without that one person being awake and available.
Run the drill occasionally with a different engineer executing it, following only the written documentation, to confirm the documentation is actually complete rather than relying on tribal knowledge that lives in one person's head.
That handoff test tends to surface gaps a fluent author never notices: a credential that lives in the original engineer's personal password manager, a step that assumes context nobody wrote down. Finding those gaps during a calm, scheduled drill is far better than finding them during a real incident.
What Good Looks Like
Good backup practice means restores are actually tested on a fixed schedule, timed, and checked for data integrity, not just backup jobs completing without error, with documentation complete enough that someone other than the person who wrote it could run a restore during a real incident.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How often should we test that our database backups actually restore?
On a fixed, automated schedule rather than whenever someone remembers, weekly for a fast-moving product and at least monthly for a stable one. A manual process that depends on someone's memory tends to lapse exactly when things get busy, which is also when a real incident is more likely.
Is a backup job reporting success enough to know it will actually restore?
No. A job can report success while producing a backup that's corrupted, incomplete, or otherwise unusable, since most backup tooling checks whether the job completed, not whether the resulting file restores cleanly. The only real proof is an actual restore, tested on a schedule.
Should we test restoring old backups, not just the most recent one?
Yes, periodically. A drill that only restores the most recent backup doesn't confirm your retention window actually contains a usable recovery point from further back, and that's often exactly the scenario you'd need during a real incident, where the most recent backups are the ones already affected.
What should a restore drill actually measure besides success or failure?
How long the restore takes, and whether the resulting data passes a basic integrity check, not just whether the process completes without an error. A restore that finishes but returns corrupted or incomplete data can be worse than an outright failure, because it looks fine until someone relies on it.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.
Related Guides
The Backup You Haven't Tested Is Just a Hope
A step-by-step way to actually verify your database backups restore cleanly, instead of trusting a green checkmark from the backup job.
A Runbook for Verifying Database Backups Actually Restore
A step-by-step runbook for proving your database backups restore cleanly, run on a schedule instead of trusted on faith until a real outage.
The Backup You've Never Restored Isn't a Backup
A nightly backup job that succeeds every night tells you almost nothing about whether you can actually recover. The drill format that closes that gap.
A Backup You Haven't Restored From Is Just a File
A backup job that succeeds every night tells you nothing about whether a restore will actually work. Here is a runbook for testing the part that matters.
Proving Your Pipeline Backups Actually Restore
A runbook for actually testing that your event pipeline's backups restore cleanly, instead of trusting a green checkmark on a backup job.
A Runbook for Proving Your Backups Actually Restore
A step-by-step way to verify database backups actually restore, on a schedule, instead of discovering a gap the first time you need a backup for real.