Distributed Systems & Enterprise ResiliencePlaybook3 min readUpdated September 2026

Why Your Backups Might Not Actually Restore

A backup job that reports success every night is not the same thing as a backup that restores cleanly. The backup process can complete without error while the file is subtly corrupted, incomplete, or simply untested against the exact restore procedure you'd need during a real incident.

The only way to know a backup actually works is to restore it, on purpose, before you need to. This is a practical way to build that habit without it becoming a burden nobody keeps up with.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

A Green Backup Job Proves Less Than You Think

Most backup tooling reports success based on the job completing, not based on the resulting file being usable. A database that locked mid-backup, a snapshot taken during an inconsistent write, or a storage target silently truncating a large file can all still produce a job that reports success while leaving you with a backup that won't actually restore.

Treat "backup completed" and "backup is restorable" as two separate claims that each need their own verification, because they fail independently of each other.

How do you build a restore drill you'll actually repeat?

Automate restoring the latest backup into an isolated environment on a fixed schedule, weekly for a fast-moving product, monthly at minimum for a stable one, rather than relying on someone remembering to test it manually. A drill that depends on a person's memory eventually stops happening the first time that person is busy.

Measure two things every time: did the restore complete without error, and does the resulting data pass a basic integrity check, like row counts matching or a known record being present and correct. A restore that completes but returns corrupted or incomplete data is arguably worse than one that fails outright, because it looks successful.

A restore drill that holds up looks like this:

  1. Restore the latest backup into an isolated environment automatically, so the drill never touches production and never depends on someone remembering to run it.
  2. Put the drill on a fixed schedule, weekly for a fast-moving product and at least monthly for a stable one.
  3. Record whether the restore completed, and how long it took from start to a usable database.
  4. Run a basic integrity check on the restored data, since a restore that finishes with incomplete data can look fine until someone relies on it.

How long should a database restore take?

How long a restore actually takes matters as much as whether it works, because that time is what you're spending against your availability commitment during a real incident. A team promising 99.9% uptime has a downtime budget of about 0.365 days a year, roughly nine hours, and if your restore drill shows the process alone takes several hours, that's a meaningful share of your entire annual budget spent on data recovery before the rest of the incident response even starts1.

Track this restore time drill over drill. A restore that gets slower as your data grows is a warning sign worth acting on well before an actual incident forces the question.

Test the Failure Case You're Actually Worried About

A drill that always restores the most recent, healthy backup doesn't tell you what happens if you need to restore from three backups ago, because the most recent one was itself corrupted or taken after a bad deploy. Periodically restore an older backup specifically to confirm your retention window actually contains a usable point to recover from.

This matters more than it sounds like it should: the scenario where you need an old backup is usually the scenario where something has already gone wrong for a while, which is exactly the situation your drills should be built to survive.

Check the restore against a point in time before whatever caused the problem, not just against the most recent snapshot before the incident. A backup taken an hour before a bad migration ran is still a bad backup, even though it's technically recent.

Write Down Who Runs This During a Real Incident

A restore drill that only one specific engineer knows how to run is a single point of failure hiding inside your disaster recovery plan. Document the exact steps, including any credentials or access needed, somewhere the whole on-call rotation can find and follow it without that one person being awake and available.

Run the drill occasionally with a different engineer executing it, following only the written documentation, to confirm the documentation is actually complete rather than relying on tribal knowledge that lives in one person's head.

That handoff test tends to surface gaps a fluent author never notices: a credential that lives in the original engineer's personal password manager, a step that assumes context nobody wrote down. Finding those gaps during a calm, scheduled drill is far better than finding them during a real incident.

Executive Capability Standard

What Good Looks Like

Good backup practice means restores are actually tested on a fixed schedule, timed, and checked for data integrity, not just backup jobs completing without error, with documentation complete enough that someone other than the person who wrote it could run a restore during a real incident.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Read your backup tooling's actual restore documentation end to end, since most teams have never read past the section on configuring the backup job itself.
2. Do Manually:Manually restore your most recent backup into an isolated environment once, timing it and checking the resulting data, to establish a baseline before automating the drill.
3. Delegate:Give a specific engineer ownership of the restore drill schedule and its documentation, with a requirement that someone else run it successfully at least once a year using only the written steps.
4. Automate:Script the restore drill to run on a fixed schedule automatically, including the integrity check, so it happens reliably without depending on anyone remembering to trigger it.
5. Buy:Bring in a managed backup or disaster recovery service once your data volume or restore time is large enough that testing and maintaining the process in-house is consistently slipping.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Vanta

A compliance automation tool like Vanta can help evidence that backup and restore testing is actually happening on schedule, which is useful documentation separate from running the drills themselves.

Visit Vanta→

Frequently Asked Questions

How often should we test that our database backups actually restore?

On a fixed, automated schedule rather than whenever someone remembers, weekly for a fast-moving product and at least monthly for a stable one. A manual process that depends on someone's memory tends to lapse exactly when things get busy, which is also when a real incident is more likely.

Is a backup job reporting success enough to know it will actually restore?

No. A job can report success while producing a backup that's corrupted, incomplete, or otherwise unusable, since most backup tooling checks whether the job completed, not whether the resulting file restores cleanly. The only real proof is an actual restore, tested on a schedule.

Should we test restoring old backups, not just the most recent one?

Yes, periodically. A drill that only restores the most recent backup doesn't confirm your retention window actually contains a usable recovery point from further back, and that's often exactly the scenario you'd need during a real incident, where the most recent backups are the ones already affected.

What should a restore drill actually measure besides success or failure?

How long the restore takes, and whether the resulting data passes a basic integrity check, not just whether the process completes without an error. A restore that finishes but returns corrupted or incomplete data can be worse than an outright failure, because it looks fine until someone relies on it.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides