Developer Productivity & Platform EngineeringPlaybook3 min readUpdated September 2026

The Backup You Haven't Tested Is Just a Hope

A backup job that completes successfully tells you a file was written somewhere. It tells you nothing about whether that file can rebuild a working database. The only way to know is to actually restore it, on a schedule, and treat that restore as the real test.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Step 1: Restore to an Isolated Environment, Not Production

Spin up a separate environment for the restore, never test against production and never test in a way that could touch real customer traffic. This sounds obvious, but under time pressure, teams cut this corner and it's exactly the corner that turns a routine drill into a real incident.

An isolated restore target also means you can run the drill on a schedule without coordinating around production maintenance windows, which is a big part of why these drills stop happening in the first place.

Step 2: Measure Actual Restore Time, Not Backup Completion Time

The number that matters during a real incident is how long it takes to get back to a working state, not how long the backup job took to run. Time the full restore from scratch: provisioning the target, pulling the backup, running it, and verifying the schema and row counts look right.

That's your real recovery time, and it should be measured against your allowed downtime budget1. Say your restore takes six hours end to end: against a target that allows only a few hours of downtime a year, that's a recovery plan that fails on paper before it's ever tested for real.

For example, the backup job reports success in about forty minutes every night, so the team assumes recovery is quick. A drill shows that provisioning a new database takes an hour, pulling the file takes another, and verification takes a third. Real recovery time is several hours, not forty minutes. Record that end-to-end number and compare it with what your uptime commitment allows. If it doesn't fit, the fix might be a warm standby or a faster restore path, and you found that out on a quiet afternoon rather than during an outage.

Step 3: Verify Data Integrity, Not Just That the Restore Ran

A restore that completes without an error can still be missing recent transactions, have corrupted rows, or be silently out of sync with what your application expects. Run a set of specific checks after every restore: row counts against known baselines, a handful of specific records you can verify by hand, and any foreign key constraints that would flag corruption.

Write these checks down once and reuse them every drill, so verification doesn't depend on someone remembering what "looks right" means for your particular schema, and so a different engineer running the drill next quarter gets the same thorough, repeatable check every single time, no matter who happens to be running the drill that quarter.

Step 4: Rotate Who Runs the Drill

If only one engineer has ever run a restore, you don't have a tested recovery process, you have one person's private knowledge that happens to work. Rotate the drill across the team so more than one person has actually done it under realistic conditions, with the runbook as the only guide, not that person's memory filling in the gaps.

This also surfaces gaps in the written runbook fast: if a different engineer gets stuck on step three, that's a real documentation problem worth fixing before an actual incident exposes it at 2am.

Step 5: Schedule It Like You Mean It

A restore drill that's "on the roadmap" without a specific date on the calendar doesn't happen. Put it on a recurring quarterly schedule, treat a skipped drill as a real gap worth flagging, not a minor miss, and track the last successful restore date somewhere visible so nobody's guessing how stale your recovery confidence actually is.

What to Do When a Drill Actually Fails

A failed drill is a good outcome in the sense that it found the gap before a real incident did. Treat it exactly like a production incident, with a written root cause and a specific fix, not a quiet retry until it happens to work.

Teams sometimes treat a failed drill as an embarrassment to fix quietly and move on from. That instinct is backwards: a documented, understood failure from a drill is far cheaper than the same failure discovered during a real outage, and it deserves the same rigor a real incident review would get, complete with an owner and a follow-up deadline.

After each drill, confirm the following before closing it out:

  • The restore ran in an isolated environment that could not touch production or real customer traffic.
  • Total restore time, from provisioning through verification, is recorded and compared with your downtime budget.
  • Row counts, hand-verified records, and foreign key checks passed against your written baseline.
  • Someone other than the usual engineer ran the drill using only the runbook.
  • The next drill has a date on the calendar, and any failure has an owner and a deadline.
Executive Capability Standard

What Good Looks Like

A real backup strategy is proven by a timed, verified restore run on a fixed schedule by more than one person, not by a green checkmark on the backup job itself.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Find out when your team last actually restored a backup, not just when the last backup job completed successfully.
2. Do Manually:Manually run a full restore to an isolated environment and time it against your actual downtime budget.
3. Delegate:Assign the quarterly restore drill to a rotating owner so more than one person has real experience running it.
4. Automate:Automate the restore verification checks (row counts, key records, constraints) so a human doesn't have to eyeball correctness each time.
5. Buy:Bring in a database reliability specialist if a restore has never been successfully tested and the data involved is genuinely critical.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How often should we actually run a restore drill?

Quarterly is a reasonable baseline for most small teams, more often for anything storing data you genuinely couldn't afford to lose or be down for long. The exact cadence matters less than actually keeping to whatever schedule you set.

Is testing backups on a staging environment with real data safe?

Only if that staging environment already meets the same access controls and data handling rules as production. If it doesn't, restore to a fully isolated, locked-down environment instead, or scrub the data as part of the restore process before anyone else can access it.

What's the biggest mistake teams make with backup verification?

Trusting the backup job's own success status as proof the backup works. A completed job confirms a file was written, not that the file can rebuild a working, correct database. Only an actual restore confirms that.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides