Engineering Leadership & Technical HiringPlaybook3 min readUpdated September 2026

A Backup You Haven't Restored From Is Just a File

Every team with a database has a backup job, and most of those teams have never actually restored from one. The backup job succeeding tells you a file was written somewhere. It tells you nothing about whether that file can rebuild a working database, how long that rebuild takes, or whether anyone on the team knows the steps well enough to run them at two in the morning.

The only way to know a backup actually works is to restore from it, on a schedule, before an outage forces you to find out for the first time under pressure.

Why a Green Backup Job Doesn't Mean a Working Restore

A backup job can succeed for weeks while writing a file that's silently corrupted, incomplete, or missing a table that got added after the backup script was last updated. Nothing about a green checkmark in your job scheduler checks any of that, because the job's own definition of success is just "the write completed without an error."

The only test that actually validates a backup is using it to rebuild a working database somewhere else, and comparing what comes out against what you expected to see. Anything short of that is trusting the file, not verifying it.

Setting a Real Recovery Time Objective, Not an Aspirational One

A recovery time objective, how long you're willing to be down while restoring, only means something once you've actually measured how long a restore takes on your real data volume, not a guess based on how big the file looks. Teams routinely set an RTO of an hour without ever having timed a full restore, and then discover during a real incident that it takes four.

If your availability target is 99.9 percent, you're defending about 8.76 hours of downtime for the entire year1, and a single restore that runs long can consume a meaningful share of that budget in one incident. Time an actual restore before you commit to a number anyone outside engineering will hold you to.

The Restore Drill: A Runbook You Can Actually Follow

Pick a recent backup and restore it into an isolated environment on a fixed schedule, monthly is reasonable for most teams. Run a specific verification query against the restored data, row counts on your most important tables, a spot check on recent records, not just "did the restore command exit with a success code."

Write down every step as you go the first time, including the parts that felt obvious in the moment, because the person running this drill during a real incident may not be the person who ran it calmly last month. A runbook written under no time pressure is far more reliable than instructions reconstructed from memory during one.

A monthly restore drill can follow these steps:

  1. Pick a recent backup and restore it into an isolated environment that mirrors production closely enough to be a fair test.
  2. Time the restore, and compare the result against the recovery time objective you have set.
  3. Compare row counts on your most critical tables against a known-recent snapshot.
  4. Spot check that a handful of recent records exist and look correct.
  5. Record anything that failed, such as a missing credential, network path or tool version, and fix it before the next drill.

What the Drill Usually Finds

The most common discovery isn't that the backup is missing data, it's that the restore process depends on a credential, a network path, or a tool version that quietly changed since the runbook was last tested. A restore script that worked fine six months ago can fail on a database version bump, a rotated credential, or a decommissioned staging environment it was pointed at.

The second most common discovery is time: a restore that takes far longer than anyone assumed, because nobody had measured it against current data volume rather than the volume from when the process was first built. Both findings are only valuable if the drill actually runs; a runbook nobody has executed recently is a document, not a tested procedure.

Making the Drill Someone's Actual Job, Not a Someday Task

Restore drills get skipped for the same reason tech debt gets skipped: nothing forces them onto the calendar, and there's always a feature that feels more urgent this week. Put the drill on a recurring calendar invite with a named owner, and treat a skipped drill the same way you'd treat a skipped deploy of a security patch, not the same way you'd treat a deferred nice to have.

Track the last successful drill date somewhere visible, the same dashboard where you'd track uptime, so it's obvious at a glance whether the team's confidence in its backups is current or stale.

Executive Capability Standard

What Good Looks Like

Good backup verification means a recent backup has actually been restored, on a schedule, into an isolated environment, with a specific check confirming the data that came out matches what you expected, not just that the restore command exited without an error.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Find out when your backups were last actually restored from, not just when the last backup job succeeded.
2. Do Manually:Manually restore your most recent backup into an isolated environment and time how long the process actually takes.
3. Delegate:Assign an engineer ownership of a recurring restore drill and a runbook written down in enough detail for someone else to follow.
4. Automate:Schedule the restore drill to run automatically and alert the team if a restore fails or exceeds your recovery time objective.
5. Buy:Bring in a database specialist to review your backup and restore process if you've never successfully tested one under realistic conditions.

How to Get Started

Frequently Asked Questions

How often should we actually run a restore drill?

Monthly is a reasonable default for most teams, and always after a database version upgrade or a meaningful schema change, since either can quietly break a restore process that worked fine before. A drill that's never repeated only tells you about the system as it existed on the day you ran it.

Do we need a separate environment to restore into?

Yes, an isolated environment that mirrors production closely enough to be a fair test, not your actual production database. Restoring into an environment that's meaningfully different in scale or configuration can hide problems that would only appear at real production volume.

What's the minimum verification after a restore, if we're short on time?

Row counts on your most critical tables compared against a known-recent snapshot, plus a spot check that a handful of recent records actually exist and look correct. That's not exhaustive, but it catches the most common failure, a restore that silently drops or truncates recent data, in a few minutes of work.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides