Databases and resilienceChecklist3 min readUpdated September 2026

Postgres Backup Checklist: What to Verify Before You Need It

A sound Postgres backup strategy has three parts: regular base backups, continuous archiving of write-ahead logs for point-in-time recovery, and a tested restore procedure. A backup you've never restored is only a hope.

Use the checklist below on a managed or self-hosted database. Each item is something you can verify in an afternoon.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What kinds of Postgres backup exist, and which do you need?

There are two families. Logical backups, made with pg_dump, export schema and data as SQL or an archive file. They're portable across versions and good for moving data, but slow to restore on large databases and only capture a moment in time. Physical backups copy the data files, and combined with archived write-ahead logs (WAL) they allow point-in-time recovery to a chosen second.

Most production systems want both: physical backups with WAL archiving for fast recovery and precise rollback, plus periodic logical dumps as a portable safety copy. Managed services such as AWS RDS and Supabase provide automated backups and, on some plans, point-in-time recovery. Confirm exactly which your plan includes and how long it retains them, rather than assuming.

How to decide your recovery targets

Two numbers drive the design. Recovery point objective (RPO) is how much recent data you can afford to lose. Recovery time objective (RTO) is how long you can be down. Nightly dumps mean an RPO of up to a day; WAL archiving can bring it down to minutes.

Put the numbers in context with your availability goals. A 99.9% availability target allows about 8.76 hours of downtime per year1. A single slow restore of a large database could consume much of that, so measure your real restore time, not the vendor's brochure figure. A disaster recovery plan sets these targets per service, and each service gets its own pair of numbers.

The checklist

Go through these and mark each yes or no:

  • Automated backups run on a schedule, and you get an alert when one fails, not just when it succeeds.
  • WAL archiving or the provider's point-in-time recovery is enabled, and you know the earliest restorable time.
  • Backups are encrypted at rest, and you know who can access the keys.
  • At least one copy lives in a separate account or project from production, so a compromised credential can't delete both.
  • Retention is written down and matches your policy and any customer or legal commitments.
  • Roles, extensions and other cluster-wide objects are captured, since pg_dump of one database doesn't include them.
  • Someone other than the person who set it up can find and follow the restore steps.

A no on the off-account copy is the most common gap and the most damaging. Ransomware and mistaken deletes both target whatever the production credentials can reach.

How to run a restore drill

Schedule a restore test at least quarterly. It takes a couple of hours the first time and gets faster.

  1. Choose a recent backup, and a target time if you're testing point-in-time recovery.
  2. Restore into a new, isolated instance, never over production.
  3. Time the restore, and time how long the database takes to become consistent and usable.
  4. Check contents: row counts on key tables, the newest record's timestamp, and a few application queries.
  5. Point a staging copy of the app at it and run a smoke test.
  6. Record the result, the duration and anything that surprised you, then update the runbook.

Also test the unpleasant case: someone deleted a table at a known time. Practice recovering to just before it and extracting the data without overwriting current rows.

Mistakes that make backups useless

The patterns behind most failed recoveries:

  • Backups stored on the same disk, volume or account as the database.
  • Never restoring, so a corrupt or incomplete backup goes unnoticed for months.
  • Logical dumps that skip large objects, roles or extensions that the app needs.
  • Assuming a managed service's backup covers you against accidental deletion. It restores the whole database, which may not be what you want.
  • Ignoring backup size growth until storage or restore time hits a limit.

Related reading: the tradeoffs between hosts in MongoDB Atlas vs AWS RDS vs Supabase, and how to size costs in the managed Postgres cost comparison.

Executive Capability Standard

What Good Looks Like

Backups run automatically with failure alerts, a copy exists outside the production account, and a timed restore has been tested within the last quarter.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Find out what your database host backs up, how often, and how long it keeps it, in writing.
2. Do Manually:Restore the latest backup into a scratch instance by hand and record how long it takes.
3. Delegate:Assign a named owner for backup health with a quarterly restore drill on their calendar.
4. Automate:Script the restore into a temporary instance with automatic checks on row counts and freshness.
5. Buy:Choose a managed database tier that includes point-in-time recovery if your team can't operate WAL archiving.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

AWS RDS

Fits when you want automated backups and snapshot management handled by the platform, with retention you configure.

Visit AWS RDS→
Supabase

Fits when you want backups bundled with a managed Postgres platform, after confirming what your plan covers.

Visit Supabase→

Frequently Asked Questions

How often should you back up a production Postgres database?

Take a base backup at least daily and archive write-ahead logs continuously if data loss must be measured in minutes. Your recovery point objective sets the interval, not a default schedule.

What is point-in-time recovery in Postgres?

It restores a base backup and then replays archived write-ahead logs up to a chosen moment, such as just before a bad deploy. It needs continuous WAL archiving or a provider feature that does it for you.

Is a managed database's automatic backup enough?

It covers hardware and many failure cases, but check retention, whether point-in-time recovery is included, and whether copies exist outside the account. Also test restores, since accidental deletes need selective recovery.

How often should you test restoring a backup?

At least quarterly, and after major changes to the database or backup configuration. Restore to an isolated instance, time it, and verify the data before updating your runbook.

Sources

Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.

  1. Allowed downtime per year by availability target. Google SRE Book, Table 1-1 Availability table, 2016.

Related Guides