For those running PostgreSQL on Kubernetes with an operator such as CloudNativePG, Zalando, or Crunchy, how do you confirm that your backups are genuinely restorable? A successful backup job alone does not prove that recovery will work.
I manage several CloudNativePG clusters and have started exploring restore testing. In a small k3s environment using CloudNativePG 1.30, the barman-cloud plugin, and S3, I deploy a temporary cluster from the newest backup, run a few SQL checks, measure how long it takes to become ready, and then remove it. For a few hundred megabytes, the process usually takes about a minute when everything goes normally.
I'm interested in how others approach this:
- Do you test restores regularly, and is the process manual, scheduled with a CronJob, part of CI, or handled by another tool?
- What do you validate after recovery: cluster readiness, row counts, recent timestamps, specific tables, application queries, or something else?
- How often do you run these tests, and do you also verify point-in-time recovery and WAL archives?
I'm also interested in hearing from teams that do not test restores, since I'd like to understand how common that is and what a practical baseline looks like.
4 Answers
The only dependable way to verify a backup is to restore it. At minimum, confirm that the recovered cluster becomes usable, expected tables exist, and representative queries return the expected data. For stronger coverage, include checks for recent records, row counts, schema migrations, extensions, and a point-in-time recovery target rather than relying only on backup-job success.
We create a temporary CloudNativePG cluster from the S3 backup location every week using a scheduled job. Once it is ready, we run a few validation queries and then delete it. One useful check is a table that receives frequent inserts: we verify that its newest timestamp is close to the expected backup or recovery point. That gives us more confidence than checking only whether the pods became ready.
We currently perform a full disaster-recovery restore about twice a year. That confirms the overall procedure, but the discussion has convinced us to add smaller automated restore checks more frequently. For small databases, restoring into a temporary Docker Compose PostgreSQL instance and comparing key tables and data against production is a reasonable lightweight option.
Our backup job launches a restore test as soon as the backup completes. It restores into a temporary PostgreSQL instance, runs basic checks, and tears the instance down afterward. The important part is making the restore test automatic, because a manual process tends to get skipped.

That matches what I’m testing now. The temporary-cluster approach is quick for our database size, so the main challenge is deciding which checks provide useful coverage without trying to compare every row.