Every team I have worked with says the same thing when I ask about backups. "Yeah, we have those. The job runs every night."
Then I ask when someone last restored from one, and the room gets quiet. Not because anyone is careless. A backup job that runs is easy to see. A restore that works is something you only learn by doing it.
So before any risky deploy, I run a small check. It takes about fifteen minutes, and it has one rule: a backup does not count until I have restored it somewhere harmless.
A backup that exists is not a backup that restores
A file sitting in a bucket proves only that a file is sitting in a bucket. It does not prove the dump is complete, that it is readable, that you have the credentials to fetch it, or that you remember the steps at 2am while the site is down.
There is also the time question. If a restore takes four hours and your plan assumed twenty minutes, you do not have a plan. You have a hope. The only way to learn the real number is to run it once and look at the clock.
The ways backups fail quietly
None of these make noise. That is the whole problem.
- The job fails and nobody reads the failure email. The last good backup is three weeks old and the dashboard still says "enabled".
- Backups live on the same server as the database. When the server goes, both go together.
- Backups live in the same cloud account. One compromised or suspended account takes the data and the copies with it.
- Only files are backed up. The uploads folder is safe, the database is not in the job at all.
- Credentials expired. The upload key was rotated months ago and the job has been writing nowhere since.
- Retention is too short. You notice a bad migration on Monday, and the only backups left are from Sunday night, after the damage.
- Nobody knows the restore steps. The one person who set it up has moved on, and the notes are in their head.
The fifteen minute check
Here is what I actually do before a risky deploy. It is boring on purpose.
- Check the last successful backup time. Not "is the job enabled", but the timestamp of the newest file or snapshot. If it is older than you expected, stop here and fix that first.
- Restore the newest backup into a throwaway database. Never into production, never over anything you care about.
- Run a couple of sanity queries. Row counts on your biggest tables, and the newest record by created date. If the newest row is from last week, you just found a problem.
- Write down how long it took. That number goes in your runbook next to the steps.
- Delete the throwaway database.
PostgreSQL version
This assumes a custom format dump made with pg_dump -Fc. The official docs for pg_restore and SQL dumps cover every flag, and I would read both once.
# fetch the newest dump first, then: createdb restore_check time pg_restore --no-owner --dbname=restore_check latest.dump psql restore_check -c "SELECT count(*) FROM orders;" psql restore_check -c "SELECT max(created_at) FROM orders;" dropdb restore_check
The time in front is the point. You get your real restore duration for free.
MongoDB version
Same idea with mongodump and mongorestore. The mongorestore docs explain the namespace options, which is how you restore into a different database name so you cannot touch the real one.
time mongorestore --uri="mongodb://localhost:27017" \
--archive=latest.archive --gzip \
--nsFrom="app.*" --nsTo="restore_check.*"
mongosh restore_check --eval "db.orders.countDocuments()"
mongosh restore_check --eval "db.orders.find().sort({createdAt:-1}).limit(1)"
mongosh restore_check --eval "db.dropDatabase()"Swap app and orders for your own names. Run it on a scratch machine or a local box, not on the production host.
If you use a managed database
Managed services take snapshots for you, which is great, but the same rule applies. Two things to check by hand. First, what is the actual point-in-time restore window on your plan? Open the provider console and read the number, do not trust your memory. Second, restore a snapshot into a new instance at least once. Many teams never find out how slow that is, or that a setting or a security group did not come along with it.
For self-managed Postgres with continuous archiving, the official continuous archiving docs are the place to start. It is more work, and it is also how you get point-in-time recovery.
Where this fits with the rest of production safety
A tested restore is one layer. I wrote about the bigger picture in Production safety is not a setting, and this is the part of that story that people skip because it feels like paperwork. If you want a second pair of eyes on your backup and deploy setup, that is part of what we do under DevOps and cloud cost.
What each step catches
| Step | Failure it exposes |
|---|---|
| Check the newest backup timestamp | Job failing quietly, expired credentials |
| Restore into a throwaway database | Corrupt or partial dumps, missing database, nobody knowing the steps |
| Sanity queries | Backups that are old or half empty |
| Note the time it took | A recovery plan that assumes minutes when reality is hours |
Notice that none of this needs special tooling. A terminal, a scratch database and a clock are enough.
What a bad result looks like
Sometimes the check fails, and that is the good outcome. You found out on a calm afternoon instead of during an outage. A restore that errors out, a newest row that is days old, or a dump that is suspiciously small are all things you can fix today.
If something fails, do not deploy yet. Fix the backup first, run the check again, and only then ship. The deploy will still be there in an hour, and a working restore is worth more than a day saved.
Write it down, then repeat it
The last step is the one that makes the rest useful. Put the restore steps where the on-call person can find them without asking you. Include where the backups live, who has access, the exact commands, and the time it took on your last run.
Then put a reminder in the calendar to run the check every quarter. Backups drift. People change, keys rotate, databases grow, and a restore that took ten minutes last year can take two hours now.
Use whatever tools you like. The habit matters more than the tool. If you take one thing away, make it this: restore your latest backup somewhere harmless this week, and write down how long it took.

