Level 22 — Backups
Everything else on the server can be rebuilt from Git in an hour. Your data cannot. This is the chapter where "production-ready" is decided.
THE ONLY BACKUP THAT COUNTS IS ONE YOU HAVE RESTORED
The universal failure story: backups ran "successfully" every night for eight months. On the day it mattered, the .dump files were 0 bytes — the password had changed and pg_dump was failing while cron dutifully wrote an empty file and exited 0.
Or: the backup existed, but it only ever contained the default postgres database. Or the disk it wrote to filled up in March. Or it was on the same server that just died.
Test a restore. Write down the date. Do it again next month. Everything else in this chapter is secondary to that one habit.
What needs backing up
| Asset | Recoverable without a backup? | Priority |
|---|---|---|
| PostgreSQL data | ❌ Never | 🔴 Critical |
| User uploads | ❌ Never | 🔴 Critical |
.env secrets | ❌ Not in Git (Level 8) | 🔴 Critical |
| Application code | ✅ From Git | 🟢 Low |
| Nginx configuration | ⚠️ Rewritable, but slowly | 🟡 Medium |
| TLS certificates | ✅ Certbot can reissue | 🟢 Low |
| PM2 ecosystem config | ✅ In Git | 🟢 Low |
| Redis data | ⚠️ Depends — see below | 🟡 Medium |
| Cron jobs | ⚠️ Rewritable | 🟡 Medium |
authorized_keys | ✅ You have the keys | 🟢 Low |
THE .env FILE IS THE MOST COMMONLY FORGOTTEN CRITICAL ASSET
It is deliberately excluded from Git, so it is absent from your code backup. Restore the database and re-clone the repo and you still cannot start the app — you no longer know the JWT secret (invalidating every session), the Stripe webhook secret, or the SMTP credentials.
Store production .env contents in a password manager as a secure note and include the file in your server backup. Verify it during the restore drill.
The 3-2-1 rule
3 copies of your data, on 2 different media/systems, with 1 copy off-site.
A BACKUP ON THE SAME SERVER IS NOT A BACKUP
It protects against exactly one scenario: DROP TABLE by accident. It does not protect against:
- The VPS being deleted (yours or the provider's mistake)
- Disk failure
- Ransomware, which encrypts local backups too — that is the first thing modern ransomware looks for
- Your account being suspended or compromised
- The datacenter having a bad day
The off-site copy is the one that matters. Everything else is convenience.
THE MODERN ADDITION: 3-2-1-1-0
- 1 copy immutable or offline (object-lock in S3, so ransomware cannot delete it)
- 0 errors on restore verification
Immutability is genuinely worth configuring — object storage with a retention lock means even a fully compromised server cannot destroy your backups. That is the difference between an incident and a company-ending event.
PostgreSQL backups
Logical backup — pg_dump
# SERVER — custom format (compressed, selective restore, parallel)
pg_dump -U myapp -h 127.0.0.1 -Fc -f /var/backups/postgres/myapp-$(date -u +%Y%m%d-%H%M%S).dump myapp_production
# Plain SQL, gzipped — human-readable
pg_dump -U myapp -h 127.0.0.1 myapp_production | gzip > backup.sql.gz
# Directory format — supports parallel dump AND parallel restore
pg_dump -U myapp -h 127.0.0.1 -Fd -j 4 -f /var/backups/postgres/myapp-$(date -u +%F) myapp_production
# Everything, including roles and permissions (run as postgres)
sudo -u postgres pg_dumpall -f /var/backups/postgres/cluster-$(date -u +%F).sql
# Docker
docker compose exec -T postgres pg_dump -U myapp -Fc myapp_production > backup.dump| Format | Flag | Restore with | Notes |
|---|---|---|---|
| Custom | -Fc | pg_restore | Recommended — compressed, selective, reorderable |
| Plain | (default) | psql | Readable, greppable, larger |
| Directory | -Fd | pg_restore -j N | Parallel dump and restore — fastest for large DBs |
| Tar | -Ft | pg_restore | Rarely useful |
pg_dump DUMPS ONE DATABASE; ROLES LIVE AT CLUSTER LEVEL
pg_dump myapp_production does not include the myapp role or its password. Restoring onto a fresh server gives you tables owned by a role that does not exist, and pg_restore throws a wall of errors.
Capture roles separately:
sudo -u postgres pg_dumpall --globals-only -f /var/backups/postgres/globals-$(date -u +%F).sqlSmall file, restored first. Skip it and your restore drill fails in a way you will not expect.
pg_dump IS CONSISTENT WITHOUT BLOCKING
It runs in a repeatable-read transaction, so the dump is a consistent snapshot from one point in time, and it does not block writes. You can run it against a live production database.
It does hold a lock that conflicts with ALTER TABLE, so avoid running migrations during a dump. And on a very large database it can hold a transaction open long enough to delay vacuum — worth knowing above ~100 GB.
Physical backup — pg_basebackup
# SERVER
sudo -u postgres pg_basebackup -D /var/backups/postgres/base-$(date -u +%F) -Ft -z -P -X streampg_dump (logical) | pg_basebackup (physical) | |
|---|---|---|
| Backs up | SQL statements to recreate data | The raw data directory |
| Size | Smaller (compressed logical) | Larger |
| Restore speed | Slower (replays SQL, rebuilds indexes) | Faster (copy files) |
| Cross-version | ✅ Yes | ❌ Same major version only |
| Cross-architecture | ✅ Yes | ❌ No |
| Selective restore | ✅ One table | ❌ All or nothing |
| Point-in-time recovery | ❌ No | ✅ With WAL archiving |
FOR A SINGLE VPS, pg_dump IS THE RIGHT DEFAULT
It is portable across PostgreSQL versions and architectures — which matters enormously during disaster recovery, when you may be restoring onto whatever server you can get quickly. Physical backups plus WAL archiving are for large databases where a dump takes too long, or where you need point-in-time recovery to the second.
The backup script
#!/usr/bin/env bash
# /home/deploy/scripts/backup.sh
set -euo pipefail
# ---- Configuration ----
BACKUP_ROOT="/var/backups/myapp"
DB_NAME="myapp_production"
DB_USER="myapp"
DB_HOST="127.0.0.1"
APP_DIR="/home/deploy/apps/myapp"
RETAIN_DAYS=14
HC_PING_URL="${HC_PING_URL:-}" # Healthchecks.io dead-man's switch
S3_BUCKET="s3://myapp-backups"
TS=$(date -u +%Y%m%d-%H%M%S)
DEST="$BACKUP_ROOT/$TS"
mkdir -p "$DEST"
# .env holds PGPASSWORD, S3 credentials, HC_PING_URL
set -a; . "$APP_DIR/.env"; set +a
export PGPASSWORD="${DATABASE_PASSWORD:?DATABASE_PASSWORD not set}"
fail() {
echo "❌ BACKUP FAILED: $1" >&2
[ -n "$HC_PING_URL" ] && curl -fsS -m 10 --retry 3 "${HC_PING_URL}/fail" > /dev/null || true
exit 1
}
trap 'fail "unexpected error on line $LINENO"' ERR
echo "==> Backup started $TS"
# ---- 1. PostgreSQL roles/globals ----
sudo -u postgres pg_dumpall --globals-only > "$DEST/globals.sql" || fail "pg_dumpall globals"
# ---- 2. Database ----
pg_dump -U "$DB_USER" -h "$DB_HOST" -Fc -f "$DEST/database.dump" "$DB_NAME" || fail "pg_dump"
# ---- 3. VERIFY the dump is real ----
SIZE=$(stat -c%s "$DEST/database.dump")
[ "$SIZE" -lt 10000 ] && fail "dump suspiciously small: ${SIZE} bytes"
pg_restore --list "$DEST/database.dump" > /dev/null 2>&1 || fail "dump is not readable by pg_restore"
TABLES=$(pg_restore --list "$DEST/database.dump" | grep -c "TABLE DATA" || true)
[ "$TABLES" -lt 1 ] && fail "dump contains no table data"
echo " database.dump: $((SIZE/1024/1024))MB, ${TABLES} tables"
# ---- 4. Application secrets and uploads ----
cp "$APP_DIR/.env" "$DEST/env.backup"
chmod 600 "$DEST/env.backup"
if [ -d "$APP_DIR/shared/uploads" ]; then
tar -czf "$DEST/uploads.tar.gz" -C "$APP_DIR/shared" uploads || fail "uploads tar"
fi
# ---- 5. System configuration ----
sudo tar -czf "$DEST/system-config.tar.gz" \
/etc/nginx/sites-available \
/etc/nginx/nginx.conf \
/etc/letsencrypt \
/etc/postgresql \
/etc/redis/redis.conf \
/etc/fail2ban/jail.local \
2>/dev/null || true
crontab -l > "$DEST/crontab.txt" 2>/dev/null || true
pm2 save && cp ~/.pm2/dump.pm2 "$DEST/pm2-dump.json" 2>/dev/null || true
dpkg --get-selections > "$DEST/packages.txt"
# ---- 6. Manifest ----
cat > "$DEST/MANIFEST.txt" <<EOF
Backup: $TS
Host: $(hostname)
Database: $DB_NAME
Commit: $(git -C "$APP_DIR" rev-parse --short HEAD 2>/dev/null || echo unknown)
PG version:$(psql -U "$DB_USER" -h "$DB_HOST" -d "$DB_NAME" -tAc 'SHOW server_version')
Tables: $TABLES
Size: $(du -sh "$DEST" | cut -f1)
EOF
# ---- 7. Checksums ----
( cd "$DEST" && sha256sum ./* > SHA256SUMS 2>/dev/null || true )
# ---- 8. Encrypt before leaving the machine ----
if command -v age > /dev/null && [ -n "${BACKUP_AGE_PUBKEY:-}" ]; then
tar -czf - -C "$BACKUP_ROOT" "$TS" | age -r "$BACKUP_AGE_PUBKEY" > "$BACKUP_ROOT/$TS.tar.gz.age"
UPLOAD="$BACKUP_ROOT/$TS.tar.gz.age"
else
tar -czf "$BACKUP_ROOT/$TS.tar.gz" -C "$BACKUP_ROOT" "$TS"
UPLOAD="$BACKUP_ROOT/$TS.tar.gz"
fi
# ---- 9. OFF-SITE ----
if command -v rclone > /dev/null; then
rclone copy "$UPLOAD" "backup-remote:myapp-backups/" || fail "off-site upload"
echo " uploaded to off-site storage"
else
echo " ⚠️ rclone not installed — NO OFF-SITE COPY"
fi
# ---- 10. Retention ----
find "$BACKUP_ROOT" -maxdepth 1 -type d -name "20*" -mtime +$RETAIN_DAYS -exec rm -rf {} + 2>/dev/null || true
find "$BACKUP_ROOT" -maxdepth 1 -type f -name "*.tar.gz*" -mtime +$RETAIN_DAYS -delete 2>/dev/null || true
# ---- 11. Success ping ----
[ -n "$HC_PING_URL" ] && curl -fsS -m 10 --retry 3 "$HC_PING_URL" > /dev/null || true
echo "==> ✅ Backup complete: $DEST ($(du -sh "$DEST" | cut -f1)), off-site uploaded"# SERVER
sudo mkdir -p /var/backups/myapp
sudo chown deploy:deploy /var/backups/myapp
sudo chmod 700 /var/backups/myapp
chmod +x /home/deploy/scripts/backup.shWhy each verification step exists
STEP 3 IS THE STEP THAT MAKES THIS A REAL BACKUP SCRIPT
[ "$SIZE" -lt 10000 ] && fail "dump suspiciously small"
pg_restore --list "$DEST/database.dump" > /dev/null 2>&1 || fail "not readable"
TABLES=$(pg_restore --list ... | grep -c "TABLE DATA")
[ "$TABLES" -lt 1 ] && fail "no table data"Without these, a pg_dump that fails on authentication writes a tiny or empty file and cron reports success. You discover it months later.
The checks catch: authentication failure, disk full mid-write, wrong database name, truncated transfer, and corruption. They cost milliseconds.
set -euo pipefail PLUS AN ERR TRAP
set -e alone does not catch everything — notably a failure inside a pipeline without pipefail, or inside a command substitution. The explicit fail() calls after each critical command are belt-and-braces, and the trap ... ERR catches the rest.
Schedule it
# SERVER
crontab -e# Daily backup at 03:15 UTC
15 3 * * * /home/deploy/scripts/backup.sh >> /home/deploy/logs/backup.log 2>&1
# Weekly restore verification, Sunday 04:00
0 4 * * 0 /home/deploy/scripts/verify-restore.sh >> /home/deploy/logs/restore-test.log 2>&1CRON RUNS WITH A MINIMAL ENVIRONMENT
No NVM, a short PATH, and $HOME may differ. The classic symptom is a script that works perfectly when you run it and silently does nothing from cron.
Use absolute paths for everything, and set PATH at the top of the crontab:
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
SHELL=/bin/bash
MAILTO=you@example.comTest exactly as cron would:
env -i /bin/bash --noprofile --norc -c '/home/deploy/scripts/backup.sh'THE DEAD-MAN'S SWITCH IS NOT OPTIONAL
A backup script's most likely failure is not running at all — cron removed, path changed, disk full, server rebooted into a broken state. In that case there is no error to notice; there is simply silence.
Healthchecks.io (free for 20 checks) alerts you when the expected ping does not arrive. Set the period to 1 day with a 2-hour grace window. (Level 21)
Off-site storage
rclone
# SERVER
sudo -v ; curl https://rclone.org/install.sh | sudo bash
rclone configWorks with S3, Backblaze B2, Cloudflare R2, Hetzner Storage Box, Google Drive, and dozens more.
# SERVER
rclone copy /var/backups/myapp/20260811.tar.gz.age backup-remote:myapp-backups/
rclone ls backup-remote:myapp-backups/
rclone check /var/backups/myapp backup-remote:myapp-backups/COST COMPARISON FOR BACKUP STORAGE
| Provider | ~Storage | Egress | Note |
|---|---|---|---|
| Cloudflare R2 | ~$0.015/GB/mo | Free | Free egress is a big deal for restores |
| Backblaze B2 | ~$0.006/GB/mo | Free to Cloudflare | Cheapest storage |
| Hetzner Storage Box | ~€3.20/mo for 1 TB | Free | Flat rate, supports SSH/rsync/borg |
| AWS S3 | ~$0.023/GB/mo | $0.09/GB | Egress cost bites during a restore |
Egress matters: with S3, downloading a 50 GB backup during an emergency costs ~$4.50 and takes as long as your link allows. R2 and B2 charge nothing.
Use a different provider than your VPS. If Hetzner has an incident, Hetzner Storage Box may be affected too.
Encrypt before uploading
BACKUPS CONTAIN EVERYTHING — USER DATA, PASSWORD HASHES, AND YOUR .env
Uploading them unencrypted means a misconfigured bucket, a leaked API key, or a provider breach exposes your entire database.
# SERVER — age is simple and modern
sudo apt install -y age
# LOCAL — generate a keypair; keep the PRIVATE key OFF the server
age-keygen -o backup-key.txt
# Public key: age1ql3z7hjy54pw3hyww5ayyfg7zqgvc7w3j2elw8zmrj2kg5sfn9aqmcac8pEncrypt with the public key on the server; only the private key (in your password manager, and printed on paper in a safe) can decrypt.
tar -czf - backup-dir | age -r age1ql3z7... > backup.tar.gz.age
age -d -i backup-key.txt backup.tar.gz.age | tar -xzf -Store the private key somewhere that survives losing the server, and test decryption during your drill. An encrypted backup you cannot decrypt is worse than no backup — it wastes the time you spend trying.
Server snapshots
Every provider offers full-disk snapshots.
| Snapshot | pg_dump | |
|---|---|---|
| Captures | Entire disk | Just the database |
| Restore | Whole server, ~5 minutes | Database only |
| Granularity | All or nothing | Selective tables |
| Database consistency | ⚠️ Crash-consistent, not transaction-consistent | ✅ Consistent |
| Cost | ~€0.01/GB/month | Storage only |
| Off-site | Same provider | Wherever you put it |
A SNAPSHOT OF A RUNNING DATABASE IS "CRASH-CONSISTENT"
It is equivalent to pulling the power cable. PostgreSQL will recover on start by replaying the WAL, and usually succeeds — but it is not guaranteed, and it will not catch application-level inconsistency.
Use both: snapshots for fast whole-server recovery, pg_dump for guaranteed-consistent data. They solve different problems.
AUTOMATE SNAPSHOTS BEFORE RISKY OPERATIONS
# LOCAL — before a major migration or an OS upgrade
hcloud server create-image --type snapshot --description "pre-migration-$(date -u +%F)" my-serverA snapshot taken five minutes before a risky change turns "we might lose everything" into "we lose five minutes".
Restoring
Full database restore
# SERVER — 1. Stop the application first
pm2 stop all
# 2. Restore roles (only needed on a fresh cluster)
sudo -u postgres psql -f /var/backups/myapp/20260811/globals.sql
# 3. Restore the database
pg_restore -U myapp -h 127.0.0.1 -d myapp_production \
--clean --if-exists --no-owner --no-privileges \
/var/backups/myapp/20260811/database.dump
# 4. Verify
psql -U myapp -h 127.0.0.1 -d myapp_production -c "SELECT count(*) FROM users;"
# 5. Restart
pm2 restart all| Flag | Meaning |
|---|---|
--clean | Drop objects before recreating them |
--if-exists | Do not error when dropping something absent |
--no-owner | Do not attempt to set ownership — useful when restoring as a different role |
--no-privileges | Skip GRANT/REVOKE statements |
-j 4 | Parallel restore (directory format only) |
--clean DROPS YOUR CURRENT DATA BEFORE RESTORING
If the dump turns out to be corrupt, you have now destroyed the live database and failed to restore. There is no undo.
Always restore into a new database first, verify, then swap:
sudo -u postgres createdb myapp_restore
pg_restore -U postgres -d myapp_restore backup.dump
psql -U postgres -d myapp_restore -c "SELECT count(*) FROM users;"
# Only when satisfied:
sudo -u postgres psql -c "ALTER DATABASE myapp_production RENAME TO myapp_broken;"
sudo -u postgres psql -c "ALTER DATABASE myapp_restore RENAME TO myapp_production;"Renames are instant and reversible. --clean is not.
Restoring a single table
# SERVER — extract just one table from the dump
pg_restore -U myapp -h 127.0.0.1 -d myapp_production -t users --data-only backup.dump
# Or into a scratch database, then copy the rows you need
sudo -u postgres createdb scratch
pg_restore -U postgres -d scratch -t users backup.dump
psql -U postgres -d scratch -c "\copy (SELECT * FROM users WHERE id='...') TO '/tmp/row.csv' CSV"Being able to recover one accidentally-deleted row without a full restore is the main practical advantage of the custom dump format.
The restore drill
THIS SECTION IS THE POINT OF THE ENTIRE CHAPTER
Schedule it. Do it monthly. Write down the date.
#!/usr/bin/env bash
# /home/deploy/scripts/verify-restore.sh
set -euo pipefail
LATEST=$(ls -1dt /var/backups/myapp/20* 2>/dev/null | head -1)
[ -z "$LATEST" ] && { echo "❌ No backup found"; exit 1; }
TEST_DB="restore_verify_$(date +%s)"
echo "==> Verifying $LATEST into $TEST_DB"
sudo -u postgres createdb "$TEST_DB"
trap 'sudo -u postgres dropdb --if-exists "$TEST_DB"' EXIT
sudo -u postgres pg_restore -d "$TEST_DB" --no-owner --no-privileges "$LATEST/database.dump" 2>&1 \
| grep -v "^pg_restore: warning" || true
# Structural check
TABLES=$(sudo -u postgres psql -d "$TEST_DB" -tAc \
"SELECT count(*) FROM information_schema.tables WHERE table_schema='public'")
[ "$TABLES" -lt 1 ] && { echo "❌ No tables restored"; exit 1; }
# Data check — adjust to your schema
USERS=$(sudo -u postgres psql -d "$TEST_DB" -tAc "SELECT count(*) FROM users" 2>/dev/null || echo 0)
[ "$USERS" -lt 1 ] && { echo "❌ users table is empty"; exit 1; }
# Freshness check — is the newest row recent?
NEWEST=$(sudo -u postgres psql -d "$TEST_DB" -tAc \
"SELECT EXTRACT(EPOCH FROM (now() - max(\"createdAt\")))/3600 FROM users" 2>/dev/null || echo 0)
echo "✅ Restore verified: ${TABLES} tables, ${USERS} users, newest record ${NEWEST%.*}h old"
echo " Verified at $(date -u +%FT%TZ)" >> /home/deploy/logs/restore-verified.logThe manual drill — quarterly, on a scratch server
Automated verification proves the dump is restorable. The full drill proves you can rebuild the service.
- [ ] Provision a brand-new VPS
- [ ] Download the latest off-site backup using only credentials from your password manager
- [ ] Decrypt it
- [ ] Install PostgreSQL, restore globals, restore the database
- [ ] Restore
.envfrom the backup - [ ] Clone the repository,
pnpm install, build - [ ] Restore Nginx configuration
- [ ] Start the application
- [ ] Verify you can log in and read real data
- [ ] Time the whole thing — that number is your true RTO
- [ ] Write down what went wrong and fix the runbook
WHAT THE DRILL ALWAYS REVEALS
Every first drill uncovers at least one of these:
- The
.envwas not in the backup - The
myapprole did not exist, so the restore threw hundreds of errors - The uploads directory was never backed up
- The decryption key was only on the server that died
- The off-site credentials were only in the
.envon that same server - The restore took 4 hours, not the 20 minutes you assumed
Better to find out on a Tuesday afternoon than during an outage.
RPO and RTO
| Term | Question | Determined by |
|---|---|---|
| RPO — Recovery Point Objective | How much data can we afford to lose? | Backup frequency |
| RTO — Recovery Time Objective | How long can we be down? | Restore speed |
| Backup schedule | RPO |
|---|---|
| Daily at 03:00 | Up to 24 hours |
| Every 6 hours | Up to 6 hours |
| Hourly | Up to 1 hour |
| Continuous WAL archiving | Seconds |
DECIDE YOUR RPO DELIBERATELY, THEN BUILD TO IT
Ask: "if we lost the last N hours of data, what actually happens?"
- A blog: 24 hours is fine.
- A SaaS with paying customers: 24 hours means a day of lost work for everyone. Go hourly.
- Anything handling payments: you need WAL archiving and point-in-time recovery.
Then be honest about whether your current setup meets it. A daily backup with a 24-hour RPO is a legitimate choice — pretending a daily backup gives you a 1-hour RPO is not.
Point-in-time recovery
For a genuinely low RPO, archive the WAL continuously:
# postgresql.conf
archive_mode = on
archive_command = 'rclone copy %p backup-remote:wal-archive/ 2>/dev/null'
archive_timeout = 300
wal_level = replicaYou can then restore to any second between your last base backup and now.
WAL ARCHIVING NEEDS CAREFUL OPERATION
If archive_command fails persistently, PostgreSQL retains WAL segments waiting to archive them, and pg_wal/ grows until the disk is full — which takes the database down. Monitor pg_stat_archiver:
SELECT archived_count, failed_count, last_failed_time FROM pg_stat_archiver;Consider pgBackRest or WAL-G rather than hand-rolling this — they handle retention, compression, verification, and parallelism properly.
Redis backups
# SERVER
redis-cli BGSAVE # background snapshot
cp /var/lib/redis/dump.rdb /var/backups/myapp/DECIDE WHETHER REDIS DATA MATTERS
| Contents | Back up? |
|---|---|
| Cache only | ❌ No — a cold cache is a performance blip |
| Sessions | ⚠️ Nice to have — losing it logs everyone out |
| BullMQ queues | ✅ Yes — losing it drops queued jobs permanently |
| Socket.IO adapter | ❌ No — pure pub/sub, nothing stored |
If Redis holds queues, enable AOF persistence and include dump.rdb/appendonly files in the backup (Level 10).
Backup security
| Rule | Why |
|---|---|
| Encrypt before upload | A leaked bucket must not expose your data |
| Keep the decryption key off the server | Ransomware encrypts everything it can reach |
| Restrict backup storage credentials to write+list | A compromised server cannot delete your history |
| Enable object lock / immutability | Even a full compromise cannot destroy backups |
chmod 700 on the local backup directory | It contains .env and password hashes |
| Test that you can decrypt | An unopenable backup is not a backup |
| Do not commit backups to Git | Obvious, and it happens |
USE WRITE-ONLY CREDENTIALS FOR THE BACKUP DESTINATION
If the server's storage credentials can delete objects, ransomware will use them. Create an API key with PutObject and ListBucket but no DeleteObject, and enable versioning with object lock.
Then even a fully compromised production server cannot destroy your ability to recover. This one configuration choice is the difference between an incident and a catastrophe.
Production Checklist — Level 22
- [ ] Daily automated PostgreSQL backup, running from cron
- [ ]
pg_dumpall --globals-onlycaptured alongside the database dump - [ ] Backup script verifies the dump: size,
pg_restore --list, table count - [ ] Script exits non-zero and alerts on any failure
- [ ]
.envincluded in the backup and stored in a password manager - [ ] User uploads backed up
- [ ] Nginx config,
/etc/letsencrypt, crontab, and package list captured - [ ] Backups encrypted before leaving the server
- [ ] Decryption key stored off the server and its location documented
- [ ] Off-site copy on a different provider
- [ ] Backup storage credentials are write-only; object lock enabled
- [ ] Retention policy defined and enforced
- [ ] Dead-man's switch alerts when the backup does not run
- [ ] Provider snapshots enabled, plus manual ones before risky changes
- [ ] Automated weekly restore verification into a scratch database
- [ ] Full manual restore drill performed on a fresh server — date recorded
- [ ] I know my measured RTO, not my assumed one
- [ ] RPO chosen deliberately and matched by backup frequency
- [ ] Redis persistence configured if it holds queues
- [ ] Restore procedure documented in the runbook (Level 28)