Disaster Recovery Plan
What to do when the server is gone, the data is corrupt, or you have been compromised.
STORE THIS DOCUMENT SOMEWHERE THAT IS NOT THE SERVER
A recovery plan that lives only on the machine you are recovering is not a plan.
Keep a copy in: your password manager, a printed page in a drawer, a private repository on a different platform, and a note on your phone. Include the credentials list below — or at least where to find it.
Before an incident — the information you will need
Fill this in and store it with the plan. Doing it now takes ten minutes; doing it during an outage is impossible.
| Item | Where it is | Who has access |
|---|---|---|
| VPS provider login | 1Password → "Hetzner" | Aziz |
| VPS provider 2FA backup codes | 1Password | Aziz |
| Domain registrar login | 1Password → "Namecheap" | Aziz |
| DNS provider login | 1Password → "Cloudflare" | Aziz |
| SSH private key | ~/.ssh/id_ed25519 + 1Password | Aziz |
Production .env contents | 1Password → "myapp production env" | Aziz |
| Backup storage credentials | 1Password → "R2 backup" | Aziz |
| Backup decryption key | 1Password → "myapp age key" + printed copy | Aziz |
| GitHub / GitLab access | 1Password | Aziz |
| Error tracker (Sentry) | 1Password | Aziz |
| Payment provider (Stripe) | 1Password | Aziz |
| Email/SMTP provider | 1Password | Aziz |
THE CIRCULAR DEPENDENCY THAT BREAKS RECOVERY
If your backup storage credentials live only in the .env on the server, and the server is gone, you cannot reach your backups. If the decryption key was only on that server, the backups are useless.
Both must exist outside the server. Verify this during your restore drill — it is the failure that most commonly turns a recoverable incident into a permanent one.
Scenario 1 — The server is gone
Deleted by accident, provider failure, or unrecoverable corruption.
Impact: total outage. Target RTO: 1–2 hours (measure yours in a drill). RPO: since the last backup.
Steps
# 1. Provision a new VPS — same region, same or larger size, Ubuntu 24.04 LTS
# Add your SSH key at creation time.# 2. Run Phases 2–12 of the runbook (system, user, SSH, firewall, fail2ban,
# Node, PostgreSQL, Redis, Nginx). See /deploy/25-from-zero-runbook# 3. LOCAL — retrieve the backup using credentials from your PASSWORD MANAGER
rclone copy backup-remote:myapp-backups/20260811-031500.tar.gz.age ./
age -d -i ~/myapp-backup-key.txt 20260811-031500.tar.gz.age | tar -xzf -# 4. Upload to the new server
scp -r 20260811-031500 deploy@NEW_IP:/tmp/restore/# SERVER — 5. Restore PostgreSQL roles, then the database
sudo -u postgres psql -f /tmp/restore/globals.sql
sudo -u postgres createdb myapp_production
pg_restore -U postgres -d myapp_production --no-owner --no-privileges \
/tmp/restore/database.dump
sudo -u postgres psql -d myapp_production -c "SELECT count(*) FROM users;"# SERVER — 6. Restore secrets and uploads
mkdir -p ~/apps/myapp/shared
cp /tmp/restore/env.backup ~/apps/myapp/shared/.env
chmod 600 ~/apps/myapp/shared/.env
tar -xzf /tmp/restore/uploads.tar.gz -C ~/apps/myapp/shared/# SERVER — 7. Deploy the application
cd ~/apps/myapp
git clone git@github.com:myorg/myapp.git repo.git --bare
~/apps/myapp/deploy.sh# SERVER — 8. Restore Nginx configuration and re-issue certificates
sudo tar -xzf /tmp/restore/system-config.tar.gz -C /
sudo nginx -t && sudo systemctl reload nginx
sudo certbot --nginx -d app.example.com -d api.example.com -d example.com -d www.example.com# 9. Update DNS to the new IP (TTL 300 means ~5 minutes)# SERVER — 10. Restore cron and PM2 startup
crontab /tmp/restore/crontab.txt
pm2 startup # run the printed command
pm2 save# 11. Verify everything
curl -I https://app.example.com
curl -s https://api.example.com/api/health | jq# SERVER — 12. Clean up
rm -rf /tmp/restoreA SNAPSHOT MAKES THIS 5 MINUTES INSTEAD OF 2 HOURS
If you have a recent provider snapshot, restore it, update DNS, and then restore just the database from your most recent dump to close the gap.
Take snapshots regularly and always before risky changes. They do not replace pg_dump (they are only crash-consistent), but they collapse the rebuild time enormously.
Scenario 2 — Database corruption or bad data
An accidental DELETE, a migration that destroyed a column, or genuine corruption.
Impact: wrong or missing data; the site may still be up. Target RTO: 30 minutes.
Steps
# SERVER — 1. STOP WRITES IMMEDIATELY
pm2 stop allSTOPPING THE APPLICATION IS THE FIRST STEP, ALWAYS
Every second the app keeps running, more writes land on top of the damage and widen the gap between your backup and reality. Take the outage; it is shorter than the alternative.
# SERVER — 2. Snapshot the CURRENT (damaged) state before doing anything
pg_dump -U myapp -h 127.0.0.1 -Fc -f /var/backups/pre-restore-$(date -u +%Y%m%d-%H%M%S).dump myapp_productionYou may need rows written after the backup but before the damage. Do not destroy the evidence.
# SERVER — 3. Restore into a NEW database, never over the live one
sudo -u postgres createdb myapp_restore
pg_restore -U postgres -d myapp_restore --no-owner --no-privileges /var/backups/myapp/LATEST/database.dump# SERVER — 4. Verify the restored data
sudo -u postgres psql -d myapp_restore -c "SELECT count(*) FROM users;"
sudo -u postgres psql -d myapp_restore -c "SELECT max(\"createdAt\") FROM orders;"# SERVER — 5. Swap by rename — instant and reversible
sudo -u postgres psql <<'SQL'
ALTER DATABASE myapp_production RENAME TO myapp_damaged;
ALTER DATABASE myapp_restore RENAME TO myapp_production;
SQL# SERVER — 6. Restart and verify
pm2 restart all
curl -s http://127.0.0.1:3001/api/health# SERVER — 7. Recover the gap from myapp_damaged if needed
# e.g. orders created after the backup but before the damageNEVER pg_restore --clean ONTO THE LIVE DATABASE
It drops your current data before restoring. If the dump turns out to be corrupt, you have now lost both. The rename approach is instant, reversible, and leaves the damaged copy available for gap recovery.
Scenario 3 — Compromise
You have evidence of unauthorised access.
Impact: potential data breach; the server cannot be trusted. Target RTO: 2–4 hours.
DO NOT TRY TO CLEAN A COMPROMISED SERVER
A competent attacker installs several persistence mechanisms. Finding all of them is not realistic. Rebuild. (Level 23)
Immediate (first 15 minutes)
- Do not power off — RAM is evidence, and a reboot may trigger persistence.
- Isolate at the cloud firewall (not UFW — you would lose your own access): block all inbound except your IP.
- Snapshot the disk for forensics.
- Assume every secret on that server is compromised.
Assess (next hour)
# SERVER
last -50
sudo grep "Accepted" /var/log/auth.log | tail -50
ps auxf
sudo ss -tupn
cat ~/.ssh/authorized_keys /root/.ssh/authorized_keys
sudo crontab -l; crontab -l; ls -la /etc/cron.d/
systemctl list-units --type=service --state=running
sudo find / -newermt "3 days ago" -type f -not -path "/proc/*" -not -path "/sys/*" 2>/dev/null | head -50Determine: entry point, time of first access, what was reachable.
Rebuild
- Provision a new server.
- Apply full hardening (Levels 3–5).
- Generate new SSH keys — the old ones are suspect.
- Restore the database from a backup taken before the compromise.
- Deploy code from Git after reviewing recent commits for injected changes.
- Rotate every secret: database passwords, JWT secrets, refresh token secrets, API keys, OAuth secrets, webhook secrets, CI/CD secrets, provider API tokens, deploy keys.
- Point DNS at the new server.
- Destroy the old server, keeping the forensic snapshot.
Notify
| Who | When | Why |
|---|---|---|
| Users | Within 72 hours if personal data was accessed | GDPR and equivalent laws |
| Payment processor | Immediately if payment data was involved | Contractual obligation |
| Hosting provider | If their infrastructure was implicated | They may have more information |
| Data protection authority | Per your jurisdiction's rules | Legal requirement |
Post-mortem
Write it down: entry point, dwell time, what was reached, and which control would have prevented it. Then implement that control.
Scenario 4 — Certificate expired
Impact: browsers show a full-page security warning. With HSTS, users cannot click through. Target RTO: 15 minutes.
# SERVER
sudo certbot certificates
sudo ufw status | grep 80 # is port 80 open?
sudo certbot renew --force-renewal
sudo systemctl reload nginxIf renewal fails:
# SERVER
sudo systemctl stop nginx
sudo certbot certonly --standalone -d app.example.com -d api.example.com
sudo systemctl start nginxThen find the root cause — usually port 80 was closed, the challenge location was removed from the Nginx config, or the reload hook is missing so Nginx serves a stale certificate from memory.
Scenario 5 — Domain or DNS failure
Impact: total outage; the server is fine but unreachable.
| Cause | Fix |
|---|---|
| Domain expired | Renew immediately; some TLDs have a grace period |
| Nameservers changed | Restore them at the registrar |
| Records deleted | Recreate; TTL 300 means ~5 minutes |
| DNS provider outage | Switch nameservers to a backup provider |
| Registrar account compromised | Contact the registrar; enable registrar lock |
KEEP A COPY OF YOUR DNS ZONE
Export your zone file and store it with this plan. Recreating 12 records from memory during an outage — including the DKIM record that makes your email work — is worse than it sounds.
Scenario 6 — Provider-wide outage
Impact: total outage, and you cannot do anything about it.
- Check the provider's status page.
- Communicate with users — a status page or a tweet beats silence.
- If the outage is prolonged and you have off-site backups, execute Scenario 1 at a different provider.
- Afterwards, decide whether multi-provider redundancy is worth the cost.
THIS IS WHY OFF-SITE BACKUPS MUST BE AT A DIFFERENT PROVIDER
Hetzner backups do not help during a Hetzner incident. Cross-provider storage is what makes Scenario 6 recoverable rather than a wait. (Level 22)
Communication during an incident
| Duration | Action |
|---|---|
| < 5 min | Fix it; no communication needed |
| 5–15 min | Post to your status channel |
| 15–60 min | Status page update with an ETA; email key customers |
| > 1 hour | Regular updates every 30 minutes |
| Data loss | Individual notification, with specifics |
Template:
[Investigating] We are aware of an issue affecting access to the application, starting at 14:20 UTC. We are investigating and will update within 30 minutes.
[Identified] The cause is a database connectivity failure. We are restoring service. Estimated resolution: 15:10 UTC.
[Resolved] Service was restored at 15:05 UTC. Total impact: 45 minutes. A post-mortem will follow within 48 hours.
SAY SOMETHING EARLY, EVEN WITHOUT ANSWERS
"We are investigating" posted at minute five buys enormous goodwill compared to silence followed by a detailed explanation at minute ninety. Users forgive outages; they do not forgive being ignored.
Recovery time expectations
| Scenario | Realistic RTO | RPO |
|---|---|---|
| Bad deploy | 5 seconds (symlink rollback) | 0 |
| App crash | Automatic (PM2) | 0 |
| Certificate expired | 15 minutes | 0 |
| DNS misconfiguration | 5–30 minutes (TTL) | 0 |
| Database corruption | 30 minutes | Since last backup |
| Server destroyed (with snapshot) | 30 minutes | Since last backup |
| Server destroyed (from backups) | 1–2 hours | Since last backup |
| Compromise | 2–4 hours | Before the compromise |
| Provider outage | Provider-dependent, or 2 hours to migrate | Since last backup |
THESE NUMBERS ARE ONLY REAL IF YOU HAVE TESTED THEM
Every number above assumes the backup is valid, the decryption key is reachable, and you have done this before. Untested, they are guesses — and the true figure is usually 3–5× longer, because you discover the missing .env and the absent database role on the day.
Do the drill. Record the actual time. That number is your RTO.
Drill schedule
| Drill | Frequency | Verifies |
|---|---|---|
| Rollback | Monthly | The rollback script works |
| Database restore into a scratch DB | Weekly (automated) | The dump is valid |
| Full rebuild on a fresh VPS | Quarterly | The entire plan, end to end |
| Compromise walkthrough (tabletop) | Annually | You know the steps |
The full drill checklist
- [ ] Provision a new VPS
- [ ] Retrieve backups using only password-manager credentials
- [ ] Decrypt them
- [ ] Restore PostgreSQL globals and the database
- [ ] Restore
.env - [ ] Deploy the application from Git
- [ ] Restore Nginx config and issue certificates
- [ ] Verify login and a core user flow with real data
- [ ] Record the elapsed time
- [ ] Note everything that was missing or wrong
- [ ] Update this document
- [ ] Destroy the test VPS
WHAT THE FIRST DRILL ALWAYS FINDS
- The
.envwas not in the backup - The database role did not exist, so the restore threw hundreds of errors
- The uploads directory was never backed up
- The decryption key was only on the server that died
- The backup credentials were only in the
.envon that same server - It took four hours, not the twenty minutes you assumed
Better on a quiet Tuesday than during a real outage.
Next: Cheat Sheets →