Skip to content

Disaster Recovery Plan ​

What to do when the server is gone, the data is corrupt, or you have been compromised.

STORE THIS DOCUMENT SOMEWHERE THAT IS NOT THE SERVER

A recovery plan that lives only on the machine you are recovering is not a plan.

Keep a copy in: your password manager, a printed page in a drawer, a private repository on a different platform, and a note on your phone. Include the credentials list below — or at least where to find it.

Before an incident — the information you will need ​

Fill this in and store it with the plan. Doing it now takes ten minutes; doing it during an outage is impossible.

ItemWhere it isWho has access
VPS provider login1Password → "Hetzner"Aziz
VPS provider 2FA backup codes1PasswordAziz
Domain registrar login1Password → "Namecheap"Aziz
DNS provider login1Password → "Cloudflare"Aziz
SSH private key~/.ssh/id_ed25519 + 1PasswordAziz
Production .env contents1Password → "myapp production env"Aziz
Backup storage credentials1Password → "R2 backup"Aziz
Backup decryption key1Password → "myapp age key" + printed copyAziz
GitHub / GitLab access1PasswordAziz
Error tracker (Sentry)1PasswordAziz
Payment provider (Stripe)1PasswordAziz
Email/SMTP provider1PasswordAziz

THE CIRCULAR DEPENDENCY THAT BREAKS RECOVERY

If your backup storage credentials live only in the .env on the server, and the server is gone, you cannot reach your backups. If the decryption key was only on that server, the backups are useless.

Both must exist outside the server. Verify this during your restore drill — it is the failure that most commonly turns a recoverable incident into a permanent one.

Scenario 1 — The server is gone ​

Deleted by accident, provider failure, or unrecoverable corruption.

Impact: total outage. Target RTO: 1–2 hours (measure yours in a drill). RPO: since the last backup.

Steps ​

bash
# 1. Provision a new VPS — same region, same or larger size, Ubuntu 24.04 LTS
#    Add your SSH key at creation time.
bash
# 2. Run Phases 2–12 of the runbook (system, user, SSH, firewall, fail2ban,
#    Node, PostgreSQL, Redis, Nginx). See /deploy/25-from-zero-runbook
bash
# 3. LOCAL — retrieve the backup using credentials from your PASSWORD MANAGER
rclone copy backup-remote:myapp-backups/20260811-031500.tar.gz.age ./
age -d -i ~/myapp-backup-key.txt 20260811-031500.tar.gz.age | tar -xzf -
bash
# 4. Upload to the new server
scp -r 20260811-031500 deploy@NEW_IP:/tmp/restore/
bash
# SERVER — 5. Restore PostgreSQL roles, then the database
sudo -u postgres psql -f /tmp/restore/globals.sql
sudo -u postgres createdb myapp_production
pg_restore -U postgres -d myapp_production --no-owner --no-privileges \
  /tmp/restore/database.dump

sudo -u postgres psql -d myapp_production -c "SELECT count(*) FROM users;"
bash
# SERVER — 6. Restore secrets and uploads
mkdir -p ~/apps/myapp/shared
cp /tmp/restore/env.backup ~/apps/myapp/shared/.env
chmod 600 ~/apps/myapp/shared/.env
tar -xzf /tmp/restore/uploads.tar.gz -C ~/apps/myapp/shared/
bash
# SERVER — 7. Deploy the application
cd ~/apps/myapp
git clone git@github.com:myorg/myapp.git repo.git --bare
~/apps/myapp/deploy.sh
bash
# SERVER — 8. Restore Nginx configuration and re-issue certificates
sudo tar -xzf /tmp/restore/system-config.tar.gz -C /
sudo nginx -t && sudo systemctl reload nginx
sudo certbot --nginx -d app.example.com -d api.example.com -d example.com -d www.example.com
bash
# 9. Update DNS to the new IP (TTL 300 means ~5 minutes)
bash
# SERVER — 10. Restore cron and PM2 startup
crontab /tmp/restore/crontab.txt
pm2 startup    # run the printed command
pm2 save
bash
# 11. Verify everything
curl -I https://app.example.com
curl -s https://api.example.com/api/health | jq
bash
# SERVER — 12. Clean up
rm -rf /tmp/restore

A SNAPSHOT MAKES THIS 5 MINUTES INSTEAD OF 2 HOURS

If you have a recent provider snapshot, restore it, update DNS, and then restore just the database from your most recent dump to close the gap.

Take snapshots regularly and always before risky changes. They do not replace pg_dump (they are only crash-consistent), but they collapse the rebuild time enormously.

Scenario 2 — Database corruption or bad data ​

An accidental DELETE, a migration that destroyed a column, or genuine corruption.

Impact: wrong or missing data; the site may still be up. Target RTO: 30 minutes.

Steps ​

bash
# SERVER — 1. STOP WRITES IMMEDIATELY
pm2 stop all

STOPPING THE APPLICATION IS THE FIRST STEP, ALWAYS

Every second the app keeps running, more writes land on top of the damage and widen the gap between your backup and reality. Take the outage; it is shorter than the alternative.

bash
# SERVER — 2. Snapshot the CURRENT (damaged) state before doing anything
pg_dump -U myapp -h 127.0.0.1 -Fc -f /var/backups/pre-restore-$(date -u +%Y%m%d-%H%M%S).dump myapp_production

You may need rows written after the backup but before the damage. Do not destroy the evidence.

bash
# SERVER — 3. Restore into a NEW database, never over the live one
sudo -u postgres createdb myapp_restore
pg_restore -U postgres -d myapp_restore --no-owner --no-privileges /var/backups/myapp/LATEST/database.dump
bash
# SERVER — 4. Verify the restored data
sudo -u postgres psql -d myapp_restore -c "SELECT count(*) FROM users;"
sudo -u postgres psql -d myapp_restore -c "SELECT max(\"createdAt\") FROM orders;"
bash
# SERVER — 5. Swap by rename — instant and reversible
sudo -u postgres psql <<'SQL'
ALTER DATABASE myapp_production RENAME TO myapp_damaged;
ALTER DATABASE myapp_restore RENAME TO myapp_production;
SQL
bash
# SERVER — 6. Restart and verify
pm2 restart all
curl -s http://127.0.0.1:3001/api/health
bash
# SERVER — 7. Recover the gap from myapp_damaged if needed
# e.g. orders created after the backup but before the damage

NEVER pg_restore --clean ONTO THE LIVE DATABASE

It drops your current data before restoring. If the dump turns out to be corrupt, you have now lost both. The rename approach is instant, reversible, and leaves the damaged copy available for gap recovery.

Scenario 3 — Compromise ​

You have evidence of unauthorised access.

Impact: potential data breach; the server cannot be trusted. Target RTO: 2–4 hours.

DO NOT TRY TO CLEAN A COMPROMISED SERVER

A competent attacker installs several persistence mechanisms. Finding all of them is not realistic. Rebuild. (Level 23)

Immediate (first 15 minutes) ​

  1. Do not power off — RAM is evidence, and a reboot may trigger persistence.
  2. Isolate at the cloud firewall (not UFW — you would lose your own access): block all inbound except your IP.
  3. Snapshot the disk for forensics.
  4. Assume every secret on that server is compromised.

Assess (next hour) ​

bash
# SERVER
last -50
sudo grep "Accepted" /var/log/auth.log | tail -50
ps auxf
sudo ss -tupn
cat ~/.ssh/authorized_keys /root/.ssh/authorized_keys
sudo crontab -l; crontab -l; ls -la /etc/cron.d/
systemctl list-units --type=service --state=running
sudo find / -newermt "3 days ago" -type f -not -path "/proc/*" -not -path "/sys/*" 2>/dev/null | head -50

Determine: entry point, time of first access, what was reachable.

Rebuild ​

  1. Provision a new server.
  2. Apply full hardening (Levels 3–5).
  3. Generate new SSH keys — the old ones are suspect.
  4. Restore the database from a backup taken before the compromise.
  5. Deploy code from Git after reviewing recent commits for injected changes.
  6. Rotate every secret: database passwords, JWT secrets, refresh token secrets, API keys, OAuth secrets, webhook secrets, CI/CD secrets, provider API tokens, deploy keys.
  7. Point DNS at the new server.
  8. Destroy the old server, keeping the forensic snapshot.

Notify ​

WhoWhenWhy
UsersWithin 72 hours if personal data was accessedGDPR and equivalent laws
Payment processorImmediately if payment data was involvedContractual obligation
Hosting providerIf their infrastructure was implicatedThey may have more information
Data protection authorityPer your jurisdiction's rulesLegal requirement

Post-mortem ​

Write it down: entry point, dwell time, what was reached, and which control would have prevented it. Then implement that control.

Scenario 4 — Certificate expired ​

Impact: browsers show a full-page security warning. With HSTS, users cannot click through. Target RTO: 15 minutes.

bash
# SERVER
sudo certbot certificates
sudo ufw status | grep 80              # is port 80 open?
sudo certbot renew --force-renewal
sudo systemctl reload nginx

If renewal fails:

bash
# SERVER
sudo systemctl stop nginx
sudo certbot certonly --standalone -d app.example.com -d api.example.com
sudo systemctl start nginx

Then find the root cause — usually port 80 was closed, the challenge location was removed from the Nginx config, or the reload hook is missing so Nginx serves a stale certificate from memory.

Scenario 5 — Domain or DNS failure ​

Impact: total outage; the server is fine but unreachable.

CauseFix
Domain expiredRenew immediately; some TLDs have a grace period
Nameservers changedRestore them at the registrar
Records deletedRecreate; TTL 300 means ~5 minutes
DNS provider outageSwitch nameservers to a backup provider
Registrar account compromisedContact the registrar; enable registrar lock

KEEP A COPY OF YOUR DNS ZONE

Export your zone file and store it with this plan. Recreating 12 records from memory during an outage — including the DKIM record that makes your email work — is worse than it sounds.

Scenario 6 — Provider-wide outage ​

Impact: total outage, and you cannot do anything about it.

  1. Check the provider's status page.
  2. Communicate with users — a status page or a tweet beats silence.
  3. If the outage is prolonged and you have off-site backups, execute Scenario 1 at a different provider.
  4. Afterwards, decide whether multi-provider redundancy is worth the cost.

THIS IS WHY OFF-SITE BACKUPS MUST BE AT A DIFFERENT PROVIDER

Hetzner backups do not help during a Hetzner incident. Cross-provider storage is what makes Scenario 6 recoverable rather than a wait. (Level 22)

Communication during an incident ​

DurationAction
< 5 minFix it; no communication needed
5–15 minPost to your status channel
15–60 minStatus page update with an ETA; email key customers
> 1 hourRegular updates every 30 minutes
Data lossIndividual notification, with specifics

Template:

[Investigating] We are aware of an issue affecting access to the application, starting at 14:20 UTC. We are investigating and will update within 30 minutes.

[Identified] The cause is a database connectivity failure. We are restoring service. Estimated resolution: 15:10 UTC.

[Resolved] Service was restored at 15:05 UTC. Total impact: 45 minutes. A post-mortem will follow within 48 hours.

SAY SOMETHING EARLY, EVEN WITHOUT ANSWERS

"We are investigating" posted at minute five buys enormous goodwill compared to silence followed by a detailed explanation at minute ninety. Users forgive outages; they do not forgive being ignored.

Recovery time expectations ​

ScenarioRealistic RTORPO
Bad deploy5 seconds (symlink rollback)0
App crashAutomatic (PM2)0
Certificate expired15 minutes0
DNS misconfiguration5–30 minutes (TTL)0
Database corruption30 minutesSince last backup
Server destroyed (with snapshot)30 minutesSince last backup
Server destroyed (from backups)1–2 hoursSince last backup
Compromise2–4 hoursBefore the compromise
Provider outageProvider-dependent, or 2 hours to migrateSince last backup

THESE NUMBERS ARE ONLY REAL IF YOU HAVE TESTED THEM

Every number above assumes the backup is valid, the decryption key is reachable, and you have done this before. Untested, they are guesses — and the true figure is usually 3–5× longer, because you discover the missing .env and the absent database role on the day.

Do the drill. Record the actual time. That number is your RTO.

Drill schedule ​

DrillFrequencyVerifies
RollbackMonthlyThe rollback script works
Database restore into a scratch DBWeekly (automated)The dump is valid
Full rebuild on a fresh VPSQuarterlyThe entire plan, end to end
Compromise walkthrough (tabletop)AnnuallyYou know the steps

The full drill checklist ​

  • [ ] Provision a new VPS
  • [ ] Retrieve backups using only password-manager credentials
  • [ ] Decrypt them
  • [ ] Restore PostgreSQL globals and the database
  • [ ] Restore .env
  • [ ] Deploy the application from Git
  • [ ] Restore Nginx config and issue certificates
  • [ ] Verify login and a core user flow with real data
  • [ ] Record the elapsed time
  • [ ] Note everything that was missing or wrong
  • [ ] Update this document
  • [ ] Destroy the test VPS

WHAT THE FIRST DRILL ALWAYS FINDS

  • The .env was not in the backup
  • The database role did not exist, so the restore threw hundreds of errors
  • The uploads directory was never backed up
  • The decryption key was only on the server that died
  • The backup credentials were only in the .env on that same server
  • It took four hours, not the twenty minutes you assumed

Better on a quiet Tuesday than during a real outage.


Next: Cheat Sheets →