Skip to content

Level 27 — Production Checklists ​

Copy these. Work through them literally. The value is in not trusting your memory.

New VPS setup ​

Follow Level 25 for the commands; this is the tracking list.

Before touching the server ​

  • [ ] SSH key pair generated with a passphrase
  • [ ] VPS created with Ubuntu 24.04 LTS, public key added at creation time
  • [ ] Cloud firewall: inbound 22, 80, 443 only
  • [ ] DNS A records created for @, app, api, www with TTL 300
  • [ ] CAA record allowing letsencrypt.org
  • [ ] Provider web console located — you know where the emergency access button is

System ​

  • [ ] apt update && apt upgrade run; rebooted if the kernel changed
  • [ ] Hostname set; /etc/hosts entry added
  • [ ] Timezone set to UTC; time sync confirmed
  • [ ] tmux, htop, curl, git, jq, ncdu installed
  • [ ] Swap configured if RAM ≤ 4 GB, with vm.swappiness=10
  • [ ] journald storage set to persistent with a size cap

Users and access ​

  • [ ] deploy user created; password saved in a password manager
  • [ ] deploy added to sudo with usermod -aG
  • [ ] SSH key installed for deploy; 700/600 permissions verified
  • [ ] ssh deploy@server and sudo whoami tested in a second terminal
  • [ ] Root SSH login disabled
  • [ ] Password authentication disabled — verified by attempting one
  • [ ] AllowUsers restricts SSH
  • [ ] Config verified with sudo sshd -T, not by reading the file

Firewall and intrusion prevention ​

  • [ ] UFW default deny incoming, allow outgoing
  • [ ] SSH allowed before ufw enable
  • [ ] Only 22, 80, 443 allowed; IPv4 and IPv6 rules present
  • [ ] fail2ban installed; sshd jail confirmed active with fail2ban-client status sshd
  • [ ] My own IP in ignoreip
  • [ ] unattended-upgrades enabled and dry-run verified

Runtime ​

  • [ ] NVM installed as deploy, not root
  • [ ] Node LTS installed; nvm alias default set
  • [ ] pnpm via Corepack; version pinned in packageManager
  • [ ] PM2 installed globally without sudo
  • [ ] ssh deploy@server "node -v && pnpm -v && pm2 -v" works from my laptop
  • [ ] Binaries symlinked into /usr/local/bin
  • [ ] .nvmrc matches CI and local development

Data services ​

  • [ ] PostgreSQL installed and enabled on boot
  • [ ] sudo ss -tulpn | grep 5432 shows 127.0.0.1
  • [ ] Application role created — not postgres
  • [ ] Strong, URL-safe password generated with openssl rand
  • [ ] ALTER DEFAULT PRIVILEGES set
  • [ ] Memory tuned for the server's RAM
  • [ ] pg_stat_statements enabled
  • [ ] OOMScoreAdjust=-500 set on PostgreSQL
  • [ ] pg_hba.conf uses scram-sha-256; no trust
  • [ ] Redis installed with bind 127.0.0.1 and requirepass
  • [ ] redis-cli ping without auth returns NOAUTH
  • [ ] Dangerous Redis commands renamed or disabled
  • [ ] maxmemory set with an appropriate eviction policy

Application ​

  • [ ] Deploy key added to the repository, read-only
  • [ ] ssh -T git@github.com authenticates
  • [ ] Release-directory layout created with shared/
  • [ ] .env created at mode 600 in shared/
  • [ ] .env contents saved to a password manager
  • [ ] All secrets generated with openssl rand, ≥32 bytes
  • [ ] Dependencies installed with --frozen-lockfile
  • [ ] prisma generate and prisma migrate deploy run
  • [ ] Backend and frontend built successfully

Process management ​

  • [ ] ecosystem.config.cjs with cluster mode, ≥2 instances
  • [ ] Both apps bound to 127.0.0.1 — verified with ss -tulpn
  • [ ] max_memory_restart, kill_timeout, min_uptime, max_restarts set
  • [ ] watch: false
  • [ ] pm2 startup run and pm2 save executed
  • [ ] pm2-logrotate installed and configured

Web and TLS ​

  • [ ] Nginx installed; default site removed
  • [ ] server_tokens off; client_max_body_size set
  • [ ] Custom log format with $request_time and $upstream_response_time
  • [ ] Per-site access and error logs
  • [ ] Proxy params snippet with Host and X-Forwarded-Proto
  • [ ] WebSocket map and long timeouts on /socket.io/
  • [ ] Static /_nuxt/ served from disk with immutable caching
  • [ ] HTML not cached
  • [ ] Security headers included in every location
  • [ ] Rate limiting on general traffic and auth endpoints
  • [ ] sudo nginx -t passes
  • [ ] Certbot --dry-run succeeded before the real run
  • [ ] Certificate obtained; HTTP redirects to HTTPS
  • [ ] certbot renew --dry-run passes
  • [ ] Renewal deploy hook reloads Nginx
  • [ ] HSTS starting at a low max-age

Automation ​

  • [ ] Deploy script with set -euo pipefail and a health check
  • [ ] Rollback script exists
  • [ ] CI/CD deploy key created and restricted with command=
  • [ ] SSH_KNOWN_HOSTS populated — no StrictHostKeyChecking=no
  • [ ] Pipeline triggers on push to production, never pull_request
  • [ ] Pipeline runs lint and typecheck before deploying
  • [ ] Third-party actions pinned to commit SHAs

Operations ​

  • [ ] Full reboot tested — everything came back automatically
  • [ ] Backup script running daily with off-site encrypted copies
  • [ ] Backup script verifies the dump before declaring success
  • [ ] Decryption key stored off the server
  • [ ] Restore verification passing
  • [ ] Dead-man's switch on the backup job
  • [ ] Health check script on cron
  • [ ] External uptime monitoring configured
  • [ ] Error tracking with secret scrubbing

Before going to production ​

Run this before the first real user arrives.

Security ​

  • [ ] sudo ss -tulpn | grep -vE "127.0.0.1|::1" shows only 22, 80, 443
  • [ ] From outside: nc -vz <ip> 5432, 6379, 3000, 3001 all fail
  • [ ] SSL Labs score of A or better
  • [ ] Security headers verified in the response
  • [ ] .env is mode 600
  • [ ] No secrets in Git history — verified with gitleaks
  • [ ] No secrets in the client bundle — grep -ri "secret\|sk_live" .output/public/
  • [ ] Nothing sensitive in runtimeConfig.public
  • [ ] Passwords hashed with argon2id or bcrypt cost ≥ 12
  • [ ] Every record access scoped to the requesting user (IDOR audit)
  • [ ] ValidationPipe with whitelist and forbidNonWhitelisted
  • [ ] CORS restricted to known origins, not *
  • [ ] Rate limiting on login, registration, and password reset
  • [ ] pnpm audit --audit-level=high clean, or exceptions documented
  • [ ] 2FA on VPS provider, registrar, DNS provider, GitHub/GitLab
  • [ ] Security audit script run and reviewed

Reliability ​

  • [ ] Reboot test passed
  • [ ] pm2 reload tested with a slow in-flight request — it completed
  • [ ] Rollback script executed as a drill
  • [ ] Health endpoint returns the deployed commit SHA
  • [ ] Graceful shutdown implemented and verified
  • [ ] All shared state in Redis or PostgreSQL, never process memory
  • [ ] Socket.IO Redis adapter configured (required with cluster mode)
  • [ ] No setInterval cron jobs running in cluster mode
  • [ ] connection_limit sized against max_connections × instances
  • [ ] Crash-loop protection configured

Performance ​

  • [ ] gzip enabled with a sensible type list
  • [ ] Static assets served by Nginx from disk
  • [ ] Hashed assets cached for a year; HTML not cached
  • [ ] Database indexes on all foreign keys and common query filters
  • [ ] log_min_duration_statement = 1000 so slow queries are visible
  • [ ] Cache hit ratio above 0.99
  • [ ] Nuxt calls the API over loopback during SSR
  • [ ] Resource usage measured under a realistic load
  • [ ] Load test run against a staging copy

Data ​

  • [ ] Backups running and restore-verified
  • [ ] Off-site copy confirmed present
  • [ ] .env included in the backup
  • [ ] User uploads backed up
  • [ ] RPO and RTO defined and achievable
  • [ ] Full manual restore drill completed on a fresh server — date recorded
  • [ ] Backup storage credentials are write-only; object lock enabled
  • [ ] Provider snapshot taken

Observability ​

  • [ ] External uptime monitoring on both hostnames
  • [ ] TLS expiry monitored independently
  • [ ] Alerts fire on two consecutive failures
  • [ ] Alerts on traffic dropping to zero, not just spikes
  • [ ] Error tracking with release set to the commit SHA
  • [ ] Structured logging with request IDs
  • [ ] Log rotation configured for every producer
  • [ ] Disk usage alerting at 85%
  • [ ] I know where every log lives
  • [ ] Privacy policy and terms published if handling personal data
  • [ ] Cookie consent if required in your jurisdiction
  • [ ] Data retention and deletion process defined
  • [ ] Incident response contact and procedure documented
  • [ ] Disaster recovery runbook stored off the server (Level 28)
  • [ ] Domain auto-renew enabled with a calendar backup reminder
  • [ ] Someone other than you can deploy and roll back

THE BUS-FACTOR QUESTION

If you were unreachable for two weeks and the site broke, could anyone else fix it?

They need: access to the server, access to the provider account, the disaster recovery runbook, and the backup decryption key. Store these somewhere a trusted colleague can reach — a shared password-manager vault, not your laptop.


After every deployment ​

Immediate (first 2 minutes) ​

  • [ ] Pipeline reported success
  • [ ] Health endpoint returns 200
  • [ ] Health endpoint shows the expected commit SHA
  • [ ] pm2 list shows all processes online with a fresh uptime
  • [ ] Restart counts have not increased
  • [ ] Frontend loads in a browser
  • [ ] Login works
  • [ ] One core user flow works end to end
bash
# LOCAL — post-deploy smoke test
curl -fsS -o /dev/null -w "app %{http_code}\n" https://app.example.com
curl -fsS https://api.example.com/api/health | jq

Short window (first 10 minutes) ​

  • [ ] Error rate has not increased
  • [ ] Response times are normal
  • [ ] No new error types in Sentry
  • [ ] No 5xx spike in Nginx logs
  • [ ] Database connection count is normal
  • [ ] Memory usage is stable, not climbing
bash
# SERVER — watch for 10 minutes
pm2 logs --err --lines 0
sudo tail -f /var/log/nginx/api.example.com.access.log | grep -E ' 5[0-9][0-9] '

STAY FOR TEN MINUTES

Most deploy-caused incidents surface within the first few minutes, on a code path your smoke test does not touch. Ten minutes of watching is the cheapest incident prevention available.

Rollback triggers ​

Roll back immediately, without debugging first, if:

  • [ ] Error rate above ~5% of requests
  • [ ] Any core user flow is broken
  • [ ] Response times more than doubled
  • [ ] The database is under unexpected load
  • [ ] Memory is climbing steadily toward the limit
bash
# SERVER
~/apps/myapp/rollback.sh

ROLL BACK FIRST, DIAGNOSE AFTERWARDS

The instinct is to debug while users are affected. Resist it. Restore service in five seconds, then investigate calmly with the logs you already have.

The exception: if the release included a destructive migration, rollback is not safe — the old code will fail against the new schema. That is precisely why migrations must be additive (Level 20).

Later that day ​

  • [ ] Backup ran successfully after the deploy
  • [ ] No new fail2ban bans that look related
  • [ ] Disk usage has not jumped
  • [ ] Any new errors triaged and ticketed

Weekly ​

  • [ ] Review error tracker for new or increasing issues
  • [ ] Check disk usage trend
  • [ ] Check memory trend
  • [ ] Review pnpm audit output
  • [ ] Check for pending security updates: apt list --upgradable | grep -i security
  • [ ] Confirm backups ran every day this week
  • [ ] Review slow query log
  • [ ] Check TLS expiry
  • [ ] Review fail2ban bans for anything unusual
bash
# SERVER — weekly review
df -h && free -h
tail -20 /home/deploy/logs/backup.log
sudo certbot certificates | grep -E "Certificate Name|Expiry"
apt list --upgradable 2>/dev/null | grep -i security
sudo fail2ban-client status sshd
sudo -u postgres psql -d myapp_production -c \
  "SELECT substring(query,1,60), calls, round(mean_exec_time::numeric,1) AS avg_ms
   FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 10;"

Monthly ​

  • [ ] Run the full security audit script (Level 23)
  • [ ] Verify a restore into a scratch database
  • [ ] Review authorized_keys on every account — remove anything unrecognised
  • [ ] Review sudo rules
  • [ ] Review who has access to the repository, CI secrets, and the provider account
  • [ ] Update dependencies (patch and minor)
  • [ ] Review and rotate any secret older than your policy
  • [ ] Check that log rotation is actually working
  • [ ] Review resource trends — do you need to resize?
  • [ ] Confirm the DR runbook is still accurate

Quarterly ​

  • [ ] Full restore drill on a fresh VPS — time it, record the RTO
  • [ ] Practise a rollback in production
  • [ ] Review and update the disaster recovery runbook
  • [ ] Rotate SSH deploy keys and CI secrets
  • [ ] Review the Node.js version against its EOL date
  • [ ] Review the Ubuntu LTS support window
  • [ ] Review PostgreSQL major version support
  • [ ] Load test against staging
  • [ ] Review the security checklist end to end
  • [ ] Confirm someone else could still take over

Annually ​

  • [ ] Plan the Ubuntu LTS upgrade if within a year of EOL
  • [ ] Plan the PostgreSQL major version upgrade
  • [ ] Rotate all long-lived secrets
  • [ ] Review whether the architecture still fits the load
  • [ ] Review costs against usage
  • [ ] Audit third-party integrations and remove unused ones
  • [ ] Renew the domain (or confirm auto-renew)

Next: Disaster Recovery Plan →