Level 27 — Production Checklists
Copy these. Work through them literally. The value is in not trusting your memory.
New VPS setup
Follow Level 25 for the commands; this is the tracking list.
Before touching the server
- [ ] SSH key pair generated with a passphrase
- [ ] VPS created with Ubuntu 24.04 LTS, public key added at creation time
- [ ] Cloud firewall: inbound 22, 80, 443 only
- [ ] DNS A records created for
@,app,api,wwwwith TTL 300 - [ ] CAA record allowing
letsencrypt.org - [ ] Provider web console located — you know where the emergency access button is
System
- [ ]
apt update && apt upgraderun; rebooted if the kernel changed - [ ] Hostname set;
/etc/hostsentry added - [ ] Timezone set to UTC; time sync confirmed
- [ ]
tmux,htop,curl,git,jq,ncduinstalled - [ ] Swap configured if RAM ≤ 4 GB, with
vm.swappiness=10 - [ ] journald storage set to persistent with a size cap
Users and access
- [ ]
deployuser created; password saved in a password manager - [ ]
deployadded tosudowithusermod -aG - [ ] SSH key installed for
deploy;700/600permissions verified - [ ]
ssh deploy@serverandsudo whoamitested in a second terminal - [ ] Root SSH login disabled
- [ ] Password authentication disabled — verified by attempting one
- [ ]
AllowUsersrestricts SSH - [ ] Config verified with
sudo sshd -T, not by reading the file
Firewall and intrusion prevention
- [ ] UFW default deny incoming, allow outgoing
- [ ] SSH allowed before
ufw enable - [ ] Only 22, 80, 443 allowed; IPv4 and IPv6 rules present
- [ ] fail2ban installed;
sshdjail confirmed active withfail2ban-client status sshd - [ ] My own IP in
ignoreip - [ ]
unattended-upgradesenabled and dry-run verified
Runtime
- [ ] NVM installed as
deploy, not root - [ ] Node LTS installed;
nvm alias defaultset - [ ] pnpm via Corepack; version pinned in
packageManager - [ ] PM2 installed globally without sudo
- [ ]
ssh deploy@server "node -v && pnpm -v && pm2 -v"works from my laptop - [ ] Binaries symlinked into
/usr/local/bin - [ ]
.nvmrcmatches CI and local development
Data services
- [ ] PostgreSQL installed and enabled on boot
- [ ]
sudo ss -tulpn | grep 5432shows127.0.0.1 - [ ] Application role created — not
postgres - [ ] Strong, URL-safe password generated with
openssl rand - [ ]
ALTER DEFAULT PRIVILEGESset - [ ] Memory tuned for the server's RAM
- [ ]
pg_stat_statementsenabled - [ ]
OOMScoreAdjust=-500set on PostgreSQL - [ ]
pg_hba.confusesscram-sha-256; notrust - [ ] Redis installed with
bind 127.0.0.1andrequirepass - [ ]
redis-cli pingwithout auth returnsNOAUTH - [ ] Dangerous Redis commands renamed or disabled
- [ ]
maxmemoryset with an appropriate eviction policy
Application
- [ ] Deploy key added to the repository, read-only
- [ ]
ssh -T git@github.comauthenticates - [ ] Release-directory layout created with
shared/ - [ ]
.envcreated at mode600inshared/ - [ ]
.envcontents saved to a password manager - [ ] All secrets generated with
openssl rand, ≥32 bytes - [ ] Dependencies installed with
--frozen-lockfile - [ ]
prisma generateandprisma migrate deployrun - [ ] Backend and frontend built successfully
Process management
- [ ]
ecosystem.config.cjswith cluster mode, ≥2 instances - [ ] Both apps bound to
127.0.0.1— verified withss -tulpn - [ ]
max_memory_restart,kill_timeout,min_uptime,max_restartsset - [ ]
watch: false - [ ]
pm2 startuprun andpm2 saveexecuted - [ ]
pm2-logrotateinstalled and configured
Web and TLS
- [ ] Nginx installed; default site removed
- [ ]
server_tokens off;client_max_body_sizeset - [ ] Custom log format with
$request_timeand$upstream_response_time - [ ] Per-site access and error logs
- [ ] Proxy params snippet with
HostandX-Forwarded-Proto - [ ] WebSocket
mapand long timeouts on/socket.io/ - [ ] Static
/_nuxt/served from disk with immutable caching - [ ] HTML not cached
- [ ] Security headers included in every location
- [ ] Rate limiting on general traffic and auth endpoints
- [ ]
sudo nginx -tpasses - [ ] Certbot
--dry-runsucceeded before the real run - [ ] Certificate obtained; HTTP redirects to HTTPS
- [ ]
certbot renew --dry-runpasses - [ ] Renewal deploy hook reloads Nginx
- [ ] HSTS starting at a low
max-age
Automation
- [ ] Deploy script with
set -euo pipefailand a health check - [ ] Rollback script exists
- [ ] CI/CD deploy key created and restricted with
command= - [ ]
SSH_KNOWN_HOSTSpopulated — noStrictHostKeyChecking=no - [ ] Pipeline triggers on
pushtoproduction, neverpull_request - [ ] Pipeline runs lint and typecheck before deploying
- [ ] Third-party actions pinned to commit SHAs
Operations
- [ ] Full reboot tested — everything came back automatically
- [ ] Backup script running daily with off-site encrypted copies
- [ ] Backup script verifies the dump before declaring success
- [ ] Decryption key stored off the server
- [ ] Restore verification passing
- [ ] Dead-man's switch on the backup job
- [ ] Health check script on cron
- [ ] External uptime monitoring configured
- [ ] Error tracking with secret scrubbing
Before going to production
Run this before the first real user arrives.
Security
- [ ]
sudo ss -tulpn | grep -vE "127.0.0.1|::1"shows only 22, 80, 443 - [ ] From outside:
nc -vz <ip> 5432,6379,3000,3001all fail - [ ] SSL Labs score of A or better
- [ ] Security headers verified in the response
- [ ]
.envis mode600 - [ ] No secrets in Git history — verified with
gitleaks - [ ] No secrets in the client bundle —
grep -ri "secret\|sk_live" .output/public/ - [ ] Nothing sensitive in
runtimeConfig.public - [ ] Passwords hashed with argon2id or bcrypt cost ≥ 12
- [ ] Every record access scoped to the requesting user (IDOR audit)
- [ ]
ValidationPipewithwhitelistandforbidNonWhitelisted - [ ] CORS restricted to known origins, not
* - [ ] Rate limiting on login, registration, and password reset
- [ ]
pnpm audit --audit-level=highclean, or exceptions documented - [ ] 2FA on VPS provider, registrar, DNS provider, GitHub/GitLab
- [ ] Security audit script run and reviewed
Reliability
- [ ] Reboot test passed
- [ ]
pm2 reloadtested with a slow in-flight request — it completed - [ ] Rollback script executed as a drill
- [ ] Health endpoint returns the deployed commit SHA
- [ ] Graceful shutdown implemented and verified
- [ ] All shared state in Redis or PostgreSQL, never process memory
- [ ] Socket.IO Redis adapter configured (required with cluster mode)
- [ ] No
setIntervalcron jobs running in cluster mode - [ ]
connection_limitsized againstmax_connections× instances - [ ] Crash-loop protection configured
Performance
- [ ] gzip enabled with a sensible type list
- [ ] Static assets served by Nginx from disk
- [ ] Hashed assets cached for a year; HTML not cached
- [ ] Database indexes on all foreign keys and common query filters
- [ ]
log_min_duration_statement = 1000so slow queries are visible - [ ] Cache hit ratio above 0.99
- [ ] Nuxt calls the API over loopback during SSR
- [ ] Resource usage measured under a realistic load
- [ ] Load test run against a staging copy
Data
- [ ] Backups running and restore-verified
- [ ] Off-site copy confirmed present
- [ ]
.envincluded in the backup - [ ] User uploads backed up
- [ ] RPO and RTO defined and achievable
- [ ] Full manual restore drill completed on a fresh server — date recorded
- [ ] Backup storage credentials are write-only; object lock enabled
- [ ] Provider snapshot taken
Observability
- [ ] External uptime monitoring on both hostnames
- [ ] TLS expiry monitored independently
- [ ] Alerts fire on two consecutive failures
- [ ] Alerts on traffic dropping to zero, not just spikes
- [ ] Error tracking with
releaseset to the commit SHA - [ ] Structured logging with request IDs
- [ ] Log rotation configured for every producer
- [ ] Disk usage alerting at 85%
- [ ] I know where every log lives
Legal and operational
- [ ] Privacy policy and terms published if handling personal data
- [ ] Cookie consent if required in your jurisdiction
- [ ] Data retention and deletion process defined
- [ ] Incident response contact and procedure documented
- [ ] Disaster recovery runbook stored off the server (Level 28)
- [ ] Domain auto-renew enabled with a calendar backup reminder
- [ ] Someone other than you can deploy and roll back
THE BUS-FACTOR QUESTION
If you were unreachable for two weeks and the site broke, could anyone else fix it?
They need: access to the server, access to the provider account, the disaster recovery runbook, and the backup decryption key. Store these somewhere a trusted colleague can reach — a shared password-manager vault, not your laptop.
After every deployment
Immediate (first 2 minutes)
- [ ] Pipeline reported success
- [ ] Health endpoint returns 200
- [ ] Health endpoint shows the expected commit SHA
- [ ]
pm2 listshows all processesonlinewith a fresh uptime - [ ] Restart counts have not increased
- [ ] Frontend loads in a browser
- [ ] Login works
- [ ] One core user flow works end to end
# LOCAL — post-deploy smoke test
curl -fsS -o /dev/null -w "app %{http_code}\n" https://app.example.com
curl -fsS https://api.example.com/api/health | jqShort window (first 10 minutes)
- [ ] Error rate has not increased
- [ ] Response times are normal
- [ ] No new error types in Sentry
- [ ] No 5xx spike in Nginx logs
- [ ] Database connection count is normal
- [ ] Memory usage is stable, not climbing
# SERVER — watch for 10 minutes
pm2 logs --err --lines 0
sudo tail -f /var/log/nginx/api.example.com.access.log | grep -E ' 5[0-9][0-9] 'STAY FOR TEN MINUTES
Most deploy-caused incidents surface within the first few minutes, on a code path your smoke test does not touch. Ten minutes of watching is the cheapest incident prevention available.
Rollback triggers
Roll back immediately, without debugging first, if:
- [ ] Error rate above ~5% of requests
- [ ] Any core user flow is broken
- [ ] Response times more than doubled
- [ ] The database is under unexpected load
- [ ] Memory is climbing steadily toward the limit
# SERVER
~/apps/myapp/rollback.shROLL BACK FIRST, DIAGNOSE AFTERWARDS
The instinct is to debug while users are affected. Resist it. Restore service in five seconds, then investigate calmly with the logs you already have.
The exception: if the release included a destructive migration, rollback is not safe — the old code will fail against the new schema. That is precisely why migrations must be additive (Level 20).
Later that day
- [ ] Backup ran successfully after the deploy
- [ ] No new fail2ban bans that look related
- [ ] Disk usage has not jumped
- [ ] Any new errors triaged and ticketed
Weekly
- [ ] Review error tracker for new or increasing issues
- [ ] Check disk usage trend
- [ ] Check memory trend
- [ ] Review
pnpm auditoutput - [ ] Check for pending security updates:
apt list --upgradable | grep -i security - [ ] Confirm backups ran every day this week
- [ ] Review slow query log
- [ ] Check TLS expiry
- [ ] Review fail2ban bans for anything unusual
# SERVER — weekly review
df -h && free -h
tail -20 /home/deploy/logs/backup.log
sudo certbot certificates | grep -E "Certificate Name|Expiry"
apt list --upgradable 2>/dev/null | grep -i security
sudo fail2ban-client status sshd
sudo -u postgres psql -d myapp_production -c \
"SELECT substring(query,1,60), calls, round(mean_exec_time::numeric,1) AS avg_ms
FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 10;"Monthly
- [ ] Run the full security audit script (Level 23)
- [ ] Verify a restore into a scratch database
- [ ] Review
authorized_keyson every account — remove anything unrecognised - [ ] Review sudo rules
- [ ] Review who has access to the repository, CI secrets, and the provider account
- [ ] Update dependencies (patch and minor)
- [ ] Review and rotate any secret older than your policy
- [ ] Check that log rotation is actually working
- [ ] Review resource trends — do you need to resize?
- [ ] Confirm the DR runbook is still accurate
Quarterly
- [ ] Full restore drill on a fresh VPS — time it, record the RTO
- [ ] Practise a rollback in production
- [ ] Review and update the disaster recovery runbook
- [ ] Rotate SSH deploy keys and CI secrets
- [ ] Review the Node.js version against its EOL date
- [ ] Review the Ubuntu LTS support window
- [ ] Review PostgreSQL major version support
- [ ] Load test against staging
- [ ] Review the security checklist end to end
- [ ] Confirm someone else could still take over
Annually
- [ ] Plan the Ubuntu LTS upgrade if within a year of EOL
- [ ] Plan the PostgreSQL major version upgrade
- [ ] Rotate all long-lived secrets
- [ ] Review whether the architecture still fits the load
- [ ] Review costs against usage
- [ ] Audit third-party integrations and remove unused ones
- [ ] Renew the domain (or confirm auto-renew)
Next: Disaster Recovery Plan →