Level 26 — Troubleshooting Handbook
Every entry follows the same structure: Problem → Possible causes → Diagnostic commands → Solution → Prevention.
The general method
Debug by walking the request path and asking at each hop: did it get this far?
THE SINGLE MOST USEFUL DIAGNOSTIC
# SERVER
curl -i http://127.0.0.1:3001/api/healthWorks → the app is fine; the problem is Nginx, DNS, TLS, or the firewall. Fails → the problem is your application or its dependencies.
One command eliminates half the possibilities. Start here for almost anything.
SSH problems
SSH cannot connect — timeout
Causes: firewall blocking 22 · wrong IP · server off · your network blocks outbound 22
Diagnose
# LOCAL
ssh -v deploy@203.0.113.10
nc -vz 203.0.113.10 22
ping -c 3 203.0.113.10No Connection established in -v output → network/firewall, not keys.
Solution — via the provider's web console:
sudo ufw allow 22/tcp && sudo ufw reload
sudo systemctl status sshAlso check the cloud firewall in the provider UI. If your network blocks outbound 22, test from a phone hotspot.
Prevent: cloud firewall as an independent layer; know where the web console is.
SSH — connection refused
Different from a timeout: the packet arrived and was rejected. Network is fine.
Causes: sshd not running · running on a different port · failed to start after a config edit
Diagnose (via web console)
sudo systemctl status ssh
sudo ss -tulpn | grep sshd
sudo sshd -t
sudo journalctl -u ssh -n 50Solution
sudo sshd -t # find the syntax error
sudo systemctl restart sshPrevent: always sudo sshd -t before reloading; use reload not restart; keep a session open.
Permission denied (publickey)
Causes: key not in authorized_keys · wrong username · wrong permissions · agent offering too many keys · blocked by AllowUsers · banned by fail2ban
Diagnose
# LOCAL
ssh -v deploy@203.0.113.10 2>&1 | grep -E "Offering|Authentications|debug1: Next"
ssh-add -l
ssh-keygen -lf ~/.ssh/id_ed25519.pub# SERVER (web console)
sudo journalctl -u ssh -f # watch while you retry from another terminal
ls -la /home/deploy /home/deploy/.ssh
ssh-keygen -lf /home/deploy/.ssh/authorized_keys
sudo sshd -T | grep -Ei "allowusers|pubkey"
sudo fail2ban-client status sshdSolution
# SERVER — the most common cause is permissions
chmod 755 /home/deploy
chmod 700 /home/deploy/.ssh
chmod 600 /home/deploy/.ssh/authorized_keys
chown -R deploy:deploy /home/deploy/.ssh# LOCAL — if too many keys are offered
ssh -o IdentitiesOnly=yes -i ~/.ssh/id_ed25519 deploy@203.0.113.10THE SERVER LOG IS ALWAYS MORE SPECIFIC THAN THE CLIENT ERROR
sudo journalctl -u ssh -f says things like Authentication refused: bad ownership or modes for directory /home/deploy — precisely the answer, which the client never sees.
Prevent: IdentitiesOnly yes in ~/.ssh/config; your IP in fail2ban's ignoreip.
Root login disabled but I need root
Solution
ssh deploy@203.0.113.10
sudo -iIf deploy was never given sudo: use the web console, log in as root with the console password, then usermod -aG sudo deploy.
Prevent: test sudo whoami as deploy before disabling root login.
UFW blocking SSH
Diagnose (web console)
sudo ufw status numbered
sudo grep 'UFW BLOCK' /var/log/syslog | grep "DPT=22" | tailSolution
sudo ufw allow 22/tcp
sudo ufw reloadPrevent: always sudo ufw allow OpenSSH and verify with ufw show added before ufw enable.
Host key verification failed
Causes: server rebuilt · IP reassigned · genuine MITM
Solution — only after confirming you rebuilt the server:
# LOCAL
ssh-keygen -R 203.0.113.10Prevent: never delete all of known_hosts; remove only the specific host.
Git problems
Git permission denied
Diagnose
# SERVER
ssh -T git@github.com
GIT_SSH_COMMAND="ssh -v" git fetch origin
cat ~/.ssh/configSolution
# SERVER
cat >> ~/.ssh/config <<'EOF'
Host github.com
HostName github.com
User git
IdentityFile ~/.ssh/github_deploy
IdentitiesOnly yes
EOF
chmod 600 ~/.ssh/config ~/.ssh/github_deployConfirm the deploy key is registered on the repository (Settings → Deploy keys), not just your account.
Prevent: one deploy key per repository, with a distinct Host alias.
Git dubious ownership
fatal: detected dubious ownership in repository at '/home/deploy/apps/myapp'Cause: the repository is owned by a different user than the one running Git — almost always because something was once run with sudo.
Diagnose
ls -la /home/deploy/apps/myapp/.git
find /home/deploy/apps/myapp ! -user deploy | headSolution
sudo chown -R deploy:deploy /home/deploy/apps/myappDO NOT USE git config --global --add safe.directory
It silences the warning without fixing the ownership, and the same mismatch will then cause EACCES failures in pnpm install and your build. Never use safe.directory '*' — that disables the protection globally.
Prevent: never run git, pnpm, or node with sudo.
Git: could not read Username for 'https://github.com'
Cause: an HTTPS remote in a non-interactive shell with no credentials.
Solution
# SERVER
git remote set-url origin git@github.com:myorg/myapp.git
git remote -vNode / tooling problems
node / pnpm / pm2: command not found
THIS IS THE #1 CI/CD FAILURE
# LOCAL — reproduce it exactly
ssh deploy@203.0.113.10 "which node pnpm pm2" # fails
ssh deploy@203.0.113.10 # then: which node → worksThe difference: ssh host "cmd" is a non-interactive shell that exits ~/.bashrc before reaching NVM.
Solution — do both:
# SERVER — 1. move NVM loading to the TOP of ~/.bashrc, above the interactive guard
nano ~/.bashrcexport NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && \. "$NVM_DIR/nvm.sh"# SERVER — 2. symlink into a system PATH
for b in node npm npx pnpm pm2; do sudo ln -sf "$(command -v $b)" "/usr/local/bin/$b"; doneVerify
# LOCAL
ssh deploy@203.0.113.10 "node -v && pnpm -v && pm2 -v"Prevent: re-run the symlink loop after every Node upgrade; test this before writing any CI workflow.
Build fails with Killed
Cause: the OOM killer. Not a build error.
Diagnose
sudo dmesg -T | grep -i "killed process"
free -hSolution
# Add swap
sudo fallocate -l 2G /swapfile && sudo chmod 600 /swapfile
sudo mkswap /swapfile && sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
# Or raise the heap limit
NODE_OPTIONS="--max-old-space-size=2048" pnpm buildPrevent: build in CI and ship the artifact (Level 18).
ERR_PNPM_OUTDATED_LOCKFILE
Solution
# LOCAL
pnpm install
git add pnpm-lock.yaml && git commit -m "update lockfile" && git pushPrevent: pin packageManager in package.json so everyone uses the same pnpm version.
Database problems
PostgreSQL connection refused
Diagnose
sudo systemctl status postgresql
sudo ss -tulpn | grep 5432
sudo -u postgres pg_isready
sudo tail -50 /var/log/postgresql/postgresql-16-main.logSolution
sudo systemctl start postgresqlIf it will not start, the log gives the reason — usually a bad config change or a full disk.
ECONNREFUSED ::1:5432 IS DIFFERENT
That is IPv6. Your DATABASE_URL says localhost, which resolved to ::1, and PostgreSQL is only listening on IPv4.
Fix: use 127.0.0.1 instead of localhost in every connection string.
Prevent: always 127.0.0.1; monitor disk usage.
PostgreSQL authentication failed
Diagnose
sudo tail -f /var/log/postgresql/postgresql-16-main.log # while retrying
sudo grep -vE '^\s*#|^\s*$' /etc/postgresql/16/main/pg_hba.conf
sudo -u postgres psql -c "SELECT rolname, substring(rolpassword,1,14) FROM pg_authid WHERE rolname='myapp';"Solution
sudo -u postgres psql -c "ALTER USER myapp WITH PASSWORD 'new-password';"
nano ~/apps/myapp/shared/.env
pm2 reload ecosystem.config.cjs --update-env # ⚠️ --update-env is requiredIF THE PASSWORD CONTAINS @ : / ? # &
The connection string parser breaks. URL-encode it (@ → %40) or regenerate without those characters.
md5 → scram-sha-256 MIGRATION
Changing the method in pg_hba.conf is not enough — existing hashes are still MD5. Set password_encryption = 'scram-sha-256', restart, then reset each password.
PostgreSQL: too many connections
Diagnose
SELECT count(*), state FROM pg_stat_activity GROUP BY state;
SHOW max_connections;Solution
DATABASE_URL="postgresql://...?connection_limit=10"Formula: connection_limit ≤ (max_connections − 10) / number_of_PM2_instances.
Prevent: set connection_limit explicitly; remember Prisma's default is per-process.
Prisma: "did not initialize yet"
Solution
pnpm --filter backend prisma generatePrevent: make it an explicit step in the deploy script — never rely on the postinstall hook.
Migration hangs
Cause: blocked waiting on a lock held by a long-running query.
Diagnose
SELECT pid, now()-query_start AS dur, state, left(query,80)
FROM pg_stat_activity WHERE state != 'idle' ORDER BY dur DESC;
SELECT * FROM pg_locks WHERE NOT granted;Solution
SELECT pg_cancel_backend(12345); -- polite
SELECT pg_terminate_backend(12345); -- forcefulPrevent: SET lock_timeout = '5s' in migrations; CREATE INDEX CONCURRENTLY.
Redis connection refused
Diagnose
sudo systemctl status redis-server
sudo ss -tulpn | grep 6379
redis-cli ping
sudo tail -50 /var/log/redis/redis-server.logSolution: sudo systemctl start redis-server
| Error | Fix |
|---|---|
NOAUTH Authentication required | Add the password to REDIS_URL — note the leading colon: redis://:pass@... |
WRONGPASS | Password mismatch; check requirepass and reload PM2 with --update-env |
MISCONF ... RDB snapshots | Background save failing — disk full or /var/lib/redis permissions |
OOM command not allowed | maxmemory reached with noeviction — raise it or add TTLs |
Docker problems
Docker permission denied
permission denied while trying to connect to the Docker daemon socketSolution
sudo usermod -aG docker $USER
newgrp docker # or log out and back in
docker psPrevent: remember group changes need a new login. And that this group is equivalent to root.
Port is already allocated
Diagnose
sudo ss -tulpn | grep :5432
docker ps -a --format "table {{.Names}}\t{{.Ports}}\t{{.Status}}"
systemctl is-active postgresqlSolution: stop the conflicting service, or change the mapping. A very common cause is a native PostgreSQL running alongside a Docker one.
Container exits immediately
Diagnose
docker logs <container>
docker inspect <container> | grep -A5 StateThe logs contain the answer nearly every time.
Data lost after docker compose down
Cause: -v was used, or no named volume was defined.
docker compose down -v DELETES NAMED VOLUMES WITH NO CONFIRMATION
Your database is gone. Restore from backup.
Prevent: never type -v after down on a machine with production data. Use named volumes, not bind mounts, for databases.
Nginx problems
nginx: configuration file test failed
Diagnose
sudo nginx -t # prints the file and line number
sudo nginx -T | less # full resolved configSolution: fix the reported line. Usually a missing semicolon or unclosed brace.
Prevent: sudo nginx -t && sudo systemctl reload nginx as one command, always.
502 Bad Gateway
Causes: app not running · wrong port in proxy_pass · app crashed · app bound to the wrong address
Diagnose — in this order
sudo tail -20 /var/log/nginx/error.log # the specific reason
pm2 list # is it running?
sudo ss -tulpn | grep -E '3000|3001' # listening? on what address?
curl -i http://127.0.0.1:3001/api/health # ⭐ the key test
pm2 logs api --err --lines 50 --nostream # why did it die?| Error log line | Cause | Fix |
|---|---|---|
connect() failed (111: Connection refused) | App not running / wrong port | Start it; check proxy_pass |
no live upstreams | All marked failed by max_fails | Fix the app; they recover after fail_timeout |
upstream prematurely closed connection | App crashed mid-request | Read app logs for the exception |
(13: Permission denied) | Socket/SELinux permissions | journalctl -u nginx |
Solution: if curl http://127.0.0.1:3001/health works, fix Nginx's proxy_pass. If it fails, fix the app.
Prevent: health check in the deploy pipeline; wait_ready in PM2.
504 Gateway Timeout
Cause: the app took longer than proxy_read_timeout (60s).
Diagnose
pm2 monit
sudo -u postgres psql -c "SELECT pid, now()-query_start AS d, left(query,80) FROM pg_stat_activity WHERE state!='idle' ORDER BY d DESC;"
sudo awk '{print $(NF-1), $7}' /var/log/nginx/access.log | sort -rn | headRAISING proxy_read_timeout IS NOT A FIX
It converts a 504 into a worker blocked for five minutes. Under load every worker ends up blocked and the whole site stops. Fix the slow query, or move the work to a background queue.
403 Forbidden on static files
Diagnose
sudo tail /var/log/nginx/error.log
namei -l /home/deploy/apps/myapp/current/frontend/.output/public/index.html
sudo -u www-data stat /home/deploy/apps/myapp/current/frontend/.output/public/namei -l shows permissions for every path component — instantly revealing which directory blocks www-data.
Solution
chmod 755 /home/deploy
chmod 755 /home/deploy/appsPrevent: /home/deploy must stay 755 so Nginx can traverse it.
413 Request Entity Too Large
Solution
client_max_body_size 20M;Default is 1M. Set it in http, server, or the specific location.
Wrong site served / default page shown
Diagnose
sudo nginx -T | grep -E "server_name|listen"
ls -la /etc/nginx/sites-enabled/
curl -H "Host: app.example.com" http://127.0.0.1 -ISolution: remove sites-enabled/default; verify the symlink exists; reload.
TLS problems
Certbot challenge failed
Diagnose
# LOCAL
dig app.example.com +short
curl -I http://app.example.com# SERVER
sudo mkdir -p /var/www/certbot/.well-known/acme-challenge
echo test | sudo tee /var/www/certbot/.well-known/acme-challenge/test
curl http://app.example.com/.well-known/acme-challenge/testIf the last command does not print test, your HTTPS redirect is intercepting the challenge path.
Solution
server {
listen 80;
server_name app.example.com;
location /.well-known/acme-challenge/ { root /var/www/certbot; }
location / { return 301 https://$host$request_uri; }
}RATE LIMITS
5 failed validations per hostname per hour. Use --dry-run while debugging.
Certificate expired despite the renewal timer
Diagnose
systemctl status certbot.timer
sudo journalctl -u certbot --since "60 days ago" | grep -Ei "error|fail"
sudo certbot renew --dry-run
sudo certbot certificates| Cause | Fix |
|---|---|
| Port 80 closed after initial issuance | sudo ufw allow 80/tcp |
| Nginx config lost the challenge location | Restore it |
| Missing reload hook — Nginx serves a stale cert from memory | Add renewal-hooks/deploy/reload-nginx.sh |
| DNS changed | Update the A record |
Prevent: run certbot renew --dry-run after every Nginx or firewall change; monitor expiry externally.
ERR_SSL_PROTOCOL_ERROR
Diagnose
sudo ss -tulpn | grep 443
sudo nginx -t
sudo openssl x509 -noout -modulus -in /etc/letsencrypt/live/app.example.com/fullchain.pem | openssl md5
sudo openssl rsa -noout -modulus -in /etc/letsencrypt/live/app.example.com/privkey.pem | openssl md5The two hashes must match. If not, the certificate and key are from different issuances.
Redirect loop
Causes: Cloudflare SSL mode set to Flexible · app forcing HTTPS without seeing X-Forwarded-Proto
Solution: set Cloudflare to Full (Strict); ensure proxy_set_header X-Forwarded-Proto $scheme and trust proxy in NestJS.
DNS problems
DNS not resolving
Diagnose
dig NS example.com +short # who is authoritative?
dig @ns1.provider.com app.example.com +short # correct at the source?
dig @1.1.1.1 app.example.com +short # propagated?
dig app.example.com +short # what do I get?
cat /etc/hosts # stale test entry?| Finding | Cause |
|---|---|
| Authoritative is wrong | Record wrong — fix it in the DNS panel |
| Authoritative right, public resolver wrong | Still propagating — wait out the TTL |
| Both right, yours wrong | Local cache or /etc/hosts |
| Edits have no effect | You are editing at the wrong provider — check the NS records |
Prevent: TTL 300 during changes; verify DNS before running Certbot.
Application problems
WebSocket not connecting
Diagnose
# LOCAL
curl -i -N -H "Connection: Upgrade" -H "Upgrade: websocket" \
-H "Sec-WebSocket-Version: 13" -H "Sec-WebSocket-Key: x3JJHMbDL1EzLkh9GBhXDw==" \
https://api.example.com/socket.io/?EIO=4\&transport=websocket# SERVER
sudo grep -i "socket.io" /var/log/nginx/api.example.com.access.log | tail
pm2 list # more than one instance?| Symptom | Cause | Fix |
|---|---|---|
| Handshake 400 | Missing proxy_http_version 1.1 or upgrade headers | Add both to the location |
| Disconnects every 60 seconds | proxy_read_timeout default | Set 7d on /socket.io/ |
Session ID unknown | No sticky sessions with cluster mode | Force transports: ['websocket'], or add sticky sessions |
| Events reach only some clients | No Redis adapter in cluster mode | Add @socket.io/redis-adapter |
| Messages arrive in bursts | proxy_buffering on | Set proxy_buffering off |
PM2 process crash-looping
Diagnose
pm2 list # restart count climbing?
pm2 describe api
pm2 logs api --err --lines 100 --nostream
pm2 env 0 | grep -E "NODE_ENV|DATABASE_URL"| Cause | Check |
|---|---|
| Missing env var | Zod validation error in the logs |
| DB unreachable | psql from the same host |
| Port in use | sudo ss -tulpn | grep 3001 |
| Build missing | ls backend/dist/main.js |
Solution: fix the cause, then pm2 reset api && pm2 restart api.
Prevent: Zod env validation (Level 8); min_uptime and max_restarts.
Out of memory
Diagnose
free -h
sudo dmesg -T | grep -i "killed process"
ps aux --sort=-%mem | head -10
pm2 monitSolution: add swap · reduce PM2 instances · lower max_memory_restart · reduce PostgreSQL shared_buffers · resize the server.
Prevent: OOMScoreAdjust=-500 on PostgreSQL; build in CI; alert at 90% memory.
Disk full
AT 100%, EVERYTHING BREAKS AT ONCE
PostgreSQL refuses writes, Nginx cannot log, builds fail, and SSH login may fail.
Diagnose
df -h && df -i
du -h --max-depth=1 / 2>/dev/null | sort -rh | head -20
docker system df
journalctl --disk-usage
du -sh ~/.pm2/logs ~/apps/myapp/releasesEmergency triage
sudo journalctl --vacuum-size=200M
pm2 flush
docker image prune -af # ⚠️ never --volumes
sudo apt clean
sudo find /var/log -name "*.gz" -mtime +7 -delete
ls -1dt ~/apps/myapp/releases/* | tail -n +4 | xargs -r rm -rfPrevent: log rotation everywhere; release pruning; alert at 85%.
CI/CD problems
CI/CD SSH failure
Diagnose
- run: ssh -vvv -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" "echo ok"| Error | Cause | Fix |
|---|---|---|
missing server host | Empty or misnamed secret | Check the name; check environment scoping |
Permission denied (publickey) | Key not installed / wrong user | Verify authorized_keys on the server |
error in libcrypto | Malformed private key secret | Re-copy including BEGIN/END and trailing newline |
Host key verification failed | known_hosts missing | ssh-keyscan into a secret |
command not found | Non-interactive shell PATH | Source NVM; symlink binaries |
GitHub Actions secrets missing
Diagnose
- run: |
[ -n "${{ secrets.SSH_HOST }}" ] || { echo "SSH_HOST is empty"; exit 1; }Causes: typo (names are case-sensitive) · created in the wrong repository · it is an environment secret but the job declares no environment: · the workflow is triggered by pull_request from a fork (secrets are deliberately withheld)
GitLab Runner failure
| Symptom | Cause | Fix |
|---|---|---|
| Pipeline stuck "pending" | No runner with matching tags, or offline | Settings → CI/CD → Runners |
| Secret empty | Variable is Protected, branch is not | Protect the branch |
| Services unreachable | Wrong hostname | Use the service alias, not 127.0.0.1 |
error in libcrypto | Key stored as a plain variable | Use a File-type variable |
Deployment succeeds but the application does not update
THE MOST CONFUSING FAILURE — DIAGNOSE IN THIS EXACT ORDER
# SERVER
cd ~/apps/myapp/current
git log -1 --oneline # 1. did the code update?
find backend/src -newer backend/dist/main.js -name "*.ts" | head # 2. did the build run?
pm2 list # 3. did PM2 restart? (check uptime)
pm2 env 0 | grep GIT_COMMIT # 4. new env loaded?
curl -s http://127.0.0.1:3001/api/health | jq .commit # 5. what does the app say?
readlink -f ~/apps/myapp/current # 6. did the symlink move?| Finding | Cause | Fix |
|---|---|---|
| Old commit at step 1 | Wrong branch, or fetch failed | Check the branch; check the deploy key |
| Source newer than dist at step 2 | Build did not run or failed silently | Add set -e to the deploy script |
| PM2 uptime in hours at step 3 | pm2 reload did not run | Check the script's exit path |
| Old SHA at step 4 | Missing --update-env | pm2 reload ... --update-env |
| Symlink unchanged at step 6 | The atomic switch failed | Check the mv -Tf step |
| All correct but the browser shows old | Browser or CDN cache | Hard refresh; check Cache-Control on HTML |
THE ROOT CAUSE IS USUALLY A MISSING set -e
Without it, a failed pnpm build does not stop the script. It proceeds to pm2 reload, which restarts the app on the old build, and the pipeline reports success. Green tick, unchanged site.
Every deploy script starts with set -euo pipefail.
Quick reference — first commands by symptom
| Symptom | First command |
|---|---|
| Site completely down | pm2 list && systemctl is-active nginx postgresql redis-server |
| 502 Bad Gateway | curl -i http://127.0.0.1:3001/api/health |
| 504 Gateway Timeout | pm2 monit then check pg_stat_activity |
| Slow site | uptime && free -h && df -h |
| SSH fails | ssh -v — look for Connection established |
| Deploy failed | pm2 logs --err --lines 100 --nostream |
| Certificate error | sudo certbot certificates |
| Database error | sudo tail -50 /var/log/postgresql/*.log |
| Something disappeared | sudo dmesg -T | grep -i "killed process" |
| Weird intermittent errors | df -h (check for a full disk) |
| Anything after a deploy | find src -newer dist/main.js |