Skip to content

Level 26 — Troubleshooting Handbook ​

Every entry follows the same structure: Problem → Possible causes → Diagnostic commands → Solution → Prevention.

The general method ​

Debug by walking the request path and asking at each hop: did it get this far?

THE SINGLE MOST USEFUL DIAGNOSTIC

bash
# SERVER
curl -i http://127.0.0.1:3001/api/health

Works → the app is fine; the problem is Nginx, DNS, TLS, or the firewall. Fails → the problem is your application or its dependencies.

One command eliminates half the possibilities. Start here for almost anything.


SSH problems ​

SSH cannot connect — timeout ​

Causes: firewall blocking 22 · wrong IP · server off · your network blocks outbound 22

Diagnose

bash
# LOCAL
ssh -v deploy@203.0.113.10
nc -vz 203.0.113.10 22
ping -c 3 203.0.113.10

No Connection established in -v output → network/firewall, not keys.

Solution — via the provider's web console:

bash
sudo ufw allow 22/tcp && sudo ufw reload
sudo systemctl status ssh

Also check the cloud firewall in the provider UI. If your network blocks outbound 22, test from a phone hotspot.

Prevent: cloud firewall as an independent layer; know where the web console is.


SSH — connection refused ​

Different from a timeout: the packet arrived and was rejected. Network is fine.

Causes: sshd not running · running on a different port · failed to start after a config edit

Diagnose (via web console)

bash
sudo systemctl status ssh
sudo ss -tulpn | grep sshd
sudo sshd -t
sudo journalctl -u ssh -n 50

Solution

bash
sudo sshd -t                       # find the syntax error
sudo systemctl restart ssh

Prevent: always sudo sshd -t before reloading; use reload not restart; keep a session open.


Permission denied (publickey) ​

Causes: key not in authorized_keys · wrong username · wrong permissions · agent offering too many keys · blocked by AllowUsers · banned by fail2ban

Diagnose

bash
# LOCAL
ssh -v deploy@203.0.113.10 2>&1 | grep -E "Offering|Authentications|debug1: Next"
ssh-add -l
ssh-keygen -lf ~/.ssh/id_ed25519.pub
bash
# SERVER (web console)
sudo journalctl -u ssh -f          # watch while you retry from another terminal
ls -la /home/deploy /home/deploy/.ssh
ssh-keygen -lf /home/deploy/.ssh/authorized_keys
sudo sshd -T | grep -Ei "allowusers|pubkey"
sudo fail2ban-client status sshd

Solution

bash
# SERVER — the most common cause is permissions
chmod 755 /home/deploy
chmod 700 /home/deploy/.ssh
chmod 600 /home/deploy/.ssh/authorized_keys
chown -R deploy:deploy /home/deploy/.ssh
bash
# LOCAL — if too many keys are offered
ssh -o IdentitiesOnly=yes -i ~/.ssh/id_ed25519 deploy@203.0.113.10

THE SERVER LOG IS ALWAYS MORE SPECIFIC THAN THE CLIENT ERROR

sudo journalctl -u ssh -f says things like Authentication refused: bad ownership or modes for directory /home/deploy — precisely the answer, which the client never sees.

Prevent: IdentitiesOnly yes in ~/.ssh/config; your IP in fail2ban's ignoreip.


Root login disabled but I need root ​

Solution

bash
ssh deploy@203.0.113.10
sudo -i

If deploy was never given sudo: use the web console, log in as root with the console password, then usermod -aG sudo deploy.

Prevent: test sudo whoami as deploy before disabling root login.


UFW blocking SSH ​

Diagnose (web console)

bash
sudo ufw status numbered
sudo grep 'UFW BLOCK' /var/log/syslog | grep "DPT=22" | tail

Solution

bash
sudo ufw allow 22/tcp
sudo ufw reload

Prevent: always sudo ufw allow OpenSSH and verify with ufw show added before ufw enable.


Host key verification failed ​

Causes: server rebuilt · IP reassigned · genuine MITM

Solution — only after confirming you rebuilt the server:

bash
# LOCAL
ssh-keygen -R 203.0.113.10

Prevent: never delete all of known_hosts; remove only the specific host.


Git problems ​

Git permission denied ​

Diagnose

bash
# SERVER
ssh -T git@github.com
GIT_SSH_COMMAND="ssh -v" git fetch origin
cat ~/.ssh/config

Solution

bash
# SERVER
cat >> ~/.ssh/config <<'EOF'
Host github.com
    HostName github.com
    User git
    IdentityFile ~/.ssh/github_deploy
    IdentitiesOnly yes
EOF
chmod 600 ~/.ssh/config ~/.ssh/github_deploy

Confirm the deploy key is registered on the repository (Settings → Deploy keys), not just your account.

Prevent: one deploy key per repository, with a distinct Host alias.


Git dubious ownership ​

fatal: detected dubious ownership in repository at '/home/deploy/apps/myapp'

Cause: the repository is owned by a different user than the one running Git — almost always because something was once run with sudo.

Diagnose

bash
ls -la /home/deploy/apps/myapp/.git
find /home/deploy/apps/myapp ! -user deploy | head

Solution

bash
sudo chown -R deploy:deploy /home/deploy/apps/myapp

DO NOT USE git config --global --add safe.directory

It silences the warning without fixing the ownership, and the same mismatch will then cause EACCES failures in pnpm install and your build. Never use safe.directory '*' — that disables the protection globally.

Prevent: never run git, pnpm, or node with sudo.


Git: could not read Username for 'https://github.com' ​

Cause: an HTTPS remote in a non-interactive shell with no credentials.

Solution

bash
# SERVER
git remote set-url origin git@github.com:myorg/myapp.git
git remote -v

Node / tooling problems ​

node / pnpm / pm2: command not found ​

THIS IS THE #1 CI/CD FAILURE

bash
# LOCAL — reproduce it exactly
ssh deploy@203.0.113.10 "which node pnpm pm2"     # fails
ssh deploy@203.0.113.10                            # then: which node   → works

The difference: ssh host "cmd" is a non-interactive shell that exits ~/.bashrc before reaching NVM.

Solution — do both:

bash
# SERVER — 1. move NVM loading to the TOP of ~/.bashrc, above the interactive guard
nano ~/.bashrc
bash
export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && \. "$NVM_DIR/nvm.sh"
bash
# SERVER — 2. symlink into a system PATH
for b in node npm npx pnpm pm2; do sudo ln -sf "$(command -v $b)" "/usr/local/bin/$b"; done

Verify

bash
# LOCAL
ssh deploy@203.0.113.10 "node -v && pnpm -v && pm2 -v"

Prevent: re-run the symlink loop after every Node upgrade; test this before writing any CI workflow.


Build fails with Killed ​

Cause: the OOM killer. Not a build error.

Diagnose

bash
sudo dmesg -T | grep -i "killed process"
free -h

Solution

bash
# Add swap
sudo fallocate -l 2G /swapfile && sudo chmod 600 /swapfile
sudo mkswap /swapfile && sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

# Or raise the heap limit
NODE_OPTIONS="--max-old-space-size=2048" pnpm build

Prevent: build in CI and ship the artifact (Level 18).


ERR_PNPM_OUTDATED_LOCKFILE ​

Solution

bash
# LOCAL
pnpm install
git add pnpm-lock.yaml && git commit -m "update lockfile" && git push

Prevent: pin packageManager in package.json so everyone uses the same pnpm version.


Database problems ​

PostgreSQL connection refused ​

Diagnose

bash
sudo systemctl status postgresql
sudo ss -tulpn | grep 5432
sudo -u postgres pg_isready
sudo tail -50 /var/log/postgresql/postgresql-16-main.log

Solution

bash
sudo systemctl start postgresql

If it will not start, the log gives the reason — usually a bad config change or a full disk.

ECONNREFUSED ::1:5432 IS DIFFERENT

That is IPv6. Your DATABASE_URL says localhost, which resolved to ::1, and PostgreSQL is only listening on IPv4.

Fix: use 127.0.0.1 instead of localhost in every connection string.

Prevent: always 127.0.0.1; monitor disk usage.


PostgreSQL authentication failed ​

Diagnose

bash
sudo tail -f /var/log/postgresql/postgresql-16-main.log   # while retrying
sudo grep -vE '^\s*#|^\s*$' /etc/postgresql/16/main/pg_hba.conf
sudo -u postgres psql -c "SELECT rolname, substring(rolpassword,1,14) FROM pg_authid WHERE rolname='myapp';"

Solution

bash
sudo -u postgres psql -c "ALTER USER myapp WITH PASSWORD 'new-password';"
nano ~/apps/myapp/shared/.env
pm2 reload ecosystem.config.cjs --update-env    # ⚠️ --update-env is required

IF THE PASSWORD CONTAINS @ : / ? # &

The connection string parser breaks. URL-encode it (@ → %40) or regenerate without those characters.

md5 → scram-sha-256 MIGRATION

Changing the method in pg_hba.conf is not enough — existing hashes are still MD5. Set password_encryption = 'scram-sha-256', restart, then reset each password.


PostgreSQL: too many connections ​

Diagnose

sql
SELECT count(*), state FROM pg_stat_activity GROUP BY state;
SHOW max_connections;

Solution

DATABASE_URL="postgresql://...?connection_limit=10"

Formula: connection_limit ≤ (max_connections − 10) / number_of_PM2_instances.

Prevent: set connection_limit explicitly; remember Prisma's default is per-process.


Prisma: "did not initialize yet" ​

Solution

bash
pnpm --filter backend prisma generate

Prevent: make it an explicit step in the deploy script — never rely on the postinstall hook.


Migration hangs ​

Cause: blocked waiting on a lock held by a long-running query.

Diagnose

sql
SELECT pid, now()-query_start AS dur, state, left(query,80)
FROM pg_stat_activity WHERE state != 'idle' ORDER BY dur DESC;

SELECT * FROM pg_locks WHERE NOT granted;

Solution

sql
SELECT pg_cancel_backend(12345);       -- polite
SELECT pg_terminate_backend(12345);    -- forceful

Prevent: SET lock_timeout = '5s' in migrations; CREATE INDEX CONCURRENTLY.


Redis connection refused ​

Diagnose

bash
sudo systemctl status redis-server
sudo ss -tulpn | grep 6379
redis-cli ping
sudo tail -50 /var/log/redis/redis-server.log

Solution: sudo systemctl start redis-server

ErrorFix
NOAUTH Authentication requiredAdd the password to REDIS_URL — note the leading colon: redis://:pass@...
WRONGPASSPassword mismatch; check requirepass and reload PM2 with --update-env
MISCONF ... RDB snapshotsBackground save failing — disk full or /var/lib/redis permissions
OOM command not allowedmaxmemory reached with noeviction — raise it or add TTLs

Docker problems ​

Docker permission denied ​

permission denied while trying to connect to the Docker daemon socket

Solution

bash
sudo usermod -aG docker $USER
newgrp docker          # or log out and back in
docker ps

Prevent: remember group changes need a new login. And that this group is equivalent to root.


Port is already allocated ​

Diagnose

bash
sudo ss -tulpn | grep :5432
docker ps -a --format "table {{.Names}}\t{{.Ports}}\t{{.Status}}"
systemctl is-active postgresql

Solution: stop the conflicting service, or change the mapping. A very common cause is a native PostgreSQL running alongside a Docker one.


Container exits immediately ​

Diagnose

bash
docker logs <container>
docker inspect <container> | grep -A5 State

The logs contain the answer nearly every time.


Data lost after docker compose down ​

Cause: -v was used, or no named volume was defined.

docker compose down -v DELETES NAMED VOLUMES WITH NO CONFIRMATION

Your database is gone. Restore from backup.

Prevent: never type -v after down on a machine with production data. Use named volumes, not bind mounts, for databases.


Nginx problems ​

nginx: configuration file test failed ​

Diagnose

bash
sudo nginx -t              # prints the file and line number
sudo nginx -T | less       # full resolved config

Solution: fix the reported line. Usually a missing semicolon or unclosed brace.

Prevent: sudo nginx -t && sudo systemctl reload nginx as one command, always.


502 Bad Gateway ​

Causes: app not running · wrong port in proxy_pass · app crashed · app bound to the wrong address

Diagnose — in this order

bash
sudo tail -20 /var/log/nginx/error.log       # the specific reason
pm2 list                                      # is it running?
sudo ss -tulpn | grep -E '3000|3001'          # listening? on what address?
curl -i http://127.0.0.1:3001/api/health      # ⭐ the key test
pm2 logs api --err --lines 50 --nostream      # why did it die?
Error log lineCauseFix
connect() failed (111: Connection refused)App not running / wrong portStart it; check proxy_pass
no live upstreamsAll marked failed by max_failsFix the app; they recover after fail_timeout
upstream prematurely closed connectionApp crashed mid-requestRead app logs for the exception
(13: Permission denied)Socket/SELinux permissionsjournalctl -u nginx

Solution: if curl http://127.0.0.1:3001/health works, fix Nginx's proxy_pass. If it fails, fix the app.

Prevent: health check in the deploy pipeline; wait_ready in PM2.


504 Gateway Timeout ​

Cause: the app took longer than proxy_read_timeout (60s).

Diagnose

bash
pm2 monit
sudo -u postgres psql -c "SELECT pid, now()-query_start AS d, left(query,80) FROM pg_stat_activity WHERE state!='idle' ORDER BY d DESC;"
sudo awk '{print $(NF-1), $7}' /var/log/nginx/access.log | sort -rn | head

RAISING proxy_read_timeout IS NOT A FIX

It converts a 504 into a worker blocked for five minutes. Under load every worker ends up blocked and the whole site stops. Fix the slow query, or move the work to a background queue.


403 Forbidden on static files ​

Diagnose

bash
sudo tail /var/log/nginx/error.log
namei -l /home/deploy/apps/myapp/current/frontend/.output/public/index.html
sudo -u www-data stat /home/deploy/apps/myapp/current/frontend/.output/public/

namei -l shows permissions for every path component — instantly revealing which directory blocks www-data.

Solution

bash
chmod 755 /home/deploy
chmod 755 /home/deploy/apps

Prevent: /home/deploy must stay 755 so Nginx can traverse it.


413 Request Entity Too Large ​

Solution

nginx
client_max_body_size 20M;

Default is 1M. Set it in http, server, or the specific location.


Wrong site served / default page shown ​

Diagnose

bash
sudo nginx -T | grep -E "server_name|listen"
ls -la /etc/nginx/sites-enabled/
curl -H "Host: app.example.com" http://127.0.0.1 -I

Solution: remove sites-enabled/default; verify the symlink exists; reload.


TLS problems ​

Certbot challenge failed ​

Diagnose

bash
# LOCAL
dig app.example.com +short
curl -I http://app.example.com
bash
# SERVER
sudo mkdir -p /var/www/certbot/.well-known/acme-challenge
echo test | sudo tee /var/www/certbot/.well-known/acme-challenge/test
curl http://app.example.com/.well-known/acme-challenge/test

If the last command does not print test, your HTTPS redirect is intercepting the challenge path.

Solution

nginx
server {
    listen 80;
    server_name app.example.com;
    location /.well-known/acme-challenge/ { root /var/www/certbot; }
    location / { return 301 https://$host$request_uri; }
}

RATE LIMITS

5 failed validations per hostname per hour. Use --dry-run while debugging.


Certificate expired despite the renewal timer ​

Diagnose

bash
systemctl status certbot.timer
sudo journalctl -u certbot --since "60 days ago" | grep -Ei "error|fail"
sudo certbot renew --dry-run
sudo certbot certificates
CauseFix
Port 80 closed after initial issuancesudo ufw allow 80/tcp
Nginx config lost the challenge locationRestore it
Missing reload hook — Nginx serves a stale cert from memoryAdd renewal-hooks/deploy/reload-nginx.sh
DNS changedUpdate the A record

Prevent: run certbot renew --dry-run after every Nginx or firewall change; monitor expiry externally.


ERR_SSL_PROTOCOL_ERROR ​

Diagnose

bash
sudo ss -tulpn | grep 443
sudo nginx -t
sudo openssl x509 -noout -modulus -in /etc/letsencrypt/live/app.example.com/fullchain.pem | openssl md5
sudo openssl rsa  -noout -modulus -in /etc/letsencrypt/live/app.example.com/privkey.pem   | openssl md5

The two hashes must match. If not, the certificate and key are from different issuances.


Redirect loop ​

Causes: Cloudflare SSL mode set to Flexible · app forcing HTTPS without seeing X-Forwarded-Proto

Solution: set Cloudflare to Full (Strict); ensure proxy_set_header X-Forwarded-Proto $scheme and trust proxy in NestJS.


DNS problems ​

DNS not resolving ​

Diagnose

bash
dig NS example.com +short                       # who is authoritative?
dig @ns1.provider.com app.example.com +short    # correct at the source?
dig @1.1.1.1 app.example.com +short             # propagated?
dig app.example.com +short                      # what do I get?
cat /etc/hosts                                   # stale test entry?
FindingCause
Authoritative is wrongRecord wrong — fix it in the DNS panel
Authoritative right, public resolver wrongStill propagating — wait out the TTL
Both right, yours wrongLocal cache or /etc/hosts
Edits have no effectYou are editing at the wrong provider — check the NS records

Prevent: TTL 300 during changes; verify DNS before running Certbot.


Application problems ​

WebSocket not connecting ​

Diagnose

bash
# LOCAL
curl -i -N -H "Connection: Upgrade" -H "Upgrade: websocket" \
     -H "Sec-WebSocket-Version: 13" -H "Sec-WebSocket-Key: x3JJHMbDL1EzLkh9GBhXDw==" \
     https://api.example.com/socket.io/?EIO=4\&transport=websocket
bash
# SERVER
sudo grep -i "socket.io" /var/log/nginx/api.example.com.access.log | tail
pm2 list        # more than one instance?
SymptomCauseFix
Handshake 400Missing proxy_http_version 1.1 or upgrade headersAdd both to the location
Disconnects every 60 secondsproxy_read_timeout defaultSet 7d on /socket.io/
Session ID unknownNo sticky sessions with cluster modeForce transports: ['websocket'], or add sticky sessions
Events reach only some clientsNo Redis adapter in cluster modeAdd @socket.io/redis-adapter
Messages arrive in burstsproxy_buffering onSet proxy_buffering off

PM2 process crash-looping ​

Diagnose

bash
pm2 list                                    # restart count climbing?
pm2 describe api
pm2 logs api --err --lines 100 --nostream
pm2 env 0 | grep -E "NODE_ENV|DATABASE_URL"
CauseCheck
Missing env varZod validation error in the logs
DB unreachablepsql from the same host
Port in usesudo ss -tulpn | grep 3001
Build missingls backend/dist/main.js

Solution: fix the cause, then pm2 reset api && pm2 restart api.

Prevent: Zod env validation (Level 8); min_uptime and max_restarts.


Out of memory ​

Diagnose

bash
free -h
sudo dmesg -T | grep -i "killed process"
ps aux --sort=-%mem | head -10
pm2 monit

Solution: add swap · reduce PM2 instances · lower max_memory_restart · reduce PostgreSQL shared_buffers · resize the server.

Prevent: OOMScoreAdjust=-500 on PostgreSQL; build in CI; alert at 90% memory.


Disk full ​

AT 100%, EVERYTHING BREAKS AT ONCE

PostgreSQL refuses writes, Nginx cannot log, builds fail, and SSH login may fail.

Diagnose

bash
df -h && df -i
du -h --max-depth=1 / 2>/dev/null | sort -rh | head -20
docker system df
journalctl --disk-usage
du -sh ~/.pm2/logs ~/apps/myapp/releases

Emergency triage

bash
sudo journalctl --vacuum-size=200M
pm2 flush
docker image prune -af           # ⚠️ never --volumes
sudo apt clean
sudo find /var/log -name "*.gz" -mtime +7 -delete
ls -1dt ~/apps/myapp/releases/* | tail -n +4 | xargs -r rm -rf

Prevent: log rotation everywhere; release pruning; alert at 85%.


CI/CD problems ​

CI/CD SSH failure ​

Diagnose

yaml
- run: ssh -vvv -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" "echo ok"
ErrorCauseFix
missing server hostEmpty or misnamed secretCheck the name; check environment scoping
Permission denied (publickey)Key not installed / wrong userVerify authorized_keys on the server
error in libcryptoMalformed private key secretRe-copy including BEGIN/END and trailing newline
Host key verification failedknown_hosts missingssh-keyscan into a secret
command not foundNon-interactive shell PATHSource NVM; symlink binaries

GitHub Actions secrets missing ​

Diagnose

yaml
- run: |
    [ -n "${{ secrets.SSH_HOST }}" ] || { echo "SSH_HOST is empty"; exit 1; }

Causes: typo (names are case-sensitive) · created in the wrong repository · it is an environment secret but the job declares no environment: · the workflow is triggered by pull_request from a fork (secrets are deliberately withheld)


GitLab Runner failure ​

SymptomCauseFix
Pipeline stuck "pending"No runner with matching tags, or offlineSettings → CI/CD → Runners
Secret emptyVariable is Protected, branch is notProtect the branch
Services unreachableWrong hostnameUse the service alias, not 127.0.0.1
error in libcryptoKey stored as a plain variableUse a File-type variable

Deployment succeeds but the application does not update ​

THE MOST CONFUSING FAILURE — DIAGNOSE IN THIS EXACT ORDER

bash
# SERVER
cd ~/apps/myapp/current

git log -1 --oneline                                    # 1. did the code update?
find backend/src -newer backend/dist/main.js -name "*.ts" | head   # 2. did the build run?
pm2 list                                                 # 3. did PM2 restart? (check uptime)
pm2 env 0 | grep GIT_COMMIT                              # 4. new env loaded?
curl -s http://127.0.0.1:3001/api/health | jq .commit    # 5. what does the app say?
readlink -f ~/apps/myapp/current                         # 6. did the symlink move?
FindingCauseFix
Old commit at step 1Wrong branch, or fetch failedCheck the branch; check the deploy key
Source newer than dist at step 2Build did not run or failed silentlyAdd set -e to the deploy script
PM2 uptime in hours at step 3pm2 reload did not runCheck the script's exit path
Old SHA at step 4Missing --update-envpm2 reload ... --update-env
Symlink unchanged at step 6The atomic switch failedCheck the mv -Tf step
All correct but the browser shows oldBrowser or CDN cacheHard refresh; check Cache-Control on HTML

THE ROOT CAUSE IS USUALLY A MISSING set -e

Without it, a failed pnpm build does not stop the script. It proceeds to pm2 reload, which restarts the app on the old build, and the pipeline reports success. Green tick, unchanged site.

Every deploy script starts with set -euo pipefail.


Quick reference — first commands by symptom ​

SymptomFirst command
Site completely downpm2 list && systemctl is-active nginx postgresql redis-server
502 Bad Gatewaycurl -i http://127.0.0.1:3001/api/health
504 Gateway Timeoutpm2 monit then check pg_stat_activity
Slow siteuptime && free -h && df -h
SSH failsssh -v — look for Connection established
Deploy failedpm2 logs --err --lines 100 --nostream
Certificate errorsudo certbot certificates
Database errorsudo tail -50 /var/log/postgresql/*.log
Something disappearedsudo dmesg -T | grep -i "killed process"
Weird intermittent errorsdf -h (check for a full disk)
Anything after a deployfind src -newer dist/main.js

Next: Level 27 — Production Checklists →