Level 13 — PM2
Your application must survive crashes, reboots, and deployments without a human watching. That is what a process manager does.
Why a Node application needs a process manager
Start your API by hand:
# SERVER
node backend/dist/main.jsEverything works — until any of the following:
| Event | Without a process manager | With PM2 |
|---|---|---|
| You close the SSH session | SIGHUP kills the process. Site down. | Runs independently of your session |
| Unhandled exception | Process exits. Site down until a human notices. | Restarted in milliseconds |
| Server reboots | Nothing starts. Site down until you SSH in. | Started automatically at boot |
| Memory leak | Grows until the OOM killer intervenes | Restarted at a configured threshold |
| Deployment | Kill and restart — visible downtime, dropped requests | reload — zero downtime |
| 4 CPU cores available | One core used | Cluster mode uses all of them |
| Need to see logs | Whatever scrolled past in your terminal | Persisted, rotated, queryable |
PM2 vs systemd
systemd is already on your server and does restarts, boot startup, and logging perfectly well. It is arguably the more "correct" Unix answer.
PM2 wins for Node specifically because of cluster mode and reload. Cluster mode runs N instances behind Node's built-in load balancer with no extra configuration. pm2 reload restarts instances one at a time, waiting for each to come up before killing the next — genuine zero-downtime deploys on a single server. Reproducing that with systemd requires socket activation or a second port plus an Nginx config swap.
PM2 also gives you pm2 logs, pm2 monit, and a readable pm2 list out of the box.
Use PM2 for Node apps. Use systemd for everything else (Nginx, PostgreSQL, Redis, and PM2 itself via the startup script).
Installing PM2
# SERVER — as deploy, NOT with sudo
npm install -g pm2
pm2 -vDO NOT sudo npm install -g pm2
With NVM, that writes root-owned files into ~/.nvm and breaks later non-sudo installs. It also means PM2 runs as root, so your app runs as root — the thing Level 3 exists to prevent.
PM2 must be installed and run as deploy.
The ecosystem file
Do not manage production processes with ad-hoc pm2 start commands. Put everything in a config file, commit it, and deploy it.
// /home/deploy/apps/myapp/ecosystem.config.cjs
module.exports = {
apps: [
{
name: 'api',
cwd: '/home/deploy/apps/myapp/backend',
script: 'dist/main.js',
// Cluster mode: N processes sharing port 3001
exec_mode: 'cluster',
instances: 2,
env: {
NODE_ENV: 'production',
PORT: 3001,
HOST: '127.0.0.1',
},
// Restart if RSS exceeds this — a leak safety net
max_memory_restart: '512M',
// Graceful shutdown
kill_timeout: 10000, // ms to wait after SIGTERM before SIGKILL
listen_timeout: 8000, // ms to wait for the new process to listen
wait_ready: true, // wait for process.send('ready')
// Crash-loop protection
min_uptime: '30s', // shorter than this counts as a failed start
max_restarts: 10, // give up after 10 failures in a row
restart_delay: 2000,
exp_backoff_restart_delay: 200,
autorestart: true,
watch: false, // NEVER true in production
// Logs
error_file: '/home/deploy/logs/api-error.log',
out_file: '/home/deploy/logs/api-out.log',
merge_logs: true,
log_date_format: 'YYYY-MM-DD HH:mm:ss Z',
node_args: '--enable-source-maps',
},
{
name: 'web',
cwd: '/home/deploy/apps/myapp/frontend',
script: '.output/server/index.mjs',
exec_mode: 'cluster',
instances: 2,
env: {
NODE_ENV: 'production',
PORT: 3000,
HOST: '127.0.0.1',
NITRO_PORT: 3000,
NITRO_HOST: '127.0.0.1',
},
max_memory_restart: '768M',
kill_timeout: 10000,
listen_timeout: 8000,
min_uptime: '30s',
max_restarts: 10,
autorestart: true,
watch: false,
error_file: '/home/deploy/logs/web-error.log',
out_file: '/home/deploy/logs/web-out.log',
merge_logs: true,
log_date_format: 'YYYY-MM-DD HH:mm:ss Z',
},
],
};USE THE .cjs EXTENSION
If your package.json has "type": "module", a file named ecosystem.config.js is treated as an ES module and module.exports fails with module is not defined. Naming it .cjs forces CommonJS and avoids the problem entirely.
Every option explained
| Option | Meaning | Guidance |
|---|---|---|
name | Identifier for pm2 restart <name> | Short and stable |
cwd | Working directory | Absolute path. Relative paths in your app resolve from here. |
script | Entrypoint, relative to cwd | dist/main.js, .output/server/index.mjs |
exec_mode | fork or cluster | See below |
instances | Process count | 2, or 'max' for one per core |
env | Environment variables | See the --update-env warning |
max_memory_restart | Restart above this RSS | Safety net for leaks, not a fix |
kill_timeout | ms between SIGTERM and SIGKILL | Longer than your slowest request |
listen_timeout | ms to wait for the new instance during reload | 8000 is safe for Nest/Nuxt |
wait_ready | Wait for process.send('ready') | Most accurate readiness signal |
min_uptime | Below this, a start counts as failed | Prevents a fast crash loop from looking healthy |
max_restarts | Consecutive failures before giving up | Stops infinite restart loops |
exp_backoff_restart_delay | Exponentially increasing delay | Reduces load during an outage |
watch | Restart on file change | Never true in production |
error_file/out_file | Log paths | Absolute paths |
node_args | Flags passed to Node | --enable-source-maps, --max-old-space-size |
watch: true IN PRODUCTION CAUSES RESTART STORMS
A deploy writes hundreds of files. Watch mode restarts on each one, so the app thrashes through dozens of restarts mid-deploy, potentially hitting max_restarts and stopping entirely. It also holds file watches on node_modules, exhausting inotify limits.
watch is a development convenience. Always false in production.
Cluster mode vs fork mode
| fork | cluster | |
|---|---|---|
| Processes | 1 | N |
| CPU cores used | 1 | N |
Zero-downtime reload | ❌ No | ✅ Yes |
| In-memory state shared | N/A | ❌ Not shared |
| Port binding | Direct | Via SO_REUSEPORT / master |
| Use for | Workers, cron, non-HTTP | HTTP servers |
CLUSTER MODE BREAKS ANYTHING THAT ASSUMES ONE PROCESS
Each instance is a completely separate process with its own memory. Things that silently break:
| Pattern | Why it breaks | Fix |
|---|---|---|
| In-memory sessions | User hits a different worker, appears logged out | Redis session store (Level 10) |
| In-memory cache | Each worker has a different view; invalidation misses | Redis |
| In-memory rate limiting | Each worker counts separately — your limit is effectively ×N | Redis |
setInterval cron | Runs N times, so N emails per user | pm2-cron, a dedicated fork-mode worker, or a distributed lock |
| Socket.IO without an adapter | Events only reach clients on the emitting worker | @socket.io/redis-adapter + sticky sessions |
let counter = 0 | Diverges per worker | Redis INCR |
The symptom is always the same: works in development (one process), intermittently broken in production. Requests randomly succeed or fail depending on which worker handles them, which makes it maddening to debug.
How many instances?
instances: 2 // explicit
instances: 'max' // one per CPU core
instances: -1 // cores minus 1instances: 'max' IS USUALLY WRONG ON A SMALL VPS
On a 2-core / 4 GB server running Nuxt and NestJS, 'max' gives 2+2 = 4 Node processes at ~150–300 MB each, plus PostgreSQL, plus Redis, plus Nginx. You are out of RAM.
Also remember Prisma's connection_limit is per process (Level 9) — more instances means more database connections.
Sensible starting point on 2 cores / 4 GB:
api: 2 instancesweb: 2 instances- Then watch
pm2 monitandfree -hand adjust.
Socket.IO with cluster mode
WEBSOCKETS NEED BOTH THE REDIS ADAPTER AND STICKY SESSIONS
Two separate problems:
1. Broadcast isolation — io.emit() on worker 1 only reaches worker 1's clients. Solved by @socket.io/redis-adapter (Level 10).
2. Handshake affinity — Socket.IO starts with HTTP long-polling, which makes several requests that must all reach the same worker. Round-robin load balancing sends them to different workers and the handshake fails with Session ID unknown. Solved by sticky sessions.
For PM2 cluster mode, enable Node's cluster affinity:
{
name: 'api',
exec_mode: 'cluster',
instances: 2,
instance_var: 'INSTANCE_ID',
env: { NODE_ENV: 'production', PORT: 3001 },
}and use @socket.io/sticky in your bootstrap, or configure ip_hash in Nginx's upstream (Level 14).
The simplest reliable alternative: force WebSocket-only transport, which has no multi-request handshake:
const io = new Server(server, { transports: ['websocket'] });Client-side too. The cost is losing the long-polling fallback for clients behind proxies that block WebSockets — increasingly rare, and usually an acceptable trade.
Starting the application
# SERVER
mkdir -p /home/deploy/logs
cd /home/deploy/apps/myapp
pm2 start ecosystem.config.cjs
pm2 list┌────┬──────┬─────────┬─────────┬──────┬────────┬──────┬────────┬──────┬──────────┐
│ id │ name │ mode │ status │ ↺ │ cpu │ mem │ user │ ... │ │
├────┼──────┼─────────┼─────────┼──────┼────────┼──────┼────────┼──────┼──────────┤
│ 0 │ api │ cluster │ online │ 0 │ 0% │ 92mb │ deploy │ │ │
│ 1 │ api │ cluster │ online │ 0 │ 0% │ 89mb │ deploy │ │ │
│ 2 │ web │ cluster │ online │ 0 │ 0% │ 145mb│ deploy │ │ │
│ 3 │ web │ cluster │ online │ 0 │ 0% │ 141mb│ deploy │ │ │
└────┴──────┴─────────┴─────────┴──────┴────────┴──────┴────────┴──────┴──────────┘The ↺ column is the restart count. A number that keeps climbing means a crash loop — check the logs immediately.
Startup on boot
# SERVER
pm2 startupIt prints a command to run with sudo:
sudo env PATH=$PATH:/home/deploy/.nvm/versions/node/v22.11.0/bin \
/home/deploy/.nvm/versions/node/v22.11.0/lib/node_modules/pm2/bin/pm2 \
startup systemd -u deploy --hp /home/deployRun exactly that. It creates a systemd unit pm2-deploy.service that starts PM2 as deploy at boot.
# SERVER — save the CURRENT process list as what should be resurrected
pm2 savepm2 save IS THE STEP EVERYONE FORGETS
pm2 startup installs the boot service. pm2 save writes the current process list to ~/.pm2/dump.pm2. Without pm2 save, the server reboots and PM2 starts with nothing running.
Run pm2 save after every change to which apps are running. Then actually test it:
sudo reboot
# wait ~60 seconds, reconnect
pm2 list # everything should be online
curl -I https://app.example.comDo this once, deliberately, on a quiet day. Discovering it during a real incident is the wrong time.
# SERVER — verify the systemd side
systemctl status pm2-deploy
systemctl is-enabled pm2-deployUPGRADING NODE BREAKS THE STARTUP SCRIPT
The generated unit hardcodes the Node path (.../v22.11.0/bin). After nvm install 24, that path may no longer exist and PM2 fails to start at boot — which you discover after the next reboot.
After any Node upgrade:
pm2 unstartup systemd
pm2 startup # run the printed command again
pm2 saveDaily commands
pm2 list # status of everything
pm2 status # same
pm2 describe api # full details for one app
pm2 monit # live dashboard: CPU, memory, logs
pm2 logs # stream all logs
pm2 logs api --lines 100 # last 100 lines of one app
pm2 logs api --err # errors only
pm2 logs --nostream --lines 50 # print and exit (for scripts)
pm2 flush # clear all log files
pm2 restart api # hard restart — brief downtime
pm2 reload api # rolling restart — zero downtime
pm2 stop api # stop but keep in the list
pm2 delete api # remove from PM2 entirely
pm2 restart all
pm2 reload all
pm2 env 0 # environment of process 0
pm2 prettylist # full JSON
pm2 reset api # reset restart countersrestart vs reload — the difference that matters
restart | reload | |
|---|---|---|
| Downtime | Yes, brief | None |
| Requires cluster mode | No | Yes |
| In-flight requests | Dropped | Completed |
| Use for | Fork-mode apps, forcing a clean start | Every normal deploy |
reload IN FORK MODE IS JUST A RESTART
With exec_mode: 'fork' there is only one process, so there is nothing to roll. PM2 silently does a restart, and you get downtime while believing you do not. Zero-downtime requires exec_mode: 'cluster' and instances >= 2.
reload WITHOUT GRACEFUL SHUTDOWN STILL DROPS REQUESTS
PM2 sends SIGTERM and waits kill_timeout ms. If your app ignores SIGTERM, PM2 SIGKILLs it after the timeout and every in-flight request dies.
Your app must handle it — app.enableShutdownHooks() in NestJS (Level 12). Test it: start a request that takes 5 seconds, run pm2 reload api, and confirm it completes.
Environment variables and --update-env
pm2 restart DOES NOT PICK UP .env CHANGES
PM2 caches the environment from when the process was created. Editing .env and running pm2 restart api leaves the old values in place.
pm2 reload ecosystem.config.cjs --update-env # ✅ correct
pm2 restart api --update-env # ✅ also works
pm2 delete api && pm2 start ecosystem.config.cjs # ✅ guaranteed clean
pm2 restart api # ❌ old environmentVerify:
pm2 env 0 | grep -E "NODE_ENV|DATABASE_URL""I changed the database password and the app still uses the old one" is this, every single time.
Log management
PM2 writes to the files in your ecosystem config and does not rotate them by default.
PM2 LOGS WILL FILL YOUR DISK
A chatty app produces gigabytes over months. When the disk hits 100%: PostgreSQL stops accepting writes, Nginx cannot log, builds fail, and SSH logins may fail. Everything breaks at once, confusingly.
Install rotation immediately:
pm2 install pm2-logrotate
pm2 set pm2-logrotate:max_size 10M
pm2 set pm2-logrotate:retain 14
pm2 set pm2-logrotate:compress true
pm2 set pm2-logrotate:rotateInterval '0 0 * * *'
pm2 set pm2-logrotate:workerInterval 300Verify: pm2 conf pm2-logrotate and check ls -la ~/.pm2/logs/ after a day.
# SERVER
du -sh /home/deploy/logs ~/.pm2/logs
pm2 flush # truncate all logs now
tail -f /home/deploy/logs/api-error.logSTRUCTURED LOGGING PAYS OFF QUICKLY
Plain console.log produces logs you can only grep. JSON logs can be queried:
import { Logger } from 'nestjs-pino';
// { "level":"error", "time":..., "reqId":"...", "msg":"payment failed", "userId":"..." }pm2 logs api --raw --nostream | jq 'select(.level=="error")'
pm2 logs api --raw --nostream | jq -r 'select(.userId=="abc") | .msg'Include a request ID in every log line so you can reconstruct a single request across services. (Level 21)
Memory management
max_memory_restart: '512M'PM2 checks RSS periodically and restarts the instance when it exceeds the threshold. In cluster mode it restarts one instance at a time, so there is no downtime.
THIS IS A SAFETY NET, NOT A FIX
If max_memory_restart fires regularly, you have a memory leak. Find it:
pm2 monit # watch memory climb
node --inspect=127.0.0.1:9229 dist/main.js # locally, take heap snapshotsCommon Node leaks: event listeners added per request without removal, an unbounded in-memory cache/Map, closures capturing large objects, and Prisma clients instantiated per request instead of once.
Set the threshold above normal peak but well below available RAM. If normal usage is 200 MB, 512M is reasonable; 2G on a 4 GB server means the OOM killer wins first.
Crash-loop protection
min_uptime: '30s',
max_restarts: 10,
exp_backoff_restart_delay: 200,If the app crashes on startup — bad .env, database unreachable, port in use — PM2 would otherwise restart it forever, burning CPU and filling logs. With these settings PM2 gives up after 10 rapid failures and marks the process errored.
A STOPPED APP IS BETTER THAN A CRASH LOOP
errored status is visible in pm2 list and in monitoring. An infinite restart loop can look "alive" while serving nothing, and it drowns the real error in a flood of repeated stack traces.
When you see errored:
pm2 describe api
pm2 logs api --err --lines 100 --nostream
pm2 reset api && pm2 restart api # after fixing the causeHealth checks
// NestJS
@Get('health')
health() {
return {
status: 'ok',
uptime: process.uptime(),
commit: process.env.GIT_COMMIT ?? 'unknown',
pid: process.pid,
instance: process.env.NODE_APP_INSTANCE,
};
}# SERVER — after every deploy
curl -fsS http://127.0.0.1:3001/api/health | jqUSE wait_ready FOR ACCURATE ROLLING RELOADS
By default PM2 considers an instance ready when the process starts, which is before Nest has connected to the database. With wait_ready: true, PM2 waits for an explicit signal:
await app.listen(env.PORT, '127.0.0.1');
if (process.send) process.send('ready');Now pm2 reload only kills the old instance once the new one is genuinely serving traffic. Without this, there is a window during reload where requests hit a process that is up but not ready — occasional 502s during deploys.
PM2 in the deploy script
#!/usr/bin/env bash
set -euo pipefail
APP_DIR="/home/deploy/apps/myapp"
export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"
cd "$APP_DIR"
git fetch origin --prune
git reset --hard origin/production
git clean -fd
pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend prisma migrate deploy
pnpm --filter backend build
pnpm --filter frontend build
export GIT_COMMIT=$(git rev-parse --short HEAD)
pm2 reload ecosystem.config.cjs --update-env
pm2 save
# Verify — fail loudly if the new version is not healthy
sleep 5
for i in {1..10}; do
if curl -fsS http://127.0.0.1:3001/api/health > /dev/null; then
echo "✅ API healthy"; break
fi
[ "$i" -eq 10 ] && { echo "❌ health check failed"; pm2 logs api --err --lines 50 --nostream; exit 1; }
sleep 2
done
curl -fsS -o /dev/null http://127.0.0.1:3000/ && echo "✅ Web healthy"
echo "✅ Deployed $GIT_COMMIT"pm2 save AFTER EVERY DEPLOY
If the deploy changed which processes exist (added an app, changed instance count), the saved dump is stale and a reboot restores the old configuration.
Rollback
# SERVER — quick rollback with the in-place layout
cd /home/deploy/apps/myapp
git reset --hard <previous-commit-sha>
pnpm install --frozen-lockfile
pnpm --filter backend build && pnpm --filter frontend build
pm2 reload ecosystem.config.cjs --update-envDATABASE MIGRATIONS DO NOT ROLL BACK WITH THE CODE
Reverting code to a commit before a migration leaves the database with the new schema and the code expecting the old one. Prisma has no automatic down-migrations.
This is why additive migrations matter (Level 9): if a migration only adds things, old code keeps working and rollback is safe. If it dropped a column, rolling back the code is not enough — you need a restore, which means downtime.
Rule: any migration that removes or renames something must be split across two deploys.
Instant rollback needs the release-directory pattern in Level 20.
Troubleshooting
| Problem | Cause | Diagnose | Fix |
|---|---|---|---|
pm2: command not found over SSH | NVM not loaded in non-interactive shell | ssh server "which pm2" | Symlink to /usr/local/bin (Level 6) |
| Apps gone after reboot | pm2 save not run, or startup not installed | systemctl status pm2-deploy | pm2 startup + pm2 save, then test a reboot |
| Restart count climbing | Crash loop | pm2 logs api --err --lines 100 | Read the error; check .env and DB connectivity |
Status errored | Hit max_restarts | pm2 describe api | Fix cause, pm2 reset api && pm2 restart api |
EADDRINUSE | Old process still bound | sudo ss -tulpn | grep 3001 | pm2 delete api; kill the stray PID |
| Env changes not applied | Missing --update-env | pm2 env 0 | pm2 reload ... --update-env |
| Memory grows steadily | Leak | pm2 monit | Heap snapshots; max_memory_restart as a stopgap |
| 502 during every deploy | No graceful shutdown, or reload in fork mode | pm2 describe api — check mode | Cluster mode + enableShutdownHooks() + wait_ready |
| WebSockets flaky | Cluster mode without adapter/sticky sessions | pm2 list — >1 instance? | Redis adapter + sticky sessions |
| Disk full | PM2 logs unrotated | du -sh ~/.pm2/logs | pm2 install pm2-logrotate |
| Cron/scheduled task runs N times | setInterval in cluster mode | — | Move to a fork-mode worker or use a distributed lock |
| PM2 daemon itself gone | Killed, or OOM | pm2 ping | pm2 resurrect |
# SERVER — general PM2 debugging
pm2 describe api
pm2 logs api --err --lines 200 --nostream
pm2 prettylist | jq '.[] | {name, status, restart_time, pm_uptime}'
cat ~/.pm2/pm2.log # the PM2 daemon's own log
pm2 ping # is the daemon alive?
pm2 resurrect # restore from the saved dump
pm2 update # reload the in-memory daemon after a PM2 upgradeProduction Checklist — Level 13
- [ ] PM2 installed as
deploy, not with sudo - [ ] All processes defined in a committed
ecosystem.config.cjs - [ ]
exec_mode: 'cluster'withinstances >= 2for HTTP apps - [ ] Instance count sized to available RAM, not blindly
'max' - [ ]
NODE_ENV=productionandHOST=127.0.0.1set for both apps - [ ]
watch: false - [ ]
max_memory_restartset below available RAM - [ ]
kill_timeoutlonger than the slowest request - [ ]
min_uptimeandmax_restartsset for crash-loop protection - [ ] Graceful shutdown implemented and tested with a slow request during reload
- [ ]
wait_ready: truewithprocess.send('ready')in the app - [ ]
pm2 startuprun andpm2 saveexecuted - [ ] A real reboot has been tested and everything came back
- [ ]
pm2-logrotateinstalled and configured - [ ] Deploys use
pm2 reload ... --update-env, never barerestart - [ ] Health endpoint exposing commit SHA, verified after every deploy
- [ ] Redis-backed sessions, cache, and rate limiting (cluster-safe)
- [ ] Socket.IO Redis adapter + sticky sessions configured
- [ ] No
setIntervalcron jobs running in cluster mode - [ ] Rollback procedure documented, with awareness that migrations do not revert
Next: Level 14 — Nginx →