Skip to content

Level 13 — PM2 ​

Your application must survive crashes, reboots, and deployments without a human watching. That is what a process manager does.

Why a Node application needs a process manager ​

Start your API by hand:

bash
# SERVER
node backend/dist/main.js

Everything works — until any of the following:

EventWithout a process managerWith PM2
You close the SSH sessionSIGHUP kills the process. Site down.Runs independently of your session
Unhandled exceptionProcess exits. Site down until a human notices.Restarted in milliseconds
Server rebootsNothing starts. Site down until you SSH in.Started automatically at boot
Memory leakGrows until the OOM killer intervenesRestarted at a configured threshold
DeploymentKill and restart — visible downtime, dropped requestsreload — zero downtime
4 CPU cores availableOne core usedCluster mode uses all of them
Need to see logsWhatever scrolled past in your terminalPersisted, rotated, queryable

PM2 vs systemd

systemd is already on your server and does restarts, boot startup, and logging perfectly well. It is arguably the more "correct" Unix answer.

PM2 wins for Node specifically because of cluster mode and reload. Cluster mode runs N instances behind Node's built-in load balancer with no extra configuration. pm2 reload restarts instances one at a time, waiting for each to come up before killing the next — genuine zero-downtime deploys on a single server. Reproducing that with systemd requires socket activation or a second port plus an Nginx config swap.

PM2 also gives you pm2 logs, pm2 monit, and a readable pm2 list out of the box.

Use PM2 for Node apps. Use systemd for everything else (Nginx, PostgreSQL, Redis, and PM2 itself via the startup script).

Installing PM2 ​

bash
# SERVER — as deploy, NOT with sudo
npm install -g pm2
pm2 -v

DO NOT sudo npm install -g pm2

With NVM, that writes root-owned files into ~/.nvm and breaks later non-sudo installs. It also means PM2 runs as root, so your app runs as root — the thing Level 3 exists to prevent.

PM2 must be installed and run as deploy.

The ecosystem file ​

Do not manage production processes with ad-hoc pm2 start commands. Put everything in a config file, commit it, and deploy it.

js
// /home/deploy/apps/myapp/ecosystem.config.cjs
module.exports = {
  apps: [
    {
      name: 'api',
      cwd: '/home/deploy/apps/myapp/backend',
      script: 'dist/main.js',

      // Cluster mode: N processes sharing port 3001
      exec_mode: 'cluster',
      instances: 2,

      env: {
        NODE_ENV: 'production',
        PORT: 3001,
        HOST: '127.0.0.1',
      },

      // Restart if RSS exceeds this — a leak safety net
      max_memory_restart: '512M',

      // Graceful shutdown
      kill_timeout: 10000,        // ms to wait after SIGTERM before SIGKILL
      listen_timeout: 8000,       // ms to wait for the new process to listen
      wait_ready: true,           // wait for process.send('ready')

      // Crash-loop protection
      min_uptime: '30s',          // shorter than this counts as a failed start
      max_restarts: 10,           // give up after 10 failures in a row
      restart_delay: 2000,
      exp_backoff_restart_delay: 200,

      autorestart: true,
      watch: false,               // NEVER true in production

      // Logs
      error_file: '/home/deploy/logs/api-error.log',
      out_file: '/home/deploy/logs/api-out.log',
      merge_logs: true,
      log_date_format: 'YYYY-MM-DD HH:mm:ss Z',

      node_args: '--enable-source-maps',
    },

    {
      name: 'web',
      cwd: '/home/deploy/apps/myapp/frontend',
      script: '.output/server/index.mjs',

      exec_mode: 'cluster',
      instances: 2,

      env: {
        NODE_ENV: 'production',
        PORT: 3000,
        HOST: '127.0.0.1',
        NITRO_PORT: 3000,
        NITRO_HOST: '127.0.0.1',
      },

      max_memory_restart: '768M',
      kill_timeout: 10000,
      listen_timeout: 8000,
      min_uptime: '30s',
      max_restarts: 10,
      autorestart: true,
      watch: false,

      error_file: '/home/deploy/logs/web-error.log',
      out_file: '/home/deploy/logs/web-out.log',
      merge_logs: true,
      log_date_format: 'YYYY-MM-DD HH:mm:ss Z',
    },
  ],
};

USE THE .cjs EXTENSION

If your package.json has "type": "module", a file named ecosystem.config.js is treated as an ES module and module.exports fails with module is not defined. Naming it .cjs forces CommonJS and avoids the problem entirely.

Every option explained ​

OptionMeaningGuidance
nameIdentifier for pm2 restart <name>Short and stable
cwdWorking directoryAbsolute path. Relative paths in your app resolve from here.
scriptEntrypoint, relative to cwddist/main.js, .output/server/index.mjs
exec_modefork or clusterSee below
instancesProcess count2, or 'max' for one per core
envEnvironment variablesSee the --update-env warning
max_memory_restartRestart above this RSSSafety net for leaks, not a fix
kill_timeoutms between SIGTERM and SIGKILLLonger than your slowest request
listen_timeoutms to wait for the new instance during reload8000 is safe for Nest/Nuxt
wait_readyWait for process.send('ready')Most accurate readiness signal
min_uptimeBelow this, a start counts as failedPrevents a fast crash loop from looking healthy
max_restartsConsecutive failures before giving upStops infinite restart loops
exp_backoff_restart_delayExponentially increasing delayReduces load during an outage
watchRestart on file changeNever true in production
error_file/out_fileLog pathsAbsolute paths
node_argsFlags passed to Node--enable-source-maps, --max-old-space-size

watch: true IN PRODUCTION CAUSES RESTART STORMS

A deploy writes hundreds of files. Watch mode restarts on each one, so the app thrashes through dozens of restarts mid-deploy, potentially hitting max_restarts and stopping entirely. It also holds file watches on node_modules, exhausting inotify limits.

watch is a development convenience. Always false in production.

Cluster mode vs fork mode ​

forkcluster
Processes1N
CPU cores used1N
Zero-downtime reload❌ No✅ Yes
In-memory state sharedN/A❌ Not shared
Port bindingDirectVia SO_REUSEPORT / master
Use forWorkers, cron, non-HTTPHTTP servers

CLUSTER MODE BREAKS ANYTHING THAT ASSUMES ONE PROCESS

Each instance is a completely separate process with its own memory. Things that silently break:

PatternWhy it breaksFix
In-memory sessionsUser hits a different worker, appears logged outRedis session store (Level 10)
In-memory cacheEach worker has a different view; invalidation missesRedis
In-memory rate limitingEach worker counts separately — your limit is effectively ×NRedis
setInterval cronRuns N times, so N emails per userpm2-cron, a dedicated fork-mode worker, or a distributed lock
Socket.IO without an adapterEvents only reach clients on the emitting worker@socket.io/redis-adapter + sticky sessions
let counter = 0Diverges per workerRedis INCR

The symptom is always the same: works in development (one process), intermittently broken in production. Requests randomly succeed or fail depending on which worker handles them, which makes it maddening to debug.

How many instances? ​

js
instances: 2          // explicit
instances: 'max'      // one per CPU core
instances: -1         // cores minus 1

instances: 'max' IS USUALLY WRONG ON A SMALL VPS

On a 2-core / 4 GB server running Nuxt and NestJS, 'max' gives 2+2 = 4 Node processes at ~150–300 MB each, plus PostgreSQL, plus Redis, plus Nginx. You are out of RAM.

Also remember Prisma's connection_limit is per process (Level 9) — more instances means more database connections.

Sensible starting point on 2 cores / 4 GB:

  • api: 2 instances
  • web: 2 instances
  • Then watch pm2 monit and free -h and adjust.

Socket.IO with cluster mode ​

WEBSOCKETS NEED BOTH THE REDIS ADAPTER AND STICKY SESSIONS

Two separate problems:

1. Broadcast isolation — io.emit() on worker 1 only reaches worker 1's clients. Solved by @socket.io/redis-adapter (Level 10).

2. Handshake affinity — Socket.IO starts with HTTP long-polling, which makes several requests that must all reach the same worker. Round-robin load balancing sends them to different workers and the handshake fails with Session ID unknown. Solved by sticky sessions.

For PM2 cluster mode, enable Node's cluster affinity:

js
{
  name: 'api',
  exec_mode: 'cluster',
  instances: 2,
  instance_var: 'INSTANCE_ID',
  env: { NODE_ENV: 'production', PORT: 3001 },
}

and use @socket.io/sticky in your bootstrap, or configure ip_hash in Nginx's upstream (Level 14).

The simplest reliable alternative: force WebSocket-only transport, which has no multi-request handshake:

ts
const io = new Server(server, { transports: ['websocket'] });

Client-side too. The cost is losing the long-polling fallback for clients behind proxies that block WebSockets — increasingly rare, and usually an acceptable trade.

Starting the application ​

bash
# SERVER
mkdir -p /home/deploy/logs
cd /home/deploy/apps/myapp
pm2 start ecosystem.config.cjs
pm2 list
┌────┬──────┬─────────┬─────────┬──────┬────────┬──────┬────────┬──────┬──────────┐
│ id │ name │ mode    │ status  │ ↺    │ cpu    │ mem  │ user   │ ...  │          │
├────┼──────┼─────────┼─────────┼──────┼────────┼──────┼────────┼──────┼──────────┤
│ 0  │ api  │ cluster │ online  │ 0    │ 0%     │ 92mb │ deploy │      │          │
│ 1  │ api  │ cluster │ online  │ 0    │ 0%     │ 89mb │ deploy │      │          │
│ 2  │ web  │ cluster │ online  │ 0    │ 0%     │ 145mb│ deploy │      │          │
│ 3  │ web  │ cluster │ online  │ 0    │ 0%     │ 141mb│ deploy │      │          │
└────┴──────┴─────────┴─────────┴──────┴────────┴──────┴────────┴──────┴──────────┘

The ↺ column is the restart count. A number that keeps climbing means a crash loop — check the logs immediately.

Startup on boot ​

bash
# SERVER
pm2 startup

It prints a command to run with sudo:

sudo env PATH=$PATH:/home/deploy/.nvm/versions/node/v22.11.0/bin \
  /home/deploy/.nvm/versions/node/v22.11.0/lib/node_modules/pm2/bin/pm2 \
  startup systemd -u deploy --hp /home/deploy

Run exactly that. It creates a systemd unit pm2-deploy.service that starts PM2 as deploy at boot.

bash
# SERVER — save the CURRENT process list as what should be resurrected
pm2 save

pm2 save IS THE STEP EVERYONE FORGETS

pm2 startup installs the boot service. pm2 save writes the current process list to ~/.pm2/dump.pm2. Without pm2 save, the server reboots and PM2 starts with nothing running.

Run pm2 save after every change to which apps are running. Then actually test it:

bash
sudo reboot
# wait ~60 seconds, reconnect
pm2 list                 # everything should be online
curl -I https://app.example.com

Do this once, deliberately, on a quiet day. Discovering it during a real incident is the wrong time.

bash
# SERVER — verify the systemd side
systemctl status pm2-deploy
systemctl is-enabled pm2-deploy

UPGRADING NODE BREAKS THE STARTUP SCRIPT

The generated unit hardcodes the Node path (.../v22.11.0/bin). After nvm install 24, that path may no longer exist and PM2 fails to start at boot — which you discover after the next reboot.

After any Node upgrade:

bash
pm2 unstartup systemd
pm2 startup           # run the printed command again
pm2 save

Daily commands ​

bash
pm2 list                       # status of everything
pm2 status                     # same
pm2 describe api               # full details for one app
pm2 monit                      # live dashboard: CPU, memory, logs
pm2 logs                       # stream all logs
pm2 logs api --lines 100       # last 100 lines of one app
pm2 logs api --err             # errors only
pm2 logs --nostream --lines 50 # print and exit (for scripts)
pm2 flush                      # clear all log files

pm2 restart api                # hard restart — brief downtime
pm2 reload api                 # rolling restart — zero downtime
pm2 stop api                   # stop but keep in the list
pm2 delete api                 # remove from PM2 entirely
pm2 restart all
pm2 reload all

pm2 env 0                      # environment of process 0
pm2 prettylist                 # full JSON
pm2 reset api                  # reset restart counters

restart vs reload — the difference that matters ​

restartreload
DowntimeYes, briefNone
Requires cluster modeNoYes
In-flight requestsDroppedCompleted
Use forFork-mode apps, forcing a clean startEvery normal deploy

reload IN FORK MODE IS JUST A RESTART

With exec_mode: 'fork' there is only one process, so there is nothing to roll. PM2 silently does a restart, and you get downtime while believing you do not. Zero-downtime requires exec_mode: 'cluster' and instances >= 2.

reload WITHOUT GRACEFUL SHUTDOWN STILL DROPS REQUESTS

PM2 sends SIGTERM and waits kill_timeout ms. If your app ignores SIGTERM, PM2 SIGKILLs it after the timeout and every in-flight request dies.

Your app must handle it — app.enableShutdownHooks() in NestJS (Level 12). Test it: start a request that takes 5 seconds, run pm2 reload api, and confirm it completes.

Environment variables and --update-env ​

pm2 restart DOES NOT PICK UP .env CHANGES

PM2 caches the environment from when the process was created. Editing .env and running pm2 restart api leaves the old values in place.

bash
pm2 reload ecosystem.config.cjs --update-env    # ✅ correct
pm2 restart api --update-env                     # ✅ also works
pm2 delete api && pm2 start ecosystem.config.cjs # ✅ guaranteed clean
pm2 restart api                                   # ❌ old environment

Verify:

bash
pm2 env 0 | grep -E "NODE_ENV|DATABASE_URL"

"I changed the database password and the app still uses the old one" is this, every single time.

Log management ​

PM2 writes to the files in your ecosystem config and does not rotate them by default.

PM2 LOGS WILL FILL YOUR DISK

A chatty app produces gigabytes over months. When the disk hits 100%: PostgreSQL stops accepting writes, Nginx cannot log, builds fail, and SSH logins may fail. Everything breaks at once, confusingly.

Install rotation immediately:

bash
pm2 install pm2-logrotate
pm2 set pm2-logrotate:max_size 10M
pm2 set pm2-logrotate:retain 14
pm2 set pm2-logrotate:compress true
pm2 set pm2-logrotate:rotateInterval '0 0 * * *'
pm2 set pm2-logrotate:workerInterval 300

Verify: pm2 conf pm2-logrotate and check ls -la ~/.pm2/logs/ after a day.

bash
# SERVER
du -sh /home/deploy/logs ~/.pm2/logs
pm2 flush                    # truncate all logs now
tail -f /home/deploy/logs/api-error.log

STRUCTURED LOGGING PAYS OFF QUICKLY

Plain console.log produces logs you can only grep. JSON logs can be queried:

ts
import { Logger } from 'nestjs-pino';
// { "level":"error", "time":..., "reqId":"...", "msg":"payment failed", "userId":"..." }
bash
pm2 logs api --raw --nostream | jq 'select(.level=="error")'
pm2 logs api --raw --nostream | jq -r 'select(.userId=="abc") | .msg'

Include a request ID in every log line so you can reconstruct a single request across services. (Level 21)

Memory management ​

js
max_memory_restart: '512M'

PM2 checks RSS periodically and restarts the instance when it exceeds the threshold. In cluster mode it restarts one instance at a time, so there is no downtime.

THIS IS A SAFETY NET, NOT A FIX

If max_memory_restart fires regularly, you have a memory leak. Find it:

bash
pm2 monit                             # watch memory climb
node --inspect=127.0.0.1:9229 dist/main.js    # locally, take heap snapshots

Common Node leaks: event listeners added per request without removal, an unbounded in-memory cache/Map, closures capturing large objects, and Prisma clients instantiated per request instead of once.

Set the threshold above normal peak but well below available RAM. If normal usage is 200 MB, 512M is reasonable; 2G on a 4 GB server means the OOM killer wins first.

Crash-loop protection ​

js
min_uptime: '30s',
max_restarts: 10,
exp_backoff_restart_delay: 200,

If the app crashes on startup — bad .env, database unreachable, port in use — PM2 would otherwise restart it forever, burning CPU and filling logs. With these settings PM2 gives up after 10 rapid failures and marks the process errored.

A STOPPED APP IS BETTER THAN A CRASH LOOP

errored status is visible in pm2 list and in monitoring. An infinite restart loop can look "alive" while serving nothing, and it drowns the real error in a flood of repeated stack traces.

When you see errored:

bash
pm2 describe api
pm2 logs api --err --lines 100 --nostream
pm2 reset api && pm2 restart api      # after fixing the cause

Health checks ​

ts
// NestJS
@Get('health')
health() {
  return {
    status: 'ok',
    uptime: process.uptime(),
    commit: process.env.GIT_COMMIT ?? 'unknown',
    pid: process.pid,
    instance: process.env.NODE_APP_INSTANCE,
  };
}
bash
# SERVER — after every deploy
curl -fsS http://127.0.0.1:3001/api/health | jq

USE wait_ready FOR ACCURATE ROLLING RELOADS

By default PM2 considers an instance ready when the process starts, which is before Nest has connected to the database. With wait_ready: true, PM2 waits for an explicit signal:

ts
await app.listen(env.PORT, '127.0.0.1');
if (process.send) process.send('ready');

Now pm2 reload only kills the old instance once the new one is genuinely serving traffic. Without this, there is a window during reload where requests hit a process that is up but not ready — occasional 502s during deploys.

PM2 in the deploy script ​

bash
#!/usr/bin/env bash
set -euo pipefail

APP_DIR="/home/deploy/apps/myapp"
export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"

cd "$APP_DIR"

git fetch origin --prune
git reset --hard origin/production
git clean -fd

pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend prisma migrate deploy
pnpm --filter backend build
pnpm --filter frontend build

export GIT_COMMIT=$(git rev-parse --short HEAD)
pm2 reload ecosystem.config.cjs --update-env
pm2 save

# Verify — fail loudly if the new version is not healthy
sleep 5
for i in {1..10}; do
  if curl -fsS http://127.0.0.1:3001/api/health > /dev/null; then
    echo "✅ API healthy"; break
  fi
  [ "$i" -eq 10 ] && { echo "❌ health check failed"; pm2 logs api --err --lines 50 --nostream; exit 1; }
  sleep 2
done

curl -fsS -o /dev/null http://127.0.0.1:3000/ && echo "✅ Web healthy"
echo "✅ Deployed $GIT_COMMIT"

pm2 save AFTER EVERY DEPLOY

If the deploy changed which processes exist (added an app, changed instance count), the saved dump is stale and a reboot restores the old configuration.

Rollback ​

bash
# SERVER — quick rollback with the in-place layout
cd /home/deploy/apps/myapp
git reset --hard <previous-commit-sha>
pnpm install --frozen-lockfile
pnpm --filter backend build && pnpm --filter frontend build
pm2 reload ecosystem.config.cjs --update-env

DATABASE MIGRATIONS DO NOT ROLL BACK WITH THE CODE

Reverting code to a commit before a migration leaves the database with the new schema and the code expecting the old one. Prisma has no automatic down-migrations.

This is why additive migrations matter (Level 9): if a migration only adds things, old code keeps working and rollback is safe. If it dropped a column, rolling back the code is not enough — you need a restore, which means downtime.

Rule: any migration that removes or renames something must be split across two deploys.

Instant rollback needs the release-directory pattern in Level 20.

Troubleshooting ​

ProblemCauseDiagnoseFix
pm2: command not found over SSHNVM not loaded in non-interactive shellssh server "which pm2"Symlink to /usr/local/bin (Level 6)
Apps gone after rebootpm2 save not run, or startup not installedsystemctl status pm2-deploypm2 startup + pm2 save, then test a reboot
Restart count climbingCrash looppm2 logs api --err --lines 100Read the error; check .env and DB connectivity
Status erroredHit max_restartspm2 describe apiFix cause, pm2 reset api && pm2 restart api
EADDRINUSEOld process still boundsudo ss -tulpn | grep 3001pm2 delete api; kill the stray PID
Env changes not appliedMissing --update-envpm2 env 0pm2 reload ... --update-env
Memory grows steadilyLeakpm2 monitHeap snapshots; max_memory_restart as a stopgap
502 during every deployNo graceful shutdown, or reload in fork modepm2 describe api — check modeCluster mode + enableShutdownHooks() + wait_ready
WebSockets flakyCluster mode without adapter/sticky sessionspm2 list — >1 instance?Redis adapter + sticky sessions
Disk fullPM2 logs unrotateddu -sh ~/.pm2/logspm2 install pm2-logrotate
Cron/scheduled task runs N timessetInterval in cluster mode—Move to a fork-mode worker or use a distributed lock
PM2 daemon itself goneKilled, or OOMpm2 pingpm2 resurrect
bash
# SERVER — general PM2 debugging
pm2 describe api
pm2 logs api --err --lines 200 --nostream
pm2 prettylist | jq '.[] | {name, status, restart_time, pm_uptime}'
cat ~/.pm2/pm2.log            # the PM2 daemon's own log
pm2 ping                       # is the daemon alive?
pm2 resurrect                  # restore from the saved dump
pm2 update                     # reload the in-memory daemon after a PM2 upgrade

Production Checklist — Level 13 ​

  • [ ] PM2 installed as deploy, not with sudo
  • [ ] All processes defined in a committed ecosystem.config.cjs
  • [ ] exec_mode: 'cluster' with instances >= 2 for HTTP apps
  • [ ] Instance count sized to available RAM, not blindly 'max'
  • [ ] NODE_ENV=production and HOST=127.0.0.1 set for both apps
  • [ ] watch: false
  • [ ] max_memory_restart set below available RAM
  • [ ] kill_timeout longer than the slowest request
  • [ ] min_uptime and max_restarts set for crash-loop protection
  • [ ] Graceful shutdown implemented and tested with a slow request during reload
  • [ ] wait_ready: true with process.send('ready') in the app
  • [ ] pm2 startup run and pm2 save executed
  • [ ] A real reboot has been tested and everything came back
  • [ ] pm2-logrotate installed and configured
  • [ ] Deploys use pm2 reload ... --update-env, never bare restart
  • [ ] Health endpoint exposing commit SHA, verified after every deploy
  • [ ] Redis-backed sessions, cache, and rate limiting (cluster-safe)
  • [ ] Socket.IO Redis adapter + sticky sessions configured
  • [ ] No setInterval cron jobs running in cluster mode
  • [ ] Rollback procedure documented, with awareness that migrations do not revert

Next: Level 14 — Nginx →