Skip to content

Level 20 — Deployment Strategies ​

How new code replaces old code. The choice determines your downtime, your rollback speed, and how much complexity you carry.

The four questions ​

Every strategy is an answer to:

  1. Is the site down during the switch?
  2. How fast can I go back?
  3. What happens to in-flight requests?
  4. How much complexity am I buying?

Strategy 1 — Simple in-place deployment ​

bash
# SERVER
cd /home/deploy/apps/myapp
git fetch origin && git reset --hard origin/production
pnpm install --frozen-lockfile
pnpm prisma migrate deploy
pnpm build
pm2 reload ecosystem.config.cjs --update-env
QuestionAnswer
DowntimeNone from pm2 reload, but the running app's files change underneath it during build
RollbackMinutes — checkout the old commit and rebuild
In-flight requestsSafe during reload; unpredictable during the build window
ComplexityMinimal

THE HIDDEN RISK OF IN-PLACE BUILDS

Between git reset and the end of pnpm build, the old process is still serving traffic from a directory whose files have already been replaced. Nuxt's SSR server lazily loads chunks; if it needs a chunk that was just overwritten by a partially-written new one, requests fail with cryptic module errors.

This window is 1–3 minutes. On a low-traffic site you may never notice. On a busy one you will see a burst of 500s in your logs on every deploy — which is exactly the confusing intermittent failure that sends people down the wrong debugging path.

WHEN THIS IS THE RIGHT CHOICE

Genuinely fine for: a personal project, an internal tool, a low-traffic site, or the first month of a new product. Do not add complexity you do not need.

Move on when: you see deploy-time 500s in your logs, a bad deploy takes more than a few minutes to undo, or you deploy often enough that the risk compounds.

Strategy 2 — Zero-downtime with PM2 reload ​

The same as above, but with the two things that make reload actually zero-downtime.

Requirement 1 — cluster mode ​

js
// ecosystem.config.cjs
{
  name: 'api',
  exec_mode: 'cluster',
  instances: 2,          // MUST be ≥ 2
}

pm2 reload IN FORK MODE IS JUST A RESTART

With one process there is nothing to roll. PM2 silently restarts it and you get downtime while believing you do not. (Level 13)

Requirement 2 — graceful shutdown ​

ts
// NestJS
const app = await NestFactory.create(AppModule);
app.enableShutdownHooks();
await app.listen(3001, '127.0.0.1');
if (process.send) process.send('ready');   // for wait_ready

WITHOUT GRACEFUL SHUTDOWN, reload STILL DROPS REQUESTS

PM2 sends SIGTERM and waits kill_timeout (default 1.6s, set it to 10000). If your app ignores SIGTERM, PM2 SIGKILLs it and every in-flight request dies mid-response.

Test it, do not assume it:

bash
# Terminal 1 — a slow endpoint
curl -w "\n%{http_code} in %{time_total}s\n" https://api.example.com/slow-endpoint
# Terminal 2 — immediately
pm2 reload api

The request must complete with 200. If you get a connection reset, graceful shutdown is not working.

MIXED VERSIONS RUN SIMULTANEOUSLY DURING A ROLL

For several seconds, worker 1 runs v1 and worker 2 runs v2 — against the same database.

This is why every migration must be backward-compatible with the previous release. If v2's migration dropped a column that v1 still selects, half your requests fail during the roll. (Level 9)

The best value-for-complexity upgrade. Build into a new directory, then flip a symlink.

Directory layout ​

/home/deploy/apps/myapp/
├── current -> releases/20260811-143022-a1b2c3d      # symlink
├── releases/
│   ├── 20260811-143022-a1b2c3d/     # current
│   ├── 20260811-101500-9f8e7d6/     # previous — instant rollback
│   ├── 20260810-183000-5c4b3a2/
│   └── ...
└── shared/
    ├── .env                          # survives every deploy
    ├── uploads/                      # user files
    └── logs/

The deploy script ​

bash
#!/usr/bin/env bash
# /home/deploy/apps/myapp/deploy.sh
set -euo pipefail

APP="/home/deploy/apps/myapp"
REPO="git@github.com:myorg/myapp.git"
BRANCH="production"
KEEP=5

export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"

mkdir -p "$APP/releases" "$APP/shared"

# --- 1. Fetch code into a NEW directory ---
TS=$(date -u +%Y%m%d-%H%M%S)
git -C "$APP/repo" fetch origin --prune 2>/dev/null || git clone --bare "$REPO" "$APP/repo"
SHA=$(git -C "$APP/repo" rev-parse --short "origin/$BRANCH")
REL="$APP/releases/${TS}-${SHA}"

echo "==> Creating release $REL"
mkdir -p "$REL"
git -C "$APP/repo" archive "origin/$BRANCH" | tar -x -C "$REL"

# --- 2. Link shared, mutable state ---
ln -sfn "$APP/shared/.env"     "$REL/.env"
ln -sfn "$APP/shared/uploads"  "$REL/uploads"
ln -sfn "$APP/shared/logs"     "$REL/logs"

# --- 3. Build (old version still serving from `current`) ---
cd "$REL"
pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend build
pnpm --filter frontend build

# --- 4. Migrate ---
pnpm --filter backend prisma migrate deploy

# --- 5. ATOMIC SWITCH ---
ln -sfn "$REL" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current"
echo "==> Switched current -> $REL"

# --- 6. Reload ---
cd "$APP/current"
GIT_COMMIT="$SHA" pm2 reload ecosystem.config.cjs --update-env
pm2 save

# --- 7. Verify ---
for i in $(seq 1 15); do
  curl -fsS http://127.0.0.1:3001/api/health >/dev/null && { echo "✅ healthy"; break; }
  [ "$i" -eq 15 ] && { echo "❌ unhealthy — rolling back"; "$APP/rollback.sh"; exit 1; }
  sleep 2
done

# --- 8. Prune old releases ---
ls -1dt "$APP/releases"/* | tail -n +$((KEEP + 1)) | xargs -r rm -rf
echo "==> Deployed $SHA"

THE ATOMIC SYMLINK SWAP

bash
ln -sfn "$REL" "$APP/current"          # ❌ NOT atomic

ln -sfn on an existing symlink does unlink then symlink — there is a window, however brief, where current does not exist. A request arriving then fails.

bash
ln -sfn "$REL" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current"    # ✅ atomic rename(2)

rename(2) is atomic at the filesystem level. current always points at exactly one valid release.

PM2 MUST BE CONFIGURED WITH THE current PATH — AND cwd RESOLUTION MATTERS

js
{ cwd: '/home/deploy/apps/myapp/current/backend' }

Node resolves symlinks at require time, so a running process keeps using the release it started from — which is exactly what you want. But pm2 reload must re-read the config through the symlink to pick up the new release. Always run it from $APP/current as the script does.

If you see the old code still running after a switch, check pm2 describe api and confirm the script path reflects the new release.

Rollback in seconds ​

bash
#!/usr/bin/env bash
# /home/deploy/apps/myapp/rollback.sh
set -euo pipefail
APP="/home/deploy/apps/myapp"

CURRENT=$(readlink -f "$APP/current")
PREV=$(ls -1dt "$APP/releases"/* | grep -v "^${CURRENT}$" | head -1)

[ -z "$PREV" ] && { echo "No previous release"; exit 1; }

echo "==> Rolling back from $(basename "$CURRENT") to $(basename "$PREV")"
ln -sfn "$PREV" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current"

cd "$APP/current"
pm2 reload ecosystem.config.cjs --update-env
pm2 save

sleep 3
curl -fsS http://127.0.0.1:3001/api/health && echo "✅ Rolled back to $(basename "$PREV")"

ROLLBACK DOES NOT UNDO MIGRATIONS

This script restores the code in about five seconds. The database schema stays wherever the forward migration left it.

If the release you are rolling back from ran an additive migration, the old code ignores the new column and everything works.

If it ran a destructive migration (dropped a column, renamed a table, added NOT NULL), the old code breaks against the new schema. Your options are then: roll forward with a fix, or restore the database from backup and lose everything written since.

This is why the additive-only rule is not a style preference — it is what makes rollback possible at all. (Level 9)

QuestionAnswer
DowntimeNone — build happens off to the side
Rollback~5 seconds
In-flight requestsSafe (with graceful shutdown)
ComplexityModerate — one script, one directory convention

THIS IS THE SWEET SPOT FOR A SINGLE VPS

It gives you the two things that matter most — no mixed-state window during build, and near-instant rollback — for perhaps 60 lines of shell. Nothing beyond this is worth adding on one server.

Strategy 4 — Blue-green deployment ​

Two complete environments. One serves traffic; you deploy to the other and switch.

bash
# SERVER — /etc/nginx/conf.d/upstream.conf
upstream app_backend { server 127.0.0.1:3001; }   # blue
# upstream app_backend { server 127.0.0.1:4001; } # green

Switching:

bash
# SERVER
sed -i 's/127.0.0.1:3001/127.0.0.1:4001/' /etc/nginx/conf.d/upstream.conf
sudo nginx -t && sudo systemctl reload nginx
QuestionAnswer
DowntimeNone
RollbackSeconds — edit the upstream back and reload
In-flight requestsSafe — Nginx reload drains gracefully
ComplexityHigh — double the RAM, double the processes, two of everything

BLUE-GREEN IS USUALLY THE WRONG CHOICE ON ONE VPS

It doubles your memory footprint — two Nuxt servers plus two NestJS servers is 600 MB–1 GB of RAM sitting idle. On a 4 GB box competing with PostgreSQL and Redis, that is a lot to pay.

And it does not solve the actual hard problem. Both environments share one database. A destructive migration breaks blue the moment green applies it, so you still need additive-only migrations — the same discipline the symlink strategy requires, at a fraction of the cost.

Blue-green earns its keep when: you need to run smoke tests against the new version on production infrastructure before any user reaches it, or you are on separate machines where the resource cost is not doubled on one box.

Strategy 5 — Rolling deployment ​

Update instances one at a time across multiple machines.

QuestionAnswer
DowntimeNone
RollbackRoll the old version back through, instance by instance
In-flight requestsSafe, with health-check-aware load balancing
ComplexityHigh — needs a real load balancer and orchestration

PM2 CLUSTER RELOAD IS A ROLLING DEPLOYMENT

On one server, pm2 reload already does exactly this across worker processes. You do not need anything more until you have multiple machines — at which point you are choosing between a load balancer with health checks, or Kubernetes.

Do not adopt Kubernetes to get rolling deployments. You already have them.

Strategy 6 — Container deployment ​

yaml
# docker-compose.prod.yml
services:
  api:
    image: ghcr.io/myorg/myapp-api:${TAG}
    restart: unless-stopped
    env_file: .env
    ports: ["127.0.0.1:3001:3001"]
    healthcheck:
      test: ["CMD", "curl", "-fsS", "http://127.0.0.1:3001/api/health"]
      interval: 10s
      retries: 5
bash
# SERVER
export TAG=a1b2c3d
docker compose -f docker-compose.prod.yml pull
docker compose -f docker-compose.prod.yml up -d

Rollback is re-tagging:

bash
export TAG=<previous-sha>
docker compose -f docker-compose.prod.yml up -d
QuestionAnswer
DowntimeBrief with plain Compose (up -d recreates); none with --scale-based rolling or a proxy that health-checks
RollbackSeconds — change the tag
In-flight requestsDepends on stop_grace_period and signal handling (Level 11)
ComplexityModerate–high — registry, image builds, tag management

docker compose up -d IS NOT ZERO-DOWNTIME BY DEFAULT

It stops the old container, then starts the new one. There is a gap of a few seconds.

To avoid it you need either a proxy that health-checks (Traefik, Caddy) and switches only when the new container is ready, or a manual scale-up/scale-down dance. Plain Compose alone does not do rolling updates — that is what orchestrators are for.

THE HONEST CASE FOR CONTAINERS HERE

The real win is immutability: the artifact tested in CI is byte-for-byte what runs in production, and rollback is a tag change with no rebuild.

The real cost is a registry, image build times, and a second mental model for logs and debugging.

If your app is plain Node with no exotic system dependencies, the symlink strategy gives you most of the same benefits with far less machinery. Choose containers when you have a reason beyond aesthetics.

Comparison ​

StrategyDowntimeRollbackExtra RAMComplexityRight for
Simple in-placeBuild-window riskMinutesNone⭐Personal projects, first month
PM2 reloadNone (during reload)MinutesNone⭐⭐Small production apps
Symlink releasesNoneSeconds~1 release of disk⭐⭐⭐Single-VPS production — recommended
Blue-greenNoneSeconds2×⭐⭐⭐⭐Pre-switch smoke testing on prod infra
RollingNoneMinutesMulti-server⭐⭐⭐⭐Several machines
ContainersBrief (plain Compose)SecondsSmall⭐⭐⭐⭐Immutable artifacts, complex deps

THE RECOMMENDED PATH

  1. Start: simple in-place. Ship the product.
  2. Soon after launch: PM2 cluster mode + graceful shutdown. Costs nothing, removes reload downtime.
  3. When deploys start to feel risky: symlink releases with a tested rollback script. This is where most single-VPS deployments should end up and stay.
  4. Only with a concrete reason: containers or multi-server.

Every step past 3 buys less than it costs on a single server.

Migration strategy — the constraint behind everything ​

Whatever deployment strategy you choose, this is the rule that makes rollback possible:

EVERY MIGRATION MUST BE BACKWARD-COMPATIBLE WITH THE PREVIOUS RELEASE

Safe (additive) — deploy freely:

  • CREATE TABLE
  • ADD COLUMN that is nullable or has a default
  • CREATE INDEX CONCURRENTLY
  • Adding an enum value

Unsafe (destructive) — must be split across two deploys:

  • DROP COLUMN / DROP TABLE
  • RENAME COLUMN / RENAME TABLE
  • ALTER COLUMN ... SET NOT NULL on existing data
  • ALTER COLUMN ... TYPE (may rewrite and lock the table)
  • Removing an enum value

Expand / contract, worked through ​

Renaming users.name to users.full_name:

Deploy 1 — expand

sql
ALTER TABLE users ADD COLUMN full_name TEXT;
ts
// Code writes BOTH, reads the old one
await prisma.user.update({ data: { name: v, full_name: v } });

Backfill in the background:

sql
UPDATE users SET full_name = name WHERE full_name IS NULL;

Deploy 2 — switch reads

ts
// Still writes both; now reads full_name

Deploy 3 — contract

ts
// Stop writing `name`
sql
ALTER TABLE users DROP COLUMN name;

Tedious, and every step is individually rollback-safe. A one-step rename is not.

HOW TO KNOW IF A MIGRATION IS SAFE

bash
# LOCAL — read the generated SQL before merging
cat prisma/migrations/*/migration.sql | grep -iE "drop|rename|not null|alter column"

Any hit means: split it, or accept that you cannot roll back this release. Make this a code-review checklist item.

Post-deployment verification ​

bash
#!/usr/bin/env bash
# smoke-test.sh
set -euo pipefail

BASE_APP="https://app.example.com"
BASE_API="https://api.example.com"

check() {
  local url=$1 expected=$2
  local code
  code=$(curl -fsS -o /dev/null -w "%{http_code}" --max-time 10 "$url" || echo "000")
  if [ "$code" = "$expected" ]; then echo "✅ $url → $code"
  else echo "❌ $url → $code (expected $expected)"; exit 1; fi
}

check "$BASE_APP/"              200
check "$BASE_API/api/health"    200
check "$BASE_APP/nonexistent"   404

# Is the deployed commit the one we expect?
DEPLOYED=$(curl -fsS "$BASE_API/api/health" | jq -r .commit)
echo "Deployed commit: $DEPLOYED"
[ "$DEPLOYED" = "${EXPECTED_SHA:-$DEPLOYED}" ] || { echo "❌ commit mismatch"; exit 1; }

# TLS still valid?
echo | openssl s_client -connect app.example.com:443 -servername app.example.com 2>/dev/null \
  | openssl x509 -noout -checkend 604800 && echo "✅ cert valid >7 days"

echo "✅ All smoke tests passed"

WATCH THE ERROR RATE FOR TEN MINUTES AFTER A DEPLOY

A smoke test proves the site responds. It does not prove the new code is correct under real traffic.

bash
# SERVER
pm2 logs --err --lines 0            # stream new errors only
sudo tail -f /var/log/nginx/api.example.com.access.log | grep -E ' 5[0-9][0-9] '

Most deploy-caused incidents surface within the first few minutes, on a code path your smoke test does not touch. Staying to watch is the cheapest incident prevention there is.

Production Checklist — Level 20 ​

  • [ ] I know which strategy I am using and why
  • [ ] PM2 in cluster mode with instances >= 2
  • [ ] Graceful shutdown implemented and tested with a slow request during reload
  • [ ] kill_timeout longer than the slowest request
  • [ ] Deploy script uses set -euo pipefail
  • [ ] If using symlink releases: swap is ln + mv -Tf, not bare ln -sfn
  • [ ] .env, uploads, and logs live in shared/ and survive deploys
  • [ ] Old releases pruned, keeping at least 5
  • [ ] Rollback script exists and has been executed at least once as a drill
  • [ ] Every migration reviewed for destructive operations before merge
  • [ ] Destructive schema changes split into expand/contract across releases
  • [ ] Health check runs after every deploy and fails the pipeline loudly
  • [ ] Automatic rollback on health-check failure, or a documented manual step
  • [ ] Smoke test hits the public URLs, not just localhost
  • [ ] Deployed commit SHA verifiable from the health endpoint
  • [ ] I watch error rates for ~10 minutes after each production deploy

Next: Level 21 — Logging and Monitoring →