Level 20 — Deployment Strategies
How new code replaces old code. The choice determines your downtime, your rollback speed, and how much complexity you carry.
The four questions
Every strategy is an answer to:
- Is the site down during the switch?
- How fast can I go back?
- What happens to in-flight requests?
- How much complexity am I buying?
Strategy 1 — Simple in-place deployment
# SERVER
cd /home/deploy/apps/myapp
git fetch origin && git reset --hard origin/production
pnpm install --frozen-lockfile
pnpm prisma migrate deploy
pnpm build
pm2 reload ecosystem.config.cjs --update-env| Question | Answer |
|---|---|
| Downtime | None from pm2 reload, but the running app's files change underneath it during build |
| Rollback | Minutes — checkout the old commit and rebuild |
| In-flight requests | Safe during reload; unpredictable during the build window |
| Complexity | Minimal |
THE HIDDEN RISK OF IN-PLACE BUILDS
Between git reset and the end of pnpm build, the old process is still serving traffic from a directory whose files have already been replaced. Nuxt's SSR server lazily loads chunks; if it needs a chunk that was just overwritten by a partially-written new one, requests fail with cryptic module errors.
This window is 1–3 minutes. On a low-traffic site you may never notice. On a busy one you will see a burst of 500s in your logs on every deploy — which is exactly the confusing intermittent failure that sends people down the wrong debugging path.
WHEN THIS IS THE RIGHT CHOICE
Genuinely fine for: a personal project, an internal tool, a low-traffic site, or the first month of a new product. Do not add complexity you do not need.
Move on when: you see deploy-time 500s in your logs, a bad deploy takes more than a few minutes to undo, or you deploy often enough that the risk compounds.
Strategy 2 — Zero-downtime with PM2 reload
The same as above, but with the two things that make reload actually zero-downtime.
Requirement 1 — cluster mode
// ecosystem.config.cjs
{
name: 'api',
exec_mode: 'cluster',
instances: 2, // MUST be ≥ 2
}pm2 reload IN FORK MODE IS JUST A RESTART
With one process there is nothing to roll. PM2 silently restarts it and you get downtime while believing you do not. (Level 13)
Requirement 2 — graceful shutdown
// NestJS
const app = await NestFactory.create(AppModule);
app.enableShutdownHooks();
await app.listen(3001, '127.0.0.1');
if (process.send) process.send('ready'); // for wait_readyWITHOUT GRACEFUL SHUTDOWN, reload STILL DROPS REQUESTS
PM2 sends SIGTERM and waits kill_timeout (default 1.6s, set it to 10000). If your app ignores SIGTERM, PM2 SIGKILLs it and every in-flight request dies mid-response.
Test it, do not assume it:
# Terminal 1 — a slow endpoint
curl -w "\n%{http_code} in %{time_total}s\n" https://api.example.com/slow-endpoint
# Terminal 2 — immediately
pm2 reload apiThe request must complete with 200. If you get a connection reset, graceful shutdown is not working.
MIXED VERSIONS RUN SIMULTANEOUSLY DURING A ROLL
For several seconds, worker 1 runs v1 and worker 2 runs v2 — against the same database.
This is why every migration must be backward-compatible with the previous release. If v2's migration dropped a column that v1 still selects, half your requests fail during the roll. (Level 9)
Strategy 3 — Atomic releases with symlinks
The best value-for-complexity upgrade. Build into a new directory, then flip a symlink.
Directory layout
/home/deploy/apps/myapp/
├── current -> releases/20260811-143022-a1b2c3d # symlink
├── releases/
│ ├── 20260811-143022-a1b2c3d/ # current
│ ├── 20260811-101500-9f8e7d6/ # previous — instant rollback
│ ├── 20260810-183000-5c4b3a2/
│ └── ...
└── shared/
├── .env # survives every deploy
├── uploads/ # user files
└── logs/The deploy script
#!/usr/bin/env bash
# /home/deploy/apps/myapp/deploy.sh
set -euo pipefail
APP="/home/deploy/apps/myapp"
REPO="git@github.com:myorg/myapp.git"
BRANCH="production"
KEEP=5
export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"
mkdir -p "$APP/releases" "$APP/shared"
# --- 1. Fetch code into a NEW directory ---
TS=$(date -u +%Y%m%d-%H%M%S)
git -C "$APP/repo" fetch origin --prune 2>/dev/null || git clone --bare "$REPO" "$APP/repo"
SHA=$(git -C "$APP/repo" rev-parse --short "origin/$BRANCH")
REL="$APP/releases/${TS}-${SHA}"
echo "==> Creating release $REL"
mkdir -p "$REL"
git -C "$APP/repo" archive "origin/$BRANCH" | tar -x -C "$REL"
# --- 2. Link shared, mutable state ---
ln -sfn "$APP/shared/.env" "$REL/.env"
ln -sfn "$APP/shared/uploads" "$REL/uploads"
ln -sfn "$APP/shared/logs" "$REL/logs"
# --- 3. Build (old version still serving from `current`) ---
cd "$REL"
pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend build
pnpm --filter frontend build
# --- 4. Migrate ---
pnpm --filter backend prisma migrate deploy
# --- 5. ATOMIC SWITCH ---
ln -sfn "$REL" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current"
echo "==> Switched current -> $REL"
# --- 6. Reload ---
cd "$APP/current"
GIT_COMMIT="$SHA" pm2 reload ecosystem.config.cjs --update-env
pm2 save
# --- 7. Verify ---
for i in $(seq 1 15); do
curl -fsS http://127.0.0.1:3001/api/health >/dev/null && { echo "✅ healthy"; break; }
[ "$i" -eq 15 ] && { echo "❌ unhealthy — rolling back"; "$APP/rollback.sh"; exit 1; }
sleep 2
done
# --- 8. Prune old releases ---
ls -1dt "$APP/releases"/* | tail -n +$((KEEP + 1)) | xargs -r rm -rf
echo "==> Deployed $SHA"THE ATOMIC SYMLINK SWAP
ln -sfn "$REL" "$APP/current" # ❌ NOT atomicln -sfn on an existing symlink does unlink then symlink — there is a window, however brief, where current does not exist. A request arriving then fails.
ln -sfn "$REL" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current" # ✅ atomic rename(2)rename(2) is atomic at the filesystem level. current always points at exactly one valid release.
PM2 MUST BE CONFIGURED WITH THE current PATH — AND cwd RESOLUTION MATTERS
{ cwd: '/home/deploy/apps/myapp/current/backend' }Node resolves symlinks at require time, so a running process keeps using the release it started from — which is exactly what you want. But pm2 reload must re-read the config through the symlink to pick up the new release. Always run it from $APP/current as the script does.
If you see the old code still running after a switch, check pm2 describe api and confirm the script path reflects the new release.
Rollback in seconds
#!/usr/bin/env bash
# /home/deploy/apps/myapp/rollback.sh
set -euo pipefail
APP="/home/deploy/apps/myapp"
CURRENT=$(readlink -f "$APP/current")
PREV=$(ls -1dt "$APP/releases"/* | grep -v "^${CURRENT}$" | head -1)
[ -z "$PREV" ] && { echo "No previous release"; exit 1; }
echo "==> Rolling back from $(basename "$CURRENT") to $(basename "$PREV")"
ln -sfn "$PREV" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current"
cd "$APP/current"
pm2 reload ecosystem.config.cjs --update-env
pm2 save
sleep 3
curl -fsS http://127.0.0.1:3001/api/health && echo "✅ Rolled back to $(basename "$PREV")"ROLLBACK DOES NOT UNDO MIGRATIONS
This script restores the code in about five seconds. The database schema stays wherever the forward migration left it.
If the release you are rolling back from ran an additive migration, the old code ignores the new column and everything works.
If it ran a destructive migration (dropped a column, renamed a table, added NOT NULL), the old code breaks against the new schema. Your options are then: roll forward with a fix, or restore the database from backup and lose everything written since.
This is why the additive-only rule is not a style preference — it is what makes rollback possible at all. (Level 9)
| Question | Answer |
|---|---|
| Downtime | None — build happens off to the side |
| Rollback | ~5 seconds |
| In-flight requests | Safe (with graceful shutdown) |
| Complexity | Moderate — one script, one directory convention |
THIS IS THE SWEET SPOT FOR A SINGLE VPS
It gives you the two things that matter most — no mixed-state window during build, and near-instant rollback — for perhaps 60 lines of shell. Nothing beyond this is worth adding on one server.
Strategy 4 — Blue-green deployment
Two complete environments. One serves traffic; you deploy to the other and switch.
# SERVER — /etc/nginx/conf.d/upstream.conf
upstream app_backend { server 127.0.0.1:3001; } # blue
# upstream app_backend { server 127.0.0.1:4001; } # greenSwitching:
# SERVER
sed -i 's/127.0.0.1:3001/127.0.0.1:4001/' /etc/nginx/conf.d/upstream.conf
sudo nginx -t && sudo systemctl reload nginx| Question | Answer |
|---|---|
| Downtime | None |
| Rollback | Seconds — edit the upstream back and reload |
| In-flight requests | Safe — Nginx reload drains gracefully |
| Complexity | High — double the RAM, double the processes, two of everything |
BLUE-GREEN IS USUALLY THE WRONG CHOICE ON ONE VPS
It doubles your memory footprint — two Nuxt servers plus two NestJS servers is 600 MB–1 GB of RAM sitting idle. On a 4 GB box competing with PostgreSQL and Redis, that is a lot to pay.
And it does not solve the actual hard problem. Both environments share one database. A destructive migration breaks blue the moment green applies it, so you still need additive-only migrations — the same discipline the symlink strategy requires, at a fraction of the cost.
Blue-green earns its keep when: you need to run smoke tests against the new version on production infrastructure before any user reaches it, or you are on separate machines where the resource cost is not doubled on one box.
Strategy 5 — Rolling deployment
Update instances one at a time across multiple machines.
| Question | Answer |
|---|---|
| Downtime | None |
| Rollback | Roll the old version back through, instance by instance |
| In-flight requests | Safe, with health-check-aware load balancing |
| Complexity | High — needs a real load balancer and orchestration |
PM2 CLUSTER RELOAD IS A ROLLING DEPLOYMENT
On one server, pm2 reload already does exactly this across worker processes. You do not need anything more until you have multiple machines — at which point you are choosing between a load balancer with health checks, or Kubernetes.
Do not adopt Kubernetes to get rolling deployments. You already have them.
Strategy 6 — Container deployment
# docker-compose.prod.yml
services:
api:
image: ghcr.io/myorg/myapp-api:${TAG}
restart: unless-stopped
env_file: .env
ports: ["127.0.0.1:3001:3001"]
healthcheck:
test: ["CMD", "curl", "-fsS", "http://127.0.0.1:3001/api/health"]
interval: 10s
retries: 5# SERVER
export TAG=a1b2c3d
docker compose -f docker-compose.prod.yml pull
docker compose -f docker-compose.prod.yml up -dRollback is re-tagging:
export TAG=<previous-sha>
docker compose -f docker-compose.prod.yml up -d| Question | Answer |
|---|---|
| Downtime | Brief with plain Compose (up -d recreates); none with --scale-based rolling or a proxy that health-checks |
| Rollback | Seconds — change the tag |
| In-flight requests | Depends on stop_grace_period and signal handling (Level 11) |
| Complexity | Moderate–high — registry, image builds, tag management |
docker compose up -d IS NOT ZERO-DOWNTIME BY DEFAULT
It stops the old container, then starts the new one. There is a gap of a few seconds.
To avoid it you need either a proxy that health-checks (Traefik, Caddy) and switches only when the new container is ready, or a manual scale-up/scale-down dance. Plain Compose alone does not do rolling updates — that is what orchestrators are for.
THE HONEST CASE FOR CONTAINERS HERE
The real win is immutability: the artifact tested in CI is byte-for-byte what runs in production, and rollback is a tag change with no rebuild.
The real cost is a registry, image build times, and a second mental model for logs and debugging.
If your app is plain Node with no exotic system dependencies, the symlink strategy gives you most of the same benefits with far less machinery. Choose containers when you have a reason beyond aesthetics.
Comparison
| Strategy | Downtime | Rollback | Extra RAM | Complexity | Right for |
|---|---|---|---|---|---|
| Simple in-place | Build-window risk | Minutes | None | ⭐ | Personal projects, first month |
| PM2 reload | None (during reload) | Minutes | None | ⭐⭐ | Small production apps |
| Symlink releases | None | Seconds | ~1 release of disk | ⭐⭐⭐ | Single-VPS production — recommended |
| Blue-green | None | Seconds | 2× | ⭐⭐⭐⭐ | Pre-switch smoke testing on prod infra |
| Rolling | None | Minutes | Multi-server | ⭐⭐⭐⭐ | Several machines |
| Containers | Brief (plain Compose) | Seconds | Small | ⭐⭐⭐⭐ | Immutable artifacts, complex deps |
THE RECOMMENDED PATH
- Start: simple in-place. Ship the product.
- Soon after launch: PM2 cluster mode + graceful shutdown. Costs nothing, removes reload downtime.
- When deploys start to feel risky: symlink releases with a tested rollback script. This is where most single-VPS deployments should end up and stay.
- Only with a concrete reason: containers or multi-server.
Every step past 3 buys less than it costs on a single server.
Migration strategy — the constraint behind everything
Whatever deployment strategy you choose, this is the rule that makes rollback possible:
EVERY MIGRATION MUST BE BACKWARD-COMPATIBLE WITH THE PREVIOUS RELEASE
Safe (additive) — deploy freely:
CREATE TABLEADD COLUMNthat is nullable or has a defaultCREATE INDEX CONCURRENTLY- Adding an enum value
Unsafe (destructive) — must be split across two deploys:
DROP COLUMN/DROP TABLERENAME COLUMN/RENAME TABLEALTER COLUMN ... SET NOT NULLon existing dataALTER COLUMN ... TYPE(may rewrite and lock the table)- Removing an enum value
Expand / contract, worked through
Renaming users.name to users.full_name:
Deploy 1 — expand
ALTER TABLE users ADD COLUMN full_name TEXT;// Code writes BOTH, reads the old one
await prisma.user.update({ data: { name: v, full_name: v } });Backfill in the background:
UPDATE users SET full_name = name WHERE full_name IS NULL;Deploy 2 — switch reads
// Still writes both; now reads full_nameDeploy 3 — contract
// Stop writing `name`ALTER TABLE users DROP COLUMN name;Tedious, and every step is individually rollback-safe. A one-step rename is not.
HOW TO KNOW IF A MIGRATION IS SAFE
# LOCAL — read the generated SQL before merging
cat prisma/migrations/*/migration.sql | grep -iE "drop|rename|not null|alter column"Any hit means: split it, or accept that you cannot roll back this release. Make this a code-review checklist item.
Post-deployment verification
#!/usr/bin/env bash
# smoke-test.sh
set -euo pipefail
BASE_APP="https://app.example.com"
BASE_API="https://api.example.com"
check() {
local url=$1 expected=$2
local code
code=$(curl -fsS -o /dev/null -w "%{http_code}" --max-time 10 "$url" || echo "000")
if [ "$code" = "$expected" ]; then echo "✅ $url → $code"
else echo "❌ $url → $code (expected $expected)"; exit 1; fi
}
check "$BASE_APP/" 200
check "$BASE_API/api/health" 200
check "$BASE_APP/nonexistent" 404
# Is the deployed commit the one we expect?
DEPLOYED=$(curl -fsS "$BASE_API/api/health" | jq -r .commit)
echo "Deployed commit: $DEPLOYED"
[ "$DEPLOYED" = "${EXPECTED_SHA:-$DEPLOYED}" ] || { echo "❌ commit mismatch"; exit 1; }
# TLS still valid?
echo | openssl s_client -connect app.example.com:443 -servername app.example.com 2>/dev/null \
| openssl x509 -noout -checkend 604800 && echo "✅ cert valid >7 days"
echo "✅ All smoke tests passed"WATCH THE ERROR RATE FOR TEN MINUTES AFTER A DEPLOY
A smoke test proves the site responds. It does not prove the new code is correct under real traffic.
# SERVER
pm2 logs --err --lines 0 # stream new errors only
sudo tail -f /var/log/nginx/api.example.com.access.log | grep -E ' 5[0-9][0-9] 'Most deploy-caused incidents surface within the first few minutes, on a code path your smoke test does not touch. Staying to watch is the cheapest incident prevention there is.
Production Checklist — Level 20
- [ ] I know which strategy I am using and why
- [ ] PM2 in cluster mode with
instances >= 2 - [ ] Graceful shutdown implemented and tested with a slow request during reload
- [ ]
kill_timeoutlonger than the slowest request - [ ] Deploy script uses
set -euo pipefail - [ ] If using symlink releases: swap is
ln+mv -Tf, not bareln -sfn - [ ]
.env, uploads, and logs live inshared/and survive deploys - [ ] Old releases pruned, keeping at least 5
- [ ] Rollback script exists and has been executed at least once as a drill
- [ ] Every migration reviewed for destructive operations before merge
- [ ] Destructive schema changes split into expand/contract across releases
- [ ] Health check runs after every deploy and fails the pipeline loudly
- [ ] Automatic rollback on health-check failure, or a documented manual step
- [ ] Smoke test hits the public URLs, not just localhost
- [ ] Deployed commit SHA verifiable from the health endpoint
- [ ] I watch error rates for ~10 minutes after each production deploy