Skip to content

Level 25 — Complete From-Zero Deployment ​

One continuous runbook. Fresh Ubuntu 24.04 VPS → production Nuxt + NestJS application with HTTPS, PostgreSQL, Redis, PM2, firewall, fail2ban, CI/CD, backups, and monitoring.

Assumptions:

ServerFresh Ubuntu 24.04 LTS, 2 vCPU / 4 GB / 40 GB
IP203.0.113.10 (replace with yours)
Domainexample.com
Frontendapp.example.com
Backendapi.example.com
Repositorygit@github.com:myorg/myapp.git
Deploy userdeploy

Estimated time: 60–90 minutes, plus DNS propagation.

WORK INSIDE tmux

Install tmux early and run everything inside it. If your connection drops mid-apt upgrade or mid-migration, the process survives and you reattach. See Phase 2.


Phase 0 — Before you touch the server ​

bash
# LOCAL — 1. Generate an SSH key if you do not have one
ssh-keygen -t ed25519 -C "aziz@laptop-2026" -f ~/.ssh/id_ed25519
cat ~/.ssh/id_ed25519.pub

2. Create the VPS, pasting that public key into the provider's SSH-keys field. Choose Ubuntu 24.04 LTS.

3. Configure the cloud firewall: inbound TCP 22, 80, 443 from anywhere; ICMP allowed; outbound all.

4. Create DNS records — do this now so propagation happens while you work:

TypeNameValueTTL
A@203.0.113.10300
Aapp203.0.113.10300
Aapi203.0.113.10300
Awww203.0.113.10300
CAA@0 issue "letsencrypt.org"3600
bash
# LOCAL — 5. Verify DNS (may take a few minutes)
dig app.example.com +short
dig api.example.com +short

6. Locate your provider's web console button. You may need it. Find it now, not during a lockout.


Phase 1 — First connection ​

bash
# LOCAL
ssh root@203.0.113.10

Accept the host key after checking the fingerprint against your provider's console.

bash
# SERVER
hostnamectl
lsb_release -a          # confirm Ubuntu 24.04
df -h && free -h        # confirm the resources you paid for

Phase 2 — System update and basics ​

bash
# SERVER
apt update && apt upgrade -y && apt autoremove -y

apt install -y \
  tmux curl wget git htop ncdu jq unzip ca-certificates \
  gnupg lsb-release software-properties-common \
  ufw fail2ban unattended-upgrades apt-listchanges

hostnamectl set-hostname prod-app-01
echo "127.0.1.1 prod-app-01" >> /etc/hosts
timedatectl set-timezone UTC
timedatectl
bash
# SERVER — start working inside tmux from here on
tmux new -s setup
bash
# SERVER — reboot if the kernel was updated
[ -f /var/run/reboot-required ] && { echo "rebooting"; reboot; }

Reconnect after ~45 seconds if it rebooted, then tmux attach -t setup.


Phase 3 — Create the deploy user ​

bash
# SERVER — as root
adduser deploy

Set a strong password and save it in your password manager — you may need it for sudo from the web console during a lockout.

bash
# SERVER
usermod -aG sudo deploy

mkdir -p /home/deploy/.ssh
cp /root/.ssh/authorized_keys /home/deploy/.ssh/authorized_keys
chown -R deploy:deploy /home/deploy/.ssh
chmod 700 /home/deploy/.ssh
chmod 600 /home/deploy/.ssh/authorized_keys
chmod 755 /home/deploy

groups deploy

TEST THE NEW USER IN A SECOND TERMINAL BEFORE PROCEEDING

bash
# LOCAL — NEW terminal, keep the root session open
ssh deploy@203.0.113.10
whoami          # deploy
sudo whoami     # root (after entering deploy's password)

Do not continue until both succeed. The root session stays open as your safety net through Phase 4.


Phase 4 — SSH hardening ​

bash
# SERVER — as deploy now
sudo nano /etc/ssh/sshd_config.d/99-hardening.conf
sshconfig
PermitRootLogin no
PasswordAuthentication no
PubkeyAuthentication yes
AuthenticationMethods publickey
PermitEmptyPasswords no
KbdInteractiveAuthentication no

AllowUsers deploy

MaxAuthTries 3
MaxSessions 5
LoginGraceTime 20
MaxStartups 10:30:60

X11Forwarding no
AllowAgentForwarding no
AllowTcpForwarding yes          # keep yes: needed for SSH tunnels to PostgreSQL
PermitUserEnvironment no

ClientAliveInterval 300
ClientAliveCountMax 2

LogLevel VERBOSE
bash
# SERVER
sudo sshd -t                                    # must be silent
sudo sshd -T | grep -Ei "permitroot|passwordauth|allowusers|maxauth"
sudo systemctl reload ssh
bash
# LOCAL — NEW terminal; verify before closing anything
ssh deploy@203.0.113.10                                        # ✅ works
ssh root@203.0.113.10                                          # ❌ rejected
ssh -o PreferredAuthentications=password deploy@203.0.113.10   # ❌ Permission denied (publickey)

Only now close the root session.


Phase 5 — Firewall ​

bash
# SERVER
sudo ufw default deny incoming
sudo ufw default allow outgoing
sudo ufw allow OpenSSH
sudo ufw limit ssh/tcp
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp

sudo ufw show added        # ⚠️ confirm SSH is listed BEFORE enabling
sudo ufw enable
sudo ufw status verbose

You should see IPv4 and (v6) rules for 22, 80, and 443 — nothing else.


Phase 6 — fail2ban ​

bash
# SERVER
sudo nano /etc/fail2ban/jail.local
ini
[DEFAULT]
ignoreip = 127.0.0.1/8 ::1
bantime  = 1h
findtime = 10m
maxretry = 5
bantime.increment = true
bantime.factor    = 2
bantime.maxtime   = 1w
banaction = nftables-multiport
banaction_allports = nftables-allports
backend = systemd

[sshd]
enabled  = true
port     = ssh
maxretry = 3
bantime  = 1h
bash
# SERVER
sudo systemctl enable --now fail2ban
sudo fail2ban-client status
sudo fail2ban-client status sshd     # ⚠️ must show the jail as active

IF THE sshd JAIL FAILS TO START

Ubuntu 24.04 uses nftables; the package default banaction sometimes mismatches. The config above sets nftables-multiport. If it still fails: sudo apt install -y iptables and restart fail2ban. A jail that failed to start is silently useless — always check.


Phase 7 — Automatic security updates ​

bash
# SERVER
sudo dpkg-reconfigure -plow unattended-upgrades      # answer Yes
sudo nano /etc/apt/apt.conf.d/50unattended-upgrades
Unattended-Upgrade::Allowed-Origins {
    "${distro_id}:${distro_codename}-security";
    "${distro_id}ESMApps:${distro_codename}-apps-security";
    "${distro_id}ESM:${distro_codename}-infra-security";
};
Unattended-Upgrade::Package-Blacklist { "postgresql-16"; "nginx"; };
Unattended-Upgrade::Remove-Unused-Kernel-Packages "true";
Unattended-Upgrade::Automatic-Reboot "false";
Unattended-Upgrade::Mail "you@example.com";
Unattended-Upgrade::MailReport "on-change";
bash
# SERVER
sudo unattended-upgrades --dry-run --debug | tail -20
systemctl is-enabled unattended-upgrades

Automatic-Reboot "false" FOR NOW

Set it to true only after you have completed Phase 17 and verified a full reboot brings everything back.


Phase 8 — Swap ​

bash
# SERVER
sudo fallocate -l 2G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
echo 'vm.swappiness=10' | sudo tee /etc/sysctl.d/99-swappiness.conf
sudo sysctl -p /etc/sysctl.d/99-swappiness.conf
free -h

Phase 9 — Node.js, pnpm, PM2 ​

bash
# SERVER — as deploy, NOT root
curl -o- https://raw.githubusercontent.com/nvm-sh/nvm/v0.40.1/install.sh | bash
source ~/.bashrc
nvm --version

nvm install --lts
nvm alias default lts/*
node -v && npm -v

corepack enable
corepack prepare pnpm@latest --activate
pnpm -v

npm install -g pm2
pm2 -v

Make the binaries visible to non-interactive shells ​

bash
# SERVER — move NVM loading ABOVE the interactive guard in .bashrc
nano ~/.bashrc

Put this at the very top of the file:

bash
export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && \. "$NVM_DIR/nvm.sh"
bash
# SERVER — and symlink into a system PATH
for b in node npm npx pnpm pm2; do
  sudo ln -sf "$(command -v $b)" "/usr/local/bin/$b"
done

VERIFY THIS NOW — IT IS THE #1 CI/CD FAILURE

bash
# LOCAL — from your laptop
ssh deploy@203.0.113.10 "node -v && pnpm -v && pm2 -v && which node pnpm pm2"

All three must print versions. If any says "command not found", your CI/CD pipeline will fail identically in Phase 16. Fix it here. (Level 6)


Phase 10 — PostgreSQL ​

bash
# SERVER
sudo apt install -y postgresql postgresql-contrib
sudo systemctl enable --now postgresql
sudo systemctl status postgresql

# ⚠️ MUST show 127.0.0.1, not 0.0.0.0
sudo ss -tulpn | grep 5432
bash
# SERVER — generate a URL-safe password and SAVE IT
openssl rand -base64 32 | tr -d '/+=' | head -c 32; echo
bash
# SERVER
sudo -u postgres psql
sql
CREATE USER myapp WITH PASSWORD 'PASTE-THE-GENERATED-PASSWORD';
CREATE DATABASE myapp_production OWNER myapp;
\c myapp_production
REVOKE ALL ON SCHEMA public FROM PUBLIC;
GRANT ALL ON SCHEMA public TO myapp;
ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO myapp;
ALTER DEFAULT PRIVILEGES IN SCHEMA public GRANT USAGE, SELECT ON SEQUENCES TO myapp;
\q

Tune for 4 GB ​

bash
# SERVER
sudo nano /etc/postgresql/16/main/postgresql.conf
conf
listen_addresses = 'localhost'
max_connections = 100
shared_buffers = 1GB
effective_cache_size = 3GB
work_mem = 16MB
maintenance_work_mem = 256MB
random_page_cost = 1.1
effective_io_concurrency = 200
max_wal_size = 2GB
log_min_duration_statement = 1000
log_checkpoints = on
log_connections = on
log_lock_waits = on
shared_preload_libraries = 'pg_stat_statements'
bash
# SERVER
sudo grep -E "^(local|host)" /etc/postgresql/16/main/pg_hba.conf     # confirm no `trust`
sudo systemctl restart postgresql
sudo -u postgres psql -d myapp_production -c "CREATE EXTENSION IF NOT EXISTS pg_stat_statements;"

# Protect the database from the OOM killer
sudo systemctl edit postgresql --force --full 2>/dev/null || true
sudo mkdir -p /etc/systemd/system/postgresql@16-main.service.d
echo -e "[Service]\nOOMScoreAdjust=-500" | sudo tee /etc/systemd/system/postgresql@16-main.service.d/oom.conf
sudo systemctl daemon-reload && sudo systemctl restart postgresql

# Verify the app can connect
psql "postgresql://myapp:PASSWORD@127.0.0.1:5432/myapp_production" -c "SELECT current_user, current_database();"

Phase 11 — Redis ​

bash
# SERVER
sudo apt install -y redis-server
openssl rand -base64 32 | tr -d '/+=' | head -c 32; echo     # save this
sudo nano /etc/redis/redis.conf

Set or confirm:

conf
bind 127.0.0.1 -::1
port 6379
protected-mode yes
requirepass PASTE-THE-GENERATED-PASSWORD

rename-command CONFIG ""
rename-command FLUSHALL ""
rename-command FLUSHDB ""
rename-command KEYS ""

maxmemory 512mb
maxmemory-policy volatile-lru

appendonly yes
appendfsync everysec

supervised systemd
bash
# SERVER
sudo chmod 640 /etc/redis/redis.conf
sudo chown redis:redis /etc/redis/redis.conf
sudo systemctl enable --now redis-server
sudo systemctl restart redis-server

sudo ss -tulpn | grep 6379          # ⚠️ must be 127.0.0.1
redis-cli ping                       # ⚠️ must return: NOAUTH Authentication required
REDISCLI_AUTH="YOUR-PASSWORD" redis-cli ping    # PONG

volatile-lru NOT allkeys-lru

This instance will hold sessions and possibly queues alongside cache. allkeys-lru would evict queued jobs. volatile-lru only evicts keys that have a TTL — so give cache keys a TTL and durable keys none. (Level 10)


Phase 12 — Nginx ​

bash
# SERVER
sudo apt install -y nginx
sudo systemctl enable --now nginx
sudo rm -f /etc/nginx/sites-enabled/default
curl -I http://127.0.0.1

Global config ​

bash
# SERVER
sudo nano /etc/nginx/nginx.conf

Inside http { } add or adjust:

nginx
    server_tokens off;
    client_max_body_size 20M;
    client_body_timeout 15s;
    client_header_timeout 15s;

    log_format main '$remote_addr - $remote_user [$time_local] "$request" '
                    '$status $body_bytes_sent "$http_referer" '
                    '"$http_user_agent" rt=$request_time urt="$upstream_response_time"';
    access_log /var/log/nginx/access.log main;

    gzip on;
    gzip_vary on;
    gzip_comp_level 6;
    gzip_min_length 1024;
    gzip_types text/plain text/css text/xml text/javascript
               application/json application/javascript application/xml+rss image/svg+xml;

    limit_req_zone  $binary_remote_addr zone=general:10m rate=30r/s;
    limit_req_zone  $binary_remote_addr zone=login:10m   rate=5r/m;
    limit_conn_zone $binary_remote_addr zone=addr:10m;
    limit_req_status 429;
    limit_conn_status 429;

Snippets ​

bash
# SERVER
sudo tee /etc/nginx/conf.d/websocket.conf > /dev/null <<'EOF'
map $http_upgrade $connection_upgrade {
    default upgrade;
    ''      close;
}
EOF

sudo tee /etc/nginx/snippets/proxy-params.conf > /dev/null <<'EOF'
proxy_http_version 1.1;
proxy_set_header Host              $host;
proxy_set_header X-Real-IP         $remote_addr;
proxy_set_header X-Forwarded-For   $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_set_header X-Forwarded-Host  $host;
proxy_set_header Upgrade           $http_upgrade;
proxy_set_header Connection        $connection_upgrade;
proxy_connect_timeout 10s;
proxy_send_timeout    60s;
proxy_read_timeout    60s;
proxy_buffering on;
proxy_buffer_size       16k;
proxy_buffers        8  16k;
proxy_redirect off;
EOF

sudo tee /etc/nginx/snippets/security-headers.conf > /dev/null <<'EOF'
add_header X-Frame-Options           "SAMEORIGIN"                              always;
add_header X-Content-Type-Options    "nosniff"                                 always;
add_header Referrer-Policy           "strict-origin-when-cross-origin"         always;
add_header Permissions-Policy        "camera=(), microphone=(), geolocation=()" always;
add_header Strict-Transport-Security "max-age=300"                             always;
EOF

sudo mkdir -p /var/www/certbot

HSTS STARTS AT max-age=300

Five minutes. Raise it to 31536000 only after HTTPS has been stable for a week. HSTS is effectively irreversible once cached (Level 16).

HTTP-only sites (Certbot adds TLS in Phase 14) ​

bash
# SERVER
sudo tee /etc/nginx/sites-available/app.example.com > /dev/null <<'EOF'
upstream nuxt_backend {
    server 127.0.0.1:3000 max_fails=3 fail_timeout=10s;
    keepalive 32;
}

server {
    listen 80;
    listen [::]:80;
    server_name app.example.com www.example.com example.com;

    location /.well-known/acme-challenge/ { root /var/www/certbot; }

    access_log /var/log/nginx/app.example.com.access.log main;
    error_log  /var/log/nginx/app.example.com.error.log warn;

    root /home/deploy/apps/myapp/current/frontend/.output/public;

    location ^~ /_nuxt/ {
        expires 1y;
        add_header Cache-Control "public, max-age=31536000, immutable" always;
        include snippets/security-headers.conf;
        access_log off;
        try_files $uri =404;
    }

    location ~* \.(ico|css|js|gif|jpe?g|png|webp|avif|svg|woff2?|ttf|map)$ {
        expires 30d;
        add_header Cache-Control "public, max-age=2592000" always;
        include snippets/security-headers.conf;
        access_log off;
        try_files $uri @nuxt;
    }

    location / {
        limit_req zone=general burst=50 nodelay;
        limit_conn addr 20;
        include snippets/proxy-params.conf;
        include snippets/security-headers.conf;
        proxy_pass http://nuxt_backend;
    }

    location @nuxt {
        include snippets/proxy-params.conf;
        proxy_pass http://nuxt_backend;
    }
}
EOF

sudo tee /etc/nginx/sites-available/api.example.com > /dev/null <<'EOF'
upstream nest_backend {
    server 127.0.0.1:3001 max_fails=3 fail_timeout=10s;
    keepalive 32;
}

server {
    listen 80;
    listen [::]:80;
    server_name api.example.com;

    location /.well-known/acme-challenge/ { root /var/www/certbot; }

    access_log /var/log/nginx/api.example.com.access.log main;
    error_log  /var/log/nginx/api.example.com.error.log warn;

    client_max_body_size 20M;

    location /socket.io/ {
        include snippets/proxy-params.conf;
        proxy_pass http://nest_backend;
        proxy_read_timeout 7d;
        proxy_send_timeout 7d;
        proxy_buffering off;
    }

    location ~ ^/api/(auth|login|register|password-reset) {
        limit_req zone=login burst=5 nodelay;
        include snippets/proxy-params.conf;
        include snippets/security-headers.conf;
        proxy_pass http://nest_backend;
    }

    location / {
        limit_req zone=general burst=40 nodelay;
        limit_conn addr 30;
        include snippets/proxy-params.conf;
        include snippets/security-headers.conf;
        proxy_pass http://nest_backend;
    }
}
EOF

sudo ln -sf /etc/nginx/sites-available/app.example.com /etc/nginx/sites-enabled/
sudo ln -sf /etc/nginx/sites-available/api.example.com /etc/nginx/sites-enabled/

sudo nginx -t && sudo systemctl reload nginx
bash
# LOCAL — should now reach Nginx (502 is expected — no app yet)
curl -I http://app.example.com

Phase 13 — Deploy the application ​

Deploy key ​

bash
# SERVER
ssh-keygen -t ed25519 -C "deploy@prod-app-01" -f ~/.ssh/github_deploy -N ""
chmod 600 ~/.ssh/github_deploy
cat ~/.ssh/github_deploy.pub

Add that public key to GitHub → your repository → Settings → Deploy keys, read-only.

bash
# SERVER
cat >> ~/.ssh/config <<'EOF'
Host github.com
    HostName github.com
    User git
    IdentityFile ~/.ssh/github_deploy
    IdentitiesOnly yes
EOF
chmod 600 ~/.ssh/config

ssh -T git@github.com     # "Hi myorg/myapp! You've successfully authenticated"

Directory layout ​

bash
# SERVER
mkdir -p ~/apps/myapp/{releases,shared/uploads,shared/logs} ~/logs ~/scripts
cd ~/apps/myapp
git clone git@github.com:myorg/myapp.git repo-src
mv repo-src "releases/$(date -u +%Y%m%d-%H%M%S)-initial"
ln -sfn "$(ls -1dt releases/* | head -1)" current
ls -la

The .env file ​

bash
# SERVER
nano ~/apps/myapp/shared/.env
bash
NODE_ENV=production

# API
PORT=3001
HOST=127.0.0.1

DATABASE_URL=postgresql://myapp:YOUR-PG-PASSWORD@127.0.0.1:5432/myapp_production?connection_limit=10&pool_timeout=20
REDIS_URL=redis://:YOUR-REDIS-PASSWORD@127.0.0.1:6379/0

JWT_SECRET=GENERATE_WITH_openssl_rand_hex_32
JWT_EXPIRES_IN=15m
REFRESH_TOKEN_SECRET=GENERATE_ANOTHER_ONE

CORS_ORIGIN=https://app.example.com

# Frontend
NUXT_PUBLIC_API_BASE=https://api.example.com
NUXT_PUBLIC_SITE_URL=https://app.example.com
NUXT_API_INTERNAL_URL=http://127.0.0.1:3001
bash
# SERVER — generate the secrets
openssl rand -hex 32      # JWT_SECRET
openssl rand -hex 32      # REFRESH_TOKEN_SECRET

chmod 600 ~/apps/myapp/shared/.env
ln -sfn ~/apps/myapp/shared/.env ~/apps/myapp/current/.env
ln -sfn ~/apps/myapp/shared/uploads ~/apps/myapp/current/uploads

SAVE THE ENTIRE .env CONTENTS TO YOUR PASSWORD MANAGER NOW

It is not in Git, so it will not be in a code backup. Losing it means losing every session and every integration secret (Level 22).

Build ​

bash
# SERVER
cd ~/apps/myapp/current
pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend prisma migrate deploy
pnpm --filter backend build
pnpm --filter frontend build

ls -la backend/dist/main.js frontend/.output/server/index.mjs

IF THE BUILD IS Killed

That is the OOM killer. Swap (Phase 8) should prevent it. If it still happens:

bash
NODE_OPTIONS="--max-old-space-size=2048" pnpm --filter frontend build

Longer term, build in CI (Level 18).


Phase 14 — PM2 ​

bash
# SERVER
nano ~/apps/myapp/current/ecosystem.config.cjs
js
module.exports = {
  apps: [
    {
      name: 'api',
      cwd: '/home/deploy/apps/myapp/current/backend',
      script: 'dist/main.js',
      exec_mode: 'cluster',
      instances: 2,
      env: { NODE_ENV: 'production', PORT: 3001, HOST: '127.0.0.1' },
      max_memory_restart: '400M',
      kill_timeout: 10000,
      listen_timeout: 8000,
      min_uptime: '30s',
      max_restarts: 10,
      exp_backoff_restart_delay: 200,
      autorestart: true,
      watch: false,
      error_file: '/home/deploy/logs/api-error.log',
      out_file: '/home/deploy/logs/api-out.log',
      merge_logs: true,
      log_date_format: 'YYYY-MM-DD HH:mm:ss Z',
      node_args: '--enable-source-maps',
    },
    {
      name: 'web',
      cwd: '/home/deploy/apps/myapp/current/frontend',
      script: '.output/server/index.mjs',
      exec_mode: 'cluster',
      instances: 2,
      env: {
        NODE_ENV: 'production',
        PORT: 3000, HOST: '127.0.0.1',
        NITRO_PORT: 3000, NITRO_HOST: '127.0.0.1',
      },
      max_memory_restart: '500M',
      kill_timeout: 10000,
      listen_timeout: 8000,
      min_uptime: '30s',
      max_restarts: 10,
      autorestart: true,
      watch: false,
      error_file: '/home/deploy/logs/web-error.log',
      out_file: '/home/deploy/logs/web-out.log',
      merge_logs: true,
      log_date_format: 'YYYY-MM-DD HH:mm:ss Z',
    },
  ],
};
bash
# SERVER
cd ~/apps/myapp/current
pm2 start ecosystem.config.cjs
pm2 list

# ⚠️ Both must show 127.0.0.1
sudo ss -tulpn | grep -E '3000|3001'

curl -i http://127.0.0.1:3001/api/health
curl -sI http://127.0.0.1:3000/ | head -1

Boot startup ​

bash
# SERVER
pm2 startup

Run the printed sudo env PATH=... command exactly.

bash
# SERVER
pm2 save
systemctl status pm2-deploy

Log rotation ​

bash
# SERVER
pm2 install pm2-logrotate
pm2 set pm2-logrotate:max_size 10M
pm2 set pm2-logrotate:retain 14
pm2 set pm2-logrotate:compress true
pm2 set pm2-logrotate:rotateInterval '0 0 * * *'
bash
# LOCAL — the site should now respond over HTTP
curl -I http://app.example.com
curl -s http://api.example.com/api/health

Phase 15 — HTTPS ​

bash
# SERVER
sudo apt install -y certbot python3-certbot-nginx
bash
# LOCAL — verify DNS before proceeding
dig app.example.com +short
dig api.example.com +short
curl -I http://app.example.com
bash
# SERVER — dry run FIRST (rate limits are unforgiving)
sudo certbot --nginx \
  -d app.example.com -d www.example.com -d example.com -d api.example.com \
  --email you@example.com --agree-tos --no-eff-email --redirect --dry-run

If it succeeds, run it for real:

bash
# SERVER
sudo certbot --nginx \
  -d app.example.com -d www.example.com -d example.com -d api.example.com \
  --email you@example.com --agree-tos --no-eff-email --redirect
bash
# SERVER
sudo certbot certificates
sudo nginx -t && sudo systemctl reload nginx

Renewal hook and verification ​

bash
# SERVER
sudo tee /etc/letsencrypt/renewal-hooks/deploy/reload-nginx.sh > /dev/null <<'EOF'
#!/usr/bin/env bash
set -e
/usr/sbin/nginx -t && /usr/bin/systemctl reload nginx
EOF
sudo chmod +x /etc/letsencrypt/renewal-hooks/deploy/reload-nginx.sh

systemctl status certbot.timer
sudo certbot renew --dry-run      # ⚠️ must succeed
bash
# LOCAL — verify
curl -I https://app.example.com
curl -s https://api.example.com/api/health
curl -I http://app.example.com | grep -i location      # 301 to https
echo | openssl s_client -connect app.example.com:443 -servername app.example.com 2>/dev/null | grep -E "^ [0-9] s:"

RAISE HSTS NOW THAT HTTPS WORKS

After a few days of stability, edit /etc/nginx/snippets/security-headers.conf:

nginx
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;

Then sudo nginx -t && sudo systemctl reload nginx.

Your application is live over HTTPS. The remaining phases are what make it production-ready.


Phase 16 — CI/CD ​

Deploy script on the server ​

bash
# SERVER
nano ~/apps/myapp/deploy.sh
bash
#!/usr/bin/env bash
set -euo pipefail

APP="/home/deploy/apps/myapp"
REPO="git@github.com:myorg/myapp.git"
BRANCH="production"
KEEP=5

export NVM_DIR="$HOME/.nvm"
[ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"

[ -d "$APP/repo.git" ] || git clone --bare "$REPO" "$APP/repo.git"
git -C "$APP/repo.git" fetch origin "+refs/heads/*:refs/heads/*" --prune

SHA=$(git -C "$APP/repo.git" rev-parse --short "$BRANCH")
REL="$APP/releases/$(date -u +%Y%m%d-%H%M%S)-$SHA"

echo "==> Building release $REL"
mkdir -p "$REL"
git -C "$APP/repo.git" archive "$BRANCH" | tar -x -C "$REL"

ln -sfn "$APP/shared/.env"    "$REL/.env"
ln -sfn "$APP/shared/uploads" "$REL/uploads"

cd "$REL"
pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend build
pnpm --filter frontend build
pnpm --filter backend prisma migrate deploy

echo "==> Switching current -> $REL"
ln -sfn "$REL" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current"

cd "$APP/current"
GIT_COMMIT="$SHA" pm2 reload ecosystem.config.cjs --update-env
pm2 save

for i in $(seq 1 15); do
  curl -fsS http://127.0.0.1:3001/api/health >/dev/null && { echo "✅ API healthy"; break; }
  [ "$i" -eq 15 ] && { echo "❌ health check failed"; pm2 logs --err --lines 50 --nostream; exit 1; }
  sleep 2
done
curl -fsS -o /dev/null http://127.0.0.1:3000/ && echo "✅ Web healthy"

ls -1dt "$APP/releases"/* | tail -n +$((KEEP + 1)) | xargs -r rm -rf
echo "==> ✅ Deployed $SHA"
bash
# SERVER
nano ~/apps/myapp/rollback.sh
bash
#!/usr/bin/env bash
set -euo pipefail
APP="/home/deploy/apps/myapp"
export NVM_DIR="$HOME/.nvm"; [ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"

CURRENT=$(readlink -f "$APP/current")
PREV=$(ls -1dt "$APP/releases"/* | grep -v "^${CURRENT}$" | head -1)
[ -z "$PREV" ] && { echo "No previous release"; exit 1; }

echo "==> Rolling back to $(basename "$PREV")"
ln -sfn "$PREV" "$APP/current.tmp"
mv -Tf "$APP/current.tmp" "$APP/current"
cd "$APP/current"
pm2 reload ecosystem.config.cjs --update-env && pm2 save
sleep 3
curl -fsS http://127.0.0.1:3001/api/health && echo "✅ Rolled back"
bash
# SERVER
chmod +x ~/apps/myapp/deploy.sh ~/apps/myapp/rollback.sh

GitHub Actions key ​

bash
# LOCAL
ssh-keygen -t ed25519 -C "github-actions-deploy" -f ~/.ssh/gh_deploy -N ""
ssh-copy-id -i ~/.ssh/gh_deploy.pub deploy@203.0.113.10
ssh-keyscan -p 22 203.0.113.10          # copy this output
cat ~/.ssh/gh_deploy                     # copy the WHOLE file

Restrict it on the server:

bash
# SERVER — prefix the github-actions key line in authorized_keys
nano ~/.ssh/authorized_keys
command="/home/deploy/apps/myapp/deploy.sh",no-agent-forwarding,no-port-forwarding,no-pty ssh-ed25519 AAAA... github-actions-deploy

Add repository secrets in GitHub: SSH_PRIVATE_KEY, SSH_HOST, SSH_USER, SSH_PORT, SSH_KNOWN_HOSTS, DEPLOY_PATH.

yaml
# .github/workflows/deploy.yml
name: Deploy
on:
  push: { branches: [production] }
  workflow_dispatch:

concurrency: { group: production-deploy, cancel-in-progress: false }
permissions: { contents: read }

jobs:
  check:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    strategy: { fail-fast: false, matrix: { task: [lint, typecheck] } }
    steps:
      - uses: actions/checkout@v4
      - uses: pnpm/action-setup@v4
      - uses: actions/setup-node@v4
        with: { node-version-file: .nvmrc, cache: 'pnpm' }
      - run: pnpm install --frozen-lockfile
      - run: pnpm --filter backend prisma generate
      - run: pnpm ${{ matrix.task }}

  deploy:
    runs-on: ubuntu-latest
    needs: [check]
    timeout-minutes: 20
    environment: { name: production, url: 'https://app.example.com' }
    steps:
      - name: Configure SSH
        run: |
          mkdir -p ~/.ssh && chmod 700 ~/.ssh
          echo "${{ secrets.SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
          chmod 600 ~/.ssh/id_ed25519
          echo "${{ secrets.SSH_KNOWN_HOSTS }}" > ~/.ssh/known_hosts

      - name: Deploy
        env:
          SSH_HOST: ${{ secrets.SSH_HOST }}
          SSH_USER: ${{ secrets.SSH_USER }}
          SSH_PORT: ${{ secrets.SSH_PORT }}
        run: ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST"

      - name: Smoke test
        run: |
          curl -fsS -o /dev/null -w "app %{http_code}\n" https://app.example.com
          curl -fsS -o /dev/null -w "api %{http_code}\n" https://api.example.com/api/health

THE command= RESTRICTION MAKES THE WORKFLOW TRIVIAL

Because the key is locked to deploy.sh, the workflow just opens an SSH connection — the server runs the script regardless of what is asked. No heredocs, no $ escaping, no NVM sourcing in YAML. It is both simpler and safer.

Test it:

bash
# LOCAL
git checkout -b production && git push -u origin production

Watch the Actions tab. Then verify the deploy: curl -s https://api.example.com/api/health | jq.


Phase 17 — Reboot test ​

DO THIS DELIBERATELY, NOW, WHILE YOU ARE WATCHING

Everything is configured to survive a reboot. Prove it before an unplanned one proves otherwise.

bash
# SERVER
sudo systemctl is-enabled nginx postgresql redis-server pm2-deploy fail2ban ufw
pm2 save
sudo reboot

Wait ~60 seconds, then:

bash
# LOCAL
ssh deploy@203.0.113.10
bash
# SERVER
uptime
systemctl is-active nginx postgresql redis-server fail2ban
pm2 list                                    # both apps online
sudo ufw status                             # active
sudo ss -tulpn | grep -vE "127.0.0.1|::1"   # only 22, 80, 443
bash
# LOCAL
curl -I https://app.example.com
curl -s https://api.example.com/api/health

If anything did not come back, fix it now and reboot again until it does. Only then set Unattended-Upgrade::Automatic-Reboot "true" if you want it.


Phase 18 — Backups ​

bash
# SERVER
sudo mkdir -p /var/backups/myapp
sudo chown deploy:deploy /var/backups/myapp
sudo chmod 700 /var/backups/myapp

sudo -v ; curl https://rclone.org/install.sh | sudo bash
rclone config          # configure your off-site remote as "backup-remote"

sudo apt install -y age
bash
# LOCAL — generate the encryption key; keep the PRIVATE half OFF the server
age-keygen -o ~/myapp-backup-key.txt
grep "public key" ~/myapp-backup-key.txt

Store ~/myapp-backup-key.txt in your password manager. Add the public key to the server .env as BACKUP_AGE_PUBKEY.

Create ~/scripts/backup.sh and ~/scripts/verify-restore.sh from Level 22, then:

bash
# SERVER
chmod +x ~/scripts/*.sh
~/scripts/backup.sh                    # run it once manually
ls -la /var/backups/myapp/
rclone ls backup-remote:myapp-backups/
~/scripts/verify-restore.sh            # ⚠️ must pass
bash
# SERVER
crontab -e
cron
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
SHELL=/bin/bash
MAILTO=you@example.com

15 3 * * *  /home/deploy/scripts/backup.sh          >> /home/deploy/logs/backup.log 2>&1
0  4 * * 0  /home/deploy/scripts/verify-restore.sh  >> /home/deploy/logs/restore-test.log 2>&1
*/10 * * * * /home/deploy/scripts/health-check.sh   >> /home/deploy/logs/health.log 2>&1

Sign up at healthchecks.io, create a check with a 1-day period and 2-hour grace, and add its ping URL to .env as HC_PING_URL.


Phase 19 — Monitoring ​

Create ~/scripts/health-check.sh from Level 21.

bash
# SERVER
chmod +x ~/scripts/health-check.sh
~/scripts/health-check.sh

Persistent journald:

bash
# SERVER
sudo mkdir -p /var/log/journal
sudo systemd-tmpfiles --create --prefix /var/log/journal
sudo tee -a /etc/systemd/journald.conf > /dev/null <<'EOF'

[Journal]
Storage=persistent
SystemMaxUse=1G
MaxRetentionSec=1month
EOF
sudo systemctl restart systemd-journald

External monitoring — sign up at UptimeRobot or Better Stack and add:

MonitorURLInterval
Frontendhttps://app.example.com5 min
API healthhttps://api.example.com/api/health5 min
TLS expiry(enable certificate monitoring)daily

THIS STEP IS NOT OPTIONAL

The cron health check runs on the server. If the server is down, it does not run and you get no alert — silence looks identical to health. Only an external monitor can tell you the server died.


Phase 20 — Final verification ​

bash
# SERVER — security posture
sudo ss -tulpn | grep -vE "127.0.0.1|::1"          # only 22, 80, 443
sudo ufw status verbose
sudo sshd -T | grep -Ei "permitroot|passwordauth|allowusers"
sudo fail2ban-client status sshd
redis-cli ping                                       # NOAUTH
sudo -u postgres psql -c "SHOW listen_addresses;"    # localhost
ls -l ~/apps/myapp/shared/.env                       # -rw-------
systemctl is-enabled unattended-upgrades

# Services and app
systemctl is-active nginx postgresql redis-server fail2ban
pm2 list
sudo certbot certificates
sudo certbot renew --dry-run

# Backups
ls -la /var/backups/myapp/
rclone ls backup-remote:myapp-backups/ | tail -3
tail -5 /home/deploy/logs/restore-verified.log
bash
# LOCAL — external verification
curl -I https://app.example.com
curl -s https://api.example.com/api/health | jq
curl -I http://app.example.com | grep -i location
curl -I https://app.example.com | grep -iE "strict-transport|x-frame|x-content"
nc -vz 203.0.113.10 5432        # must fail
nc -vz 203.0.113.10 6379        # must fail
nc -vz 203.0.113.10 3000        # must fail

Then run the full SSL Labs test at ssllabs.com/ssltest — aim for A or better.


Completion checklist ​

Security

  • [ ] SSH: keys only, no root, AllowUsers deploy, verified from a new terminal
  • [ ] UFW active: 22, 80, 443 only, IPv4 and IPv6
  • [ ] Cloud firewall configured independently
  • [ ] fail2ban sshd jail confirmed active
  • [ ] unattended-upgrades enabled and dry-run verified
  • [ ] PostgreSQL and Redis on 127.0.0.1, verified unreachable from outside
  • [ ] Redis requires authentication
  • [ ] .env is mode 600 and backed up to a password manager

Application

  • [ ] PM2 cluster mode, 2 instances per app
  • [ ] Apps bound to 127.0.0.1:3000 and 127.0.0.1:3001
  • [ ] pm2 startup + pm2 save done
  • [ ] pm2-logrotate installed
  • [ ] Health endpoint returns the deployed commit SHA

Web

  • [ ] Nginx routes both hostnames correctly
  • [ ] HTTPS with a valid certificate, HTTP redirecting
  • [ ] certbot renew --dry-run passes
  • [ ] Renewal hook reloads Nginx
  • [ ] Security headers present
  • [ ] Rate limiting on auth endpoints
  • [ ] Static assets served from disk with long cache headers

Automation

  • [ ] CI/CD deploys on push to production
  • [ ] CI key restricted with command=
  • [ ] Deploy script uses set -euo pipefail and verifies health
  • [ ] Rollback script exists and has been tested

Operations

  • [ ] Full reboot tested — everything came back
  • [ ] Daily backups running, with off-site encrypted copies
  • [ ] Restore verification passing
  • [ ] Dead-man's switch configured
  • [ ] External uptime monitoring active
  • [ ] journald persistent
  • [ ] Decryption key stored off the server

What to do next ​

  1. Practise a rollback — run ~/apps/myapp/rollback.sh deliberately and confirm the site still works.
  2. Do a full restore drill on a scratch VPS (Level 22). Time it. That number is your real RTO.
  3. Write your disaster recovery runbook (Level 28) and store it somewhere that is not on this server.
  4. Raise HSTS to a year once HTTPS has been stable for a week.
  5. Set a calendar reminder for the monthly security audit and quarterly restore drill.

Next: Level 26 — Troubleshooting Handbook →