Skip to content

Level 18 — GitHub Actions ​

A complete, production-ready deployment workflow, explained line by line, followed by every error you are likely to hit.

The vocabulary ​

TermMeaning
WorkflowA YAML file in .github/workflows/ defining an automated process
Event / triggerWhat starts it — push, pull_request, schedule, workflow_dispatch
JobA set of steps running on one runner. Jobs are parallel unless needs: is set.
StepOne command (run:) or one reusable action (uses:)
RunnerThe VM executing a job — GitHub-hosted or self-hosted
ActionA packaged reusable step, e.g. actions/checkout@v4
ArtifactA file uploaded from one job and downloaded by another
SecretAn encrypted value injected as an environment variable
EnvironmentA named target (production) with its own secrets and protection rules

Workflows live in .github/workflows/*.yml and every file there is evaluated on every event.

Setting up the deploy key ​

The CI runner needs to SSH into your VPS. Generate a dedicated key pair for this — not your personal key.

bash
# LOCAL — no passphrase; no human is present to type one
ssh-keygen -t ed25519 -C "github-actions-deploy" -f ~/.ssh/gh_deploy -N ""
bash
# LOCAL — install the public half on the server
ssh-copy-id -i ~/.ssh/gh_deploy.pub deploy@203.0.113.10

Then restrict it on the server:

bash
# SERVER
nano ~/.ssh/authorized_keys
from="140.82.112.0/20,143.55.64.0/20",no-agent-forwarding,no-port-forwarding,no-X11-forwarding ssh-ed25519 AAAAC3... github-actions-deploy

THE STRONGEST RESTRICTION IS command=

command="/home/deploy/apps/myapp/deploy.sh",no-agent-forwarding,no-port-forwarding,no-pty ssh-ed25519 AAAA... github-actions-deploy

This key can run only that one script, regardless of what the client asks for. A leaked CI key can trigger a deploy; it cannot get a shell, read .env, or install anything.

The trade-off is that your workflow can no longer run arbitrary remote commands — everything must be inside the script. That is usually a good constraint.

from= WITH GITHUB'S IP RANGES IS FRAGILE

GitHub-hosted runners use large, changing IP ranges (fetch them from https://api.github.com/meta). Pinning them means your deploys break whenever GitHub adds a range. Use from= only with self-hosted runners at a fixed address; rely on command= restriction otherwise.

Add the secrets ​

Repository → Settings → Secrets and variables → Actions → New repository secret.

SecretValueHow to get it
SSH_PRIVATE_KEYContents of ~/.ssh/gh_deploycat ~/.ssh/gh_deploy — the whole file, including BEGIN/END lines
SSH_HOST203.0.113.10Your server IP
SSH_USERdeployThe deploy user
SSH_PORT22Your SSH port
DEPLOY_PATH/home/deploy/apps/myappAbsolute path
SSH_KNOWN_HOSTSServer host keyssh-keyscan -p 22 203.0.113.10

COPY THE ENTIRE PRIVATE KEY, INCLUDING HEADER AND FOOTER AND THE TRAILING NEWLINE

-----BEGIN OPENSSH PRIVATE KEY-----
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
...
-----END OPENSSH PRIVATE KEY-----

A missing trailing newline produces error in libcrypto or invalid format, which is a genuinely confusing error message. Use:

bash
# macOS
pbcopy < ~/.ssh/gh_deploy
# Linux
xclip -sel clip < ~/.ssh/gh_deploy

rather than selecting text in a terminal.

SSH_KNOWN_HOSTS PREVENTS A REAL ATTACK

Without it, workflows use StrictHostKeyChecking=no, which accepts any host key. An attacker who can redirect your runner's traffic receives your deploy key.

bash
ssh-keyscan -p 22 203.0.113.10

Paste all output lines into the secret. Then verify the fingerprint matches what you see when connecting normally.

The workflow ​

yaml
# .github/workflows/deploy.yml
name: Deploy to Production

on:
  push:
    branches: [production]
  workflow_dispatch:

concurrency:
  group: production-deploy
  cancel-in-progress: false

permissions:
  contents: read

env:
  NODE_VERSION_FILE: .nvmrc

jobs:
  # ---------------------------------------------------------------
  # Quality gates — run in parallel
  # ---------------------------------------------------------------
  lint:
    name: Lint
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - uses: pnpm/action-setup@v4
        with:
          run_install: false

      - uses: actions/setup-node@v4
        with:
          node-version-file: ${{ env.NODE_VERSION_FILE }}
          cache: 'pnpm'

      - run: pnpm install --frozen-lockfile
      - run: pnpm lint

  typecheck:
    name: Typecheck
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: pnpm/action-setup@v4
      - uses: actions/setup-node@v4
        with:
          node-version-file: ${{ env.NODE_VERSION_FILE }}
          cache: 'pnpm'
      - run: pnpm install --frozen-lockfile
      - run: pnpm --filter backend prisma generate
      - run: pnpm typecheck

  test:
    name: Test
    runs-on: ubuntu-latest

    services:
      postgres:
        image: postgres:16-alpine
        env:
          POSTGRES_USER: test
          POSTGRES_PASSWORD: test
          POSTGRES_DB: test_db
        ports: ['5432:5432']
        options: >-
          --health-cmd pg_isready
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5
      redis:
        image: redis:7-alpine
        ports: ['6379:6379']
        options: >-
          --health-cmd "redis-cli ping"
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5

    env:
      DATABASE_URL: postgresql://test:test@127.0.0.1:5432/test_db
      REDIS_URL: redis://127.0.0.1:6379/0
      JWT_SECRET: test-secret-at-least-32-characters-long-xx
      NODE_ENV: test

    steps:
      - uses: actions/checkout@v4
      - uses: pnpm/action-setup@v4
      - uses: actions/setup-node@v4
        with:
          node-version-file: ${{ env.NODE_VERSION_FILE }}
          cache: 'pnpm'
      - run: pnpm install --frozen-lockfile
      - run: pnpm --filter backend prisma generate
      - run: pnpm --filter backend prisma migrate deploy
      - run: pnpm test

  # ---------------------------------------------------------------
  # Build — proves it compiles before we touch the server
  # ---------------------------------------------------------------
  build:
    name: Build
    runs-on: ubuntu-latest
    needs: [lint, typecheck, test]
    steps:
      - uses: actions/checkout@v4
      - uses: pnpm/action-setup@v4
      - uses: actions/setup-node@v4
        with:
          node-version-file: ${{ env.NODE_VERSION_FILE }}
          cache: 'pnpm'
      - run: pnpm install --frozen-lockfile
      - run: pnpm --filter backend prisma generate
      - name: Build
        env:
          NUXT_PUBLIC_API_BASE: https://api.example.com
          NUXT_PUBLIC_SITE_URL: https://app.example.com
        run: pnpm build

  # ---------------------------------------------------------------
  # Deploy
  # ---------------------------------------------------------------
  deploy:
    name: Deploy to VPS
    runs-on: ubuntu-latest
    needs: [build]
    environment:
      name: production
      url: https://app.example.com

    steps:
      - name: Configure SSH
        run: |
          mkdir -p ~/.ssh
          chmod 700 ~/.ssh
          echo "${{ secrets.SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
          chmod 600 ~/.ssh/id_ed25519
          echo "${{ secrets.SSH_KNOWN_HOSTS }}" > ~/.ssh/known_hosts
          chmod 644 ~/.ssh/known_hosts

      - name: Deploy
        env:
          SSH_HOST: ${{ secrets.SSH_HOST }}
          SSH_USER: ${{ secrets.SSH_USER }}
          SSH_PORT: ${{ secrets.SSH_PORT }}
          DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
          GIT_SHA: ${{ github.sha }}
        run: |
          ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<EOF
            set -euo pipefail

            export NVM_DIR="\$HOME/.nvm"
            [ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"

            cd "$DEPLOY_PATH"

            echo "==> Previous commit: \$(git rev-parse --short HEAD)"

            git fetch origin --prune
            git checkout production
            git reset --hard origin/production
            git clean -fd

            echo "==> New commit: \$(git rev-parse --short HEAD)"

            pnpm install --frozen-lockfile
            pnpm --filter backend prisma generate
            pnpm --filter backend prisma migrate deploy
            pnpm --filter backend build
            pnpm --filter frontend build

            export GIT_COMMIT="$GIT_SHA"
            pm2 reload ecosystem.config.cjs --update-env
            pm2 save
          EOF

      - name: Health check
        env:
          SSH_HOST: ${{ secrets.SSH_HOST }}
          SSH_USER: ${{ secrets.SSH_USER }}
          SSH_PORT: ${{ secrets.SSH_PORT }}
        run: |
          ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<'EOF'
            set -euo pipefail
            for i in $(seq 1 15); do
              if curl -fsS http://127.0.0.1:3001/api/health > /dev/null; then
                echo "✅ API healthy"
                curl -fsS -o /dev/null http://127.0.0.1:3000/ && echo "✅ Web healthy"
                exit 0
              fi
              echo "waiting... ($i/15)"
              sleep 3
            done
            echo "❌ Health check failed"
            pm2 logs --err --lines 60 --nostream
            exit 1
          EOF

      - name: Public smoke test
        run: |
          curl -fsS -o /dev/null -w "app: %{http_code}\n" https://app.example.com
          curl -fsS -o /dev/null -w "api: %{http_code}\n" https://api.example.com/api/health

      - name: Notify failure
        if: failure()
        run: echo "::error::Deployment failed for ${{ github.sha }}"

Line-by-line explanation ​

Triggers and concurrency ​

yaml
on:
  push:
    branches: [production]
  workflow_dispatch:

Runs on every push to production. workflow_dispatch adds a "Run workflow" button for manual re-runs — essential when a deploy fails for a transient reason.

NEVER ADD pull_request HERE

Fork PRs would then run deployment code. GitHub withholds secrets from fork PRs specifically to prevent this; do not attempt to work around it with pull_request_target. (Level 17)

yaml
concurrency:
  group: production-deploy
  cancel-in-progress: false

Only one production deploy at a time. cancel-in-progress: false is deliberate:

DO NOT CANCEL AN IN-PROGRESS PRODUCTION DEPLOY

Cancelling mid-prisma migrate deploy can leave the database in a partially-migrated state that neither the old nor new code understands. Let it finish; queue the next one.

cancel-in-progress: true is right for PR test runs, where cancelling a superseded run saves runner time and breaks nothing.

yaml
permissions:
  contents: read

The GITHUB_TOKEN gets read-only repository access. Default permissions are broader than any deploy workflow needs; least privilege applies here too.

Setting up Node and pnpm ​

yaml
      - uses: pnpm/action-setup@v4
        with:
          run_install: false

      - uses: actions/setup-node@v4
        with:
          node-version-file: .nvmrc
          cache: 'pnpm'

ORDER MATTERS: pnpm BEFORE setup-node

cache: 'pnpm' in setup-node needs the pnpm binary to already exist so it can locate the store. Reversing them gives Error: Unable to locate executable file: pnpm.

node-version-file: .nvmrc reads the same file your server uses (Level 6) — one source of truth for the Node version across laptop, CI, and production.

cache: 'pnpm' caches the pnpm store keyed on pnpm-lock.yaml, saving 30–60 seconds per job.

Service containers for tests ​

yaml
    services:
      postgres:
        image: postgres:16-alpine
        ports: ['5432:5432']
        options: >-
          --health-cmd pg_isready
          --health-interval 10s

GitHub starts real PostgreSQL and Redis containers alongside the job, reachable at 127.0.0.1.

WITHOUT --health-cmd, TESTS RACE THE DATABASE

The container starts in milliseconds, but PostgreSQL takes a few seconds to accept connections. Without a healthcheck, your first test run fails with ECONNREFUSED roughly half the time — the classic flaky CI failure.

needs: — the dependency graph ​

yaml
  build:
    needs: [lint, typecheck, test]
  deploy:
    needs: [build]

The three gates run in parallel, so the gate phase takes as long as the slowest one rather than the sum. build waits for all three; deploy waits for build.

Environments ​

yaml
    environment:
      name: production
      url: https://app.example.com

Repository → Settings → Environments → New environment → production. There you can add:

Protection ruleEffect
Required reviewersA human must approve before the deploy job runs
Wait timerDelay N minutes, allowing a cancel
Deployment branchesOnly production may deploy to this environment
Environment secretsSecrets only this environment can read

THIS IS HOW YOU GET CONTINUOUS DELIVERY WITH A GATE

Add yourself as a required reviewer. The pipeline runs every check automatically and then waits. You click approve when ready. All the automation, none of the surprise. (Level 17)

Configuring SSH ​

yaml
      - name: Configure SSH
        run: |
          mkdir -p ~/.ssh
          chmod 700 ~/.ssh
          echo "${{ secrets.SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
          chmod 600 ~/.ssh/id_ed25519
          echo "${{ secrets.SSH_KNOWN_HOSTS }}" > ~/.ssh/known_hosts
          chmod 644 ~/.ssh/known_hosts

The permissions are not optional — SSH refuses to use a private key readable by others (Level 1).

DO NOT USE StrictHostKeyChecking=no

Most tutorials do. It accepts any host key, which means a runner whose traffic is redirected hands your deploy key to whoever answers. Populating known_hosts from a secret costs one extra line and closes the hole.

THE ALTERNATIVE: webfactory/ssh-agent

yaml
      - uses: webfactory/ssh-agent@v0.9.0
        with:
          ssh-private-key: ${{ secrets.SSH_PRIVATE_KEY }}

Loads the key into an agent rather than writing it to disk. Slightly safer and less boilerplate. Third-party actions should be pinned to a commit SHA rather than a tag for supply-chain safety.

The heredoc ​

yaml
          ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<EOF
            set -euo pipefail
            export NVM_DIR="\$HOME/.nvm"
            [ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"
            cd "$DEPLOY_PATH"
            ...
          EOF

bash -s reads the script from stdin, so the heredoc becomes the remote script.

THE $ ESCAPING RULE IS WHERE THIS GOES WRONG

With an unquoted <<EOF, the runner's shell expands variables before sending.

  • $DEPLOY_PATH — expanded locally. Correct: the runner substitutes the secret value.
  • \$HOME — escaped, sent literally. Correct: expanded on the server, giving /home/deploy.

Get this backwards and you either send a literal $DEPLOY_PATH to the server (which does not have that variable) or expand $HOME to the runner's home directory.

With a quoted <<'EOF' (as in the health-check step), nothing is expanded locally — the whole block is sent verbatim. Use that when the script needs no values from the runner. It is simpler and safer; prefer it where possible.

set -euo pipefail IS ESSENTIAL IN THE REMOTE SCRIPT

Without -e, a failed pnpm build does not stop the script. It carries on to pm2 reload, which restarts your app pointing at a stale or missing build — and the workflow reports success. You get a green tick and a broken site. (Level 7)

The NVM line ​

bash
export NVM_DIR="\$HOME/.nvm"
[ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"

THIS IS THE #1 CAUSE OF GITHUB ACTIONS DEPLOY FAILURES

ssh server "command" runs a non-interactive shell, which exits ~/.bashrc before reaching the NVM setup (Level 6). Result: pnpm: command not found, even though the identical command works when you SSH in manually.

Two fixes, and you should do both:

  1. Source NVM explicitly, as above.
  2. Symlink the binaries into /usr/local/bin on the server so they are always on PATH.

Test before writing the workflow:

bash
# LOCAL
ssh deploy@203.0.113.10 "node -v && pnpm -v && pm2 -v"

If this fails, your workflow will fail identically.

Health check and smoke test ​

The health check runs on the server against 127.0.0.1, verifying the app itself. The smoke test runs from the runner against the public URL, verifying DNS, TLS, Nginx, and the app together.

A DEPLOY THAT DOES NOT VERIFY IS NOT A DEPLOY

Without these steps, a workflow reports success as long as pm2 reload returned 0 — which it does even if the app then crash-loops. The retry loop (15 attempts, 3 seconds apart) allows for slow startup while still failing within a minute.

On failure it prints pm2 logs, so the error is in the workflow output and you do not need to SSH in to find out what happened.

Model B — build in CI, ship the artifact ​

When server-side builds OOM or take too long:

yaml
  build:
    runs-on: ubuntu-latest
    needs: [lint, typecheck, test]
    steps:
      - uses: actions/checkout@v4
      - uses: pnpm/action-setup@v4
      - uses: actions/setup-node@v4
        with: { node-version-file: .nvmrc, cache: 'pnpm' }
      - run: pnpm install --frozen-lockfile
      - run: pnpm --filter backend prisma generate
      - run: pnpm build
      - run: pnpm prune --prod

      - name: Package
        run: |
          tar -czf release.tar.gz \
            backend/dist backend/prisma backend/package.json \
            frontend/.output \
            node_modules package.json pnpm-lock.yaml ecosystem.config.cjs

      - uses: actions/upload-artifact@v4
        with:
          name: release
          path: release.tar.gz
          retention-days: 7

  deploy:
    runs-on: ubuntu-latest
    needs: [build]
    environment: production
    steps:
      - uses: actions/download-artifact@v4
        with: { name: release }

      - name: Configure SSH
        run: |
          mkdir -p ~/.ssh && chmod 700 ~/.ssh
          echo "${{ secrets.SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
          chmod 600 ~/.ssh/id_ed25519
          echo "${{ secrets.SSH_KNOWN_HOSTS }}" > ~/.ssh/known_hosts

      - name: Upload and activate release
        env:
          SSH_HOST: ${{ secrets.SSH_HOST }}
          SSH_USER: ${{ secrets.SSH_USER }}
          SHA: ${{ github.sha }}
        run: |
          REL="/home/deploy/apps/myapp/releases/$(date -u +%Y%m%d%H%M%S)-${SHA:0:7}"
          scp release.tar.gz "$SSH_USER@$SSH_HOST:/tmp/"
          ssh "$SSH_USER@$SSH_HOST" bash -s <<EOF
            set -euo pipefail
            export NVM_DIR="\$HOME/.nvm"; . "\$NVM_DIR/nvm.sh"

            mkdir -p "$REL"
            tar -xzf /tmp/release.tar.gz -C "$REL"
            rm /tmp/release.tar.gz

            ln -sfn /home/deploy/apps/myapp/shared/.env "$REL/.env"

            cd "$REL" && pnpm --filter backend prisma migrate deploy

            ln -sfn "$REL" /home/deploy/apps/myapp/current.tmp
            mv -Tf /home/deploy/apps/myapp/current.tmp /home/deploy/apps/myapp/current

            cd /home/deploy/apps/myapp/current
            pm2 reload ecosystem.config.cjs --update-env
            pm2 save

            # Keep the last 5 releases
            ls -1dt /home/deploy/apps/myapp/releases/* | tail -n +6 | xargs -r rm -rf
          EOF

node_modules IN A TARBALL CAN BE LARGE AND ARCHITECTURE-SPECIFIC

Native modules compiled on the runner must match the server's architecture and libc. Both being ubuntu-latest/x86-64 makes this fine; an ARM server would not be. If you hit this, run pnpm install --frozen-lockfile --prod on the server instead of shipping node_modules.

This gives instant rollback — the previous release directory is still there:

bash
# SERVER
ln -sfn /home/deploy/apps/myapp/releases/<previous> current.tmp && mv -Tf current.tmp current
pm2 reload ecosystem.config.cjs --update-env

Full details in Level 20.

A separate CI workflow for pull requests ​

yaml
# .github/workflows/ci.yml
name: CI

on:
  pull_request:
    branches: [main, production]
  push:
    branches: [main]

concurrency:
  group: ci-${{ github.ref }}
  cancel-in-progress: true      # ✅ safe here — no deployment

permissions:
  contents: read

jobs:
  check:
    runs-on: ubuntu-latest
    strategy:
      fail-fast: false
      matrix:
        task: [lint, typecheck, test]
    steps:
      - uses: actions/checkout@v4
      - uses: pnpm/action-setup@v4
      - uses: actions/setup-node@v4
        with: { node-version-file: .nvmrc, cache: 'pnpm' }
      - run: pnpm install --frozen-lockfile
      - run: pnpm --filter backend prisma generate
      - run: pnpm ${{ matrix.task }}

The matrix runs three parallel jobs from one definition. fail-fast: false means a lint failure does not cancel the typecheck job — you see all the problems in one run.

Then require these checks in Settings → Branches → Branch protection rules for production.

Common errors ​

missing server host ​

Using an SSH action and a secret is empty or misnamed.

yaml
      - run: |
          [ -n "${{ secrets.SSH_HOST }}" ] || { echo "SSH_HOST is empty"; exit 1; }

Causes: typo in the secret name (they are case-sensitive), the secret was created in the wrong repository, or it is an environment secret while the job declares no environment:.

Permission denied (publickey) ​

yaml
      - run: ssh -vvv -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" "echo ok"

Checklist:

CheckHow
Public key in server's authorized_keys?grep "github-actions" ~/.ssh/authorized_keys on the server
Right user?SSH_USER should be deploy, not root
Private key complete?Must include BEGIN/END lines and a trailing newline
Permissions on the runner?chmod 600 ~/.ssh/id_ed25519
Server-side permissions?700 on ~/.ssh, 600 on authorized_keys, 755 on /home/deploy
Blocked by AllowUsers?sudo sshd -T | grep -i allowusers
Blocked by fail2ban?sudo fail2ban-client status sshd

Host key verification failed ​

SSH_KNOWN_HOSTS is missing or wrong. Regenerate:

bash
ssh-keyscan -p 22 203.0.113.10

If the server was rebuilt, its host key changed — update the secret.

error in libcrypto / invalid format ​

The private key secret is malformed — usually a missing trailing newline or copied without the BEGIN/END lines. Re-copy with pbcopy/xclip.

pnpm: command not found / node: command not found / pm2: command not found ​

The non-interactive shell problem. Source NVM in the remote script and symlink into /usr/local/bin (Level 6).

bash
# LOCAL — reproduce it exactly
ssh deploy@203.0.113.10 "which node pnpm pm2"

fatal: detected dubious ownership ​

The repository is owned by a different user than the one running Git — usually because something was once run with sudo.

bash
# SERVER — the correct fix
sudo chown -R deploy:deploy /home/deploy/apps/myapp

Not git config --global --add safe.directory (Level 7).

ERR_PNPM_OUTDATED_LOCKFILE ​

pnpm-lock.yaml does not match package.json. Someone edited dependencies without running install, or committed a lockfile from a different pnpm version.

bash
# LOCAL
pnpm install
git add pnpm-lock.yaml && git commit -m "update lockfile"

Pin the pnpm version with packageManager in package.json to prevent recurrence.

Deploy succeeds but the site is unchanged ​

The single most confusing failure. Diagnose in this order:

bash
# SERVER
cd /home/deploy/apps/myapp
git log -1 --oneline                              # did the code update?
find backend/src -newer backend/dist/main.js -name "*.ts" | head   # did the build run?
pm2 list                                           # did PM2 restart? check uptime
pm2 env 0 | grep GIT_COMMIT                        # is the new commit loaded?
curl -s http://127.0.0.1:3001/api/health | jq .commit
FindingCause
git log shows the old commitWrong branch, or fetch failed
Source newer than distBuild did not run or failed silently — missing set -e
PM2 uptime is hourspm2 reload did not run
Health shows the old SHAMissing --update-env
Everything correct but browser shows oldBrowser or CDN cache — hard refresh

Workflow does not trigger ​

  • Wrong branch name in on: push: branches:
  • File is not in .github/workflows/ on the default branch for some event types
  • YAML syntax error — check the Actions tab for a parse failure
  • Actions disabled for the repository, or the workflow was auto-disabled after 60 days of inactivity on a scheduled trigger

Timeout ​

yaml
    timeout-minutes: 20

Default is 360 minutes — a hung deploy would burn six hours of runner time and hold the concurrency lock. Set a realistic limit on every job.

Hardening the workflow ​

PracticeWhy
permissions: contents: readLeast privilege for GITHUB_TOKEN
Pin third-party actions to a commit SHAA compromised tag can be moved; a SHA cannot
Environment protection rulesHuman approval for production
Never echo a secretMasking is best-effort
timeout-minutes on every jobBound the damage from a hang
Dedicated, restricted deploy keyRevocable, limited blast radius
Enable Dependabot for github-actionsActions have CVEs too

PIN THIRD-PARTY ACTIONS BY SHA

yaml
- uses: some-org/some-action@v1                                    # ❌ tag can be moved
- uses: some-org/some-action@a1b2c3d4e5f6...                        # ✅ immutable

@v1 is a mutable Git tag. If the action's repository is compromised, the attacker repoints v1 and every workflow using it runs their code — with your secrets. This has happened to widely-used actions more than once.

GitHub's own actions/* are lower risk; third-party actions should be SHA-pinned.

Production Checklist — Level 18 ​

  • [ ] Dedicated CI SSH key, not a personal key
  • [ ] CI key restricted in authorized_keys (command= and/or no-port-forwarding)
  • [ ] SSH_KNOWN_HOSTS populated — no StrictHostKeyChecking=no
  • [ ] All secrets present and verified non-empty
  • [ ] ssh deploy@server "node -v && pnpm -v && pm2 -v" works from a laptop
  • [ ] NVM sourced explicitly in the remote script
  • [ ] set -euo pipefail in every remote script
  • [ ] Correct $ vs \$ escaping in the heredoc
  • [ ] Triggers on push to production; never on pull_request
  • [ ] concurrency with cancel-in-progress: false for deploys
  • [ ] permissions: contents: read
  • [ ] timeout-minutes on every job
  • [ ] Lint, typecheck, and test run in parallel; build gates on all three
  • [ ] environment: production with required reviewers if you want a gate
  • [ ] Health check on the server and a public smoke test
  • [ ] Health check failure prints pm2 logs and fails the workflow
  • [ ] Third-party actions pinned to commit SHAs
  • [ ] Branch protection requires the CI checks to pass
  • [ ] Deployed commit SHA verifiable via the health endpoint
  • [ ] Rollback procedure documented and tested

Next: Level 19 — GitLab CI/CD and GitLab Runner →