Level 18 — GitHub Actions
A complete, production-ready deployment workflow, explained line by line, followed by every error you are likely to hit.
The vocabulary
| Term | Meaning |
|---|---|
| Workflow | A YAML file in .github/workflows/ defining an automated process |
| Event / trigger | What starts it — push, pull_request, schedule, workflow_dispatch |
| Job | A set of steps running on one runner. Jobs are parallel unless needs: is set. |
| Step | One command (run:) or one reusable action (uses:) |
| Runner | The VM executing a job — GitHub-hosted or self-hosted |
| Action | A packaged reusable step, e.g. actions/checkout@v4 |
| Artifact | A file uploaded from one job and downloaded by another |
| Secret | An encrypted value injected as an environment variable |
| Environment | A named target (production) with its own secrets and protection rules |
Workflows live in .github/workflows/*.yml and every file there is evaluated on every event.
Setting up the deploy key
The CI runner needs to SSH into your VPS. Generate a dedicated key pair for this — not your personal key.
# LOCAL — no passphrase; no human is present to type one
ssh-keygen -t ed25519 -C "github-actions-deploy" -f ~/.ssh/gh_deploy -N ""# LOCAL — install the public half on the server
ssh-copy-id -i ~/.ssh/gh_deploy.pub deploy@203.0.113.10Then restrict it on the server:
# SERVER
nano ~/.ssh/authorized_keysfrom="140.82.112.0/20,143.55.64.0/20",no-agent-forwarding,no-port-forwarding,no-X11-forwarding ssh-ed25519 AAAAC3... github-actions-deployTHE STRONGEST RESTRICTION IS command=
command="/home/deploy/apps/myapp/deploy.sh",no-agent-forwarding,no-port-forwarding,no-pty ssh-ed25519 AAAA... github-actions-deployThis key can run only that one script, regardless of what the client asks for. A leaked CI key can trigger a deploy; it cannot get a shell, read .env, or install anything.
The trade-off is that your workflow can no longer run arbitrary remote commands — everything must be inside the script. That is usually a good constraint.
from= WITH GITHUB'S IP RANGES IS FRAGILE
GitHub-hosted runners use large, changing IP ranges (fetch them from https://api.github.com/meta). Pinning them means your deploys break whenever GitHub adds a range. Use from= only with self-hosted runners at a fixed address; rely on command= restriction otherwise.
Add the secrets
Repository → Settings → Secrets and variables → Actions → New repository secret.
| Secret | Value | How to get it |
|---|---|---|
SSH_PRIVATE_KEY | Contents of ~/.ssh/gh_deploy | cat ~/.ssh/gh_deploy — the whole file, including BEGIN/END lines |
SSH_HOST | 203.0.113.10 | Your server IP |
SSH_USER | deploy | The deploy user |
SSH_PORT | 22 | Your SSH port |
DEPLOY_PATH | /home/deploy/apps/myapp | Absolute path |
SSH_KNOWN_HOSTS | Server host key | ssh-keyscan -p 22 203.0.113.10 |
COPY THE ENTIRE PRIVATE KEY, INCLUDING HEADER AND FOOTER AND THE TRAILING NEWLINE
-----BEGIN OPENSSH PRIVATE KEY-----
b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW
...
-----END OPENSSH PRIVATE KEY-----A missing trailing newline produces error in libcrypto or invalid format, which is a genuinely confusing error message. Use:
# macOS
pbcopy < ~/.ssh/gh_deploy
# Linux
xclip -sel clip < ~/.ssh/gh_deployrather than selecting text in a terminal.
SSH_KNOWN_HOSTS PREVENTS A REAL ATTACK
Without it, workflows use StrictHostKeyChecking=no, which accepts any host key. An attacker who can redirect your runner's traffic receives your deploy key.
ssh-keyscan -p 22 203.0.113.10Paste all output lines into the secret. Then verify the fingerprint matches what you see when connecting normally.
The workflow
# .github/workflows/deploy.yml
name: Deploy to Production
on:
push:
branches: [production]
workflow_dispatch:
concurrency:
group: production-deploy
cancel-in-progress: false
permissions:
contents: read
env:
NODE_VERSION_FILE: .nvmrc
jobs:
# ---------------------------------------------------------------
# Quality gates — run in parallel
# ---------------------------------------------------------------
lint:
name: Lint
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
with:
run_install: false
- uses: actions/setup-node@v4
with:
node-version-file: ${{ env.NODE_VERSION_FILE }}
cache: 'pnpm'
- run: pnpm install --frozen-lockfile
- run: pnpm lint
typecheck:
name: Typecheck
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version-file: ${{ env.NODE_VERSION_FILE }}
cache: 'pnpm'
- run: pnpm install --frozen-lockfile
- run: pnpm --filter backend prisma generate
- run: pnpm typecheck
test:
name: Test
runs-on: ubuntu-latest
services:
postgres:
image: postgres:16-alpine
env:
POSTGRES_USER: test
POSTGRES_PASSWORD: test
POSTGRES_DB: test_db
ports: ['5432:5432']
options: >-
--health-cmd pg_isready
--health-interval 10s
--health-timeout 5s
--health-retries 5
redis:
image: redis:7-alpine
ports: ['6379:6379']
options: >-
--health-cmd "redis-cli ping"
--health-interval 10s
--health-timeout 5s
--health-retries 5
env:
DATABASE_URL: postgresql://test:test@127.0.0.1:5432/test_db
REDIS_URL: redis://127.0.0.1:6379/0
JWT_SECRET: test-secret-at-least-32-characters-long-xx
NODE_ENV: test
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version-file: ${{ env.NODE_VERSION_FILE }}
cache: 'pnpm'
- run: pnpm install --frozen-lockfile
- run: pnpm --filter backend prisma generate
- run: pnpm --filter backend prisma migrate deploy
- run: pnpm test
# ---------------------------------------------------------------
# Build — proves it compiles before we touch the server
# ---------------------------------------------------------------
build:
name: Build
runs-on: ubuntu-latest
needs: [lint, typecheck, test]
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version-file: ${{ env.NODE_VERSION_FILE }}
cache: 'pnpm'
- run: pnpm install --frozen-lockfile
- run: pnpm --filter backend prisma generate
- name: Build
env:
NUXT_PUBLIC_API_BASE: https://api.example.com
NUXT_PUBLIC_SITE_URL: https://app.example.com
run: pnpm build
# ---------------------------------------------------------------
# Deploy
# ---------------------------------------------------------------
deploy:
name: Deploy to VPS
runs-on: ubuntu-latest
needs: [build]
environment:
name: production
url: https://app.example.com
steps:
- name: Configure SSH
run: |
mkdir -p ~/.ssh
chmod 700 ~/.ssh
echo "${{ secrets.SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
chmod 600 ~/.ssh/id_ed25519
echo "${{ secrets.SSH_KNOWN_HOSTS }}" > ~/.ssh/known_hosts
chmod 644 ~/.ssh/known_hosts
- name: Deploy
env:
SSH_HOST: ${{ secrets.SSH_HOST }}
SSH_USER: ${{ secrets.SSH_USER }}
SSH_PORT: ${{ secrets.SSH_PORT }}
DEPLOY_PATH: ${{ secrets.DEPLOY_PATH }}
GIT_SHA: ${{ github.sha }}
run: |
ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<EOF
set -euo pipefail
export NVM_DIR="\$HOME/.nvm"
[ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"
cd "$DEPLOY_PATH"
echo "==> Previous commit: \$(git rev-parse --short HEAD)"
git fetch origin --prune
git checkout production
git reset --hard origin/production
git clean -fd
echo "==> New commit: \$(git rev-parse --short HEAD)"
pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend prisma migrate deploy
pnpm --filter backend build
pnpm --filter frontend build
export GIT_COMMIT="$GIT_SHA"
pm2 reload ecosystem.config.cjs --update-env
pm2 save
EOF
- name: Health check
env:
SSH_HOST: ${{ secrets.SSH_HOST }}
SSH_USER: ${{ secrets.SSH_USER }}
SSH_PORT: ${{ secrets.SSH_PORT }}
run: |
ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<'EOF'
set -euo pipefail
for i in $(seq 1 15); do
if curl -fsS http://127.0.0.1:3001/api/health > /dev/null; then
echo "✅ API healthy"
curl -fsS -o /dev/null http://127.0.0.1:3000/ && echo "✅ Web healthy"
exit 0
fi
echo "waiting... ($i/15)"
sleep 3
done
echo "❌ Health check failed"
pm2 logs --err --lines 60 --nostream
exit 1
EOF
- name: Public smoke test
run: |
curl -fsS -o /dev/null -w "app: %{http_code}\n" https://app.example.com
curl -fsS -o /dev/null -w "api: %{http_code}\n" https://api.example.com/api/health
- name: Notify failure
if: failure()
run: echo "::error::Deployment failed for ${{ github.sha }}"Line-by-line explanation
Triggers and concurrency
on:
push:
branches: [production]
workflow_dispatch:Runs on every push to production. workflow_dispatch adds a "Run workflow" button for manual re-runs — essential when a deploy fails for a transient reason.
NEVER ADD pull_request HERE
Fork PRs would then run deployment code. GitHub withholds secrets from fork PRs specifically to prevent this; do not attempt to work around it with pull_request_target. (Level 17)
concurrency:
group: production-deploy
cancel-in-progress: falseOnly one production deploy at a time. cancel-in-progress: false is deliberate:
DO NOT CANCEL AN IN-PROGRESS PRODUCTION DEPLOY
Cancelling mid-prisma migrate deploy can leave the database in a partially-migrated state that neither the old nor new code understands. Let it finish; queue the next one.
cancel-in-progress: true is right for PR test runs, where cancelling a superseded run saves runner time and breaks nothing.
permissions:
contents: readThe GITHUB_TOKEN gets read-only repository access. Default permissions are broader than any deploy workflow needs; least privilege applies here too.
Setting up Node and pnpm
- uses: pnpm/action-setup@v4
with:
run_install: false
- uses: actions/setup-node@v4
with:
node-version-file: .nvmrc
cache: 'pnpm'ORDER MATTERS: pnpm BEFORE setup-node
cache: 'pnpm' in setup-node needs the pnpm binary to already exist so it can locate the store. Reversing them gives Error: Unable to locate executable file: pnpm.
node-version-file: .nvmrc reads the same file your server uses (Level 6) — one source of truth for the Node version across laptop, CI, and production.
cache: 'pnpm' caches the pnpm store keyed on pnpm-lock.yaml, saving 30–60 seconds per job.
Service containers for tests
services:
postgres:
image: postgres:16-alpine
ports: ['5432:5432']
options: >-
--health-cmd pg_isready
--health-interval 10sGitHub starts real PostgreSQL and Redis containers alongside the job, reachable at 127.0.0.1.
WITHOUT --health-cmd, TESTS RACE THE DATABASE
The container starts in milliseconds, but PostgreSQL takes a few seconds to accept connections. Without a healthcheck, your first test run fails with ECONNREFUSED roughly half the time — the classic flaky CI failure.
needs: — the dependency graph
build:
needs: [lint, typecheck, test]
deploy:
needs: [build]The three gates run in parallel, so the gate phase takes as long as the slowest one rather than the sum. build waits for all three; deploy waits for build.
Environments
environment:
name: production
url: https://app.example.comRepository → Settings → Environments → New environment → production. There you can add:
| Protection rule | Effect |
|---|---|
| Required reviewers | A human must approve before the deploy job runs |
| Wait timer | Delay N minutes, allowing a cancel |
| Deployment branches | Only production may deploy to this environment |
| Environment secrets | Secrets only this environment can read |
THIS IS HOW YOU GET CONTINUOUS DELIVERY WITH A GATE
Add yourself as a required reviewer. The pipeline runs every check automatically and then waits. You click approve when ready. All the automation, none of the surprise. (Level 17)
Configuring SSH
- name: Configure SSH
run: |
mkdir -p ~/.ssh
chmod 700 ~/.ssh
echo "${{ secrets.SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
chmod 600 ~/.ssh/id_ed25519
echo "${{ secrets.SSH_KNOWN_HOSTS }}" > ~/.ssh/known_hosts
chmod 644 ~/.ssh/known_hostsThe permissions are not optional — SSH refuses to use a private key readable by others (Level 1).
DO NOT USE StrictHostKeyChecking=no
Most tutorials do. It accepts any host key, which means a runner whose traffic is redirected hands your deploy key to whoever answers. Populating known_hosts from a secret costs one extra line and closes the hole.
THE ALTERNATIVE: webfactory/ssh-agent
- uses: webfactory/ssh-agent@v0.9.0
with:
ssh-private-key: ${{ secrets.SSH_PRIVATE_KEY }}Loads the key into an agent rather than writing it to disk. Slightly safer and less boilerplate. Third-party actions should be pinned to a commit SHA rather than a tag for supply-chain safety.
The heredoc
ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<EOF
set -euo pipefail
export NVM_DIR="\$HOME/.nvm"
[ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"
cd "$DEPLOY_PATH"
...
EOFbash -s reads the script from stdin, so the heredoc becomes the remote script.
THE $ ESCAPING RULE IS WHERE THIS GOES WRONG
With an unquoted <<EOF, the runner's shell expands variables before sending.
$DEPLOY_PATH— expanded locally. Correct: the runner substitutes the secret value.\$HOME— escaped, sent literally. Correct: expanded on the server, giving/home/deploy.
Get this backwards and you either send a literal $DEPLOY_PATH to the server (which does not have that variable) or expand $HOME to the runner's home directory.
With a quoted <<'EOF' (as in the health-check step), nothing is expanded locally — the whole block is sent verbatim. Use that when the script needs no values from the runner. It is simpler and safer; prefer it where possible.
set -euo pipefail IS ESSENTIAL IN THE REMOTE SCRIPT
Without -e, a failed pnpm build does not stop the script. It carries on to pm2 reload, which restarts your app pointing at a stale or missing build — and the workflow reports success. You get a green tick and a broken site. (Level 7)
The NVM line
export NVM_DIR="\$HOME/.nvm"
[ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"THIS IS THE #1 CAUSE OF GITHUB ACTIONS DEPLOY FAILURES
ssh server "command" runs a non-interactive shell, which exits ~/.bashrc before reaching the NVM setup (Level 6). Result: pnpm: command not found, even though the identical command works when you SSH in manually.
Two fixes, and you should do both:
- Source NVM explicitly, as above.
- Symlink the binaries into
/usr/local/binon the server so they are always onPATH.
Test before writing the workflow:
# LOCAL
ssh deploy@203.0.113.10 "node -v && pnpm -v && pm2 -v"If this fails, your workflow will fail identically.
Health check and smoke test
The health check runs on the server against 127.0.0.1, verifying the app itself. The smoke test runs from the runner against the public URL, verifying DNS, TLS, Nginx, and the app together.
A DEPLOY THAT DOES NOT VERIFY IS NOT A DEPLOY
Without these steps, a workflow reports success as long as pm2 reload returned 0 — which it does even if the app then crash-loops. The retry loop (15 attempts, 3 seconds apart) allows for slow startup while still failing within a minute.
On failure it prints pm2 logs, so the error is in the workflow output and you do not need to SSH in to find out what happened.
Model B — build in CI, ship the artifact
When server-side builds OOM or take too long:
build:
runs-on: ubuntu-latest
needs: [lint, typecheck, test]
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with: { node-version-file: .nvmrc, cache: 'pnpm' }
- run: pnpm install --frozen-lockfile
- run: pnpm --filter backend prisma generate
- run: pnpm build
- run: pnpm prune --prod
- name: Package
run: |
tar -czf release.tar.gz \
backend/dist backend/prisma backend/package.json \
frontend/.output \
node_modules package.json pnpm-lock.yaml ecosystem.config.cjs
- uses: actions/upload-artifact@v4
with:
name: release
path: release.tar.gz
retention-days: 7
deploy:
runs-on: ubuntu-latest
needs: [build]
environment: production
steps:
- uses: actions/download-artifact@v4
with: { name: release }
- name: Configure SSH
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
echo "${{ secrets.SSH_PRIVATE_KEY }}" > ~/.ssh/id_ed25519
chmod 600 ~/.ssh/id_ed25519
echo "${{ secrets.SSH_KNOWN_HOSTS }}" > ~/.ssh/known_hosts
- name: Upload and activate release
env:
SSH_HOST: ${{ secrets.SSH_HOST }}
SSH_USER: ${{ secrets.SSH_USER }}
SHA: ${{ github.sha }}
run: |
REL="/home/deploy/apps/myapp/releases/$(date -u +%Y%m%d%H%M%S)-${SHA:0:7}"
scp release.tar.gz "$SSH_USER@$SSH_HOST:/tmp/"
ssh "$SSH_USER@$SSH_HOST" bash -s <<EOF
set -euo pipefail
export NVM_DIR="\$HOME/.nvm"; . "\$NVM_DIR/nvm.sh"
mkdir -p "$REL"
tar -xzf /tmp/release.tar.gz -C "$REL"
rm /tmp/release.tar.gz
ln -sfn /home/deploy/apps/myapp/shared/.env "$REL/.env"
cd "$REL" && pnpm --filter backend prisma migrate deploy
ln -sfn "$REL" /home/deploy/apps/myapp/current.tmp
mv -Tf /home/deploy/apps/myapp/current.tmp /home/deploy/apps/myapp/current
cd /home/deploy/apps/myapp/current
pm2 reload ecosystem.config.cjs --update-env
pm2 save
# Keep the last 5 releases
ls -1dt /home/deploy/apps/myapp/releases/* | tail -n +6 | xargs -r rm -rf
EOFnode_modules IN A TARBALL CAN BE LARGE AND ARCHITECTURE-SPECIFIC
Native modules compiled on the runner must match the server's architecture and libc. Both being ubuntu-latest/x86-64 makes this fine; an ARM server would not be. If you hit this, run pnpm install --frozen-lockfile --prod on the server instead of shipping node_modules.
This gives instant rollback — the previous release directory is still there:
# SERVER
ln -sfn /home/deploy/apps/myapp/releases/<previous> current.tmp && mv -Tf current.tmp current
pm2 reload ecosystem.config.cjs --update-envFull details in Level 20.
A separate CI workflow for pull requests
# .github/workflows/ci.yml
name: CI
on:
pull_request:
branches: [main, production]
push:
branches: [main]
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true # ✅ safe here — no deployment
permissions:
contents: read
jobs:
check:
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix:
task: [lint, typecheck, test]
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with: { node-version-file: .nvmrc, cache: 'pnpm' }
- run: pnpm install --frozen-lockfile
- run: pnpm --filter backend prisma generate
- run: pnpm ${{ matrix.task }}The matrix runs three parallel jobs from one definition. fail-fast: false means a lint failure does not cancel the typecheck job — you see all the problems in one run.
Then require these checks in Settings → Branches → Branch protection rules for production.
Common errors
missing server host
Using an SSH action and a secret is empty or misnamed.
- run: |
[ -n "${{ secrets.SSH_HOST }}" ] || { echo "SSH_HOST is empty"; exit 1; }Causes: typo in the secret name (they are case-sensitive), the secret was created in the wrong repository, or it is an environment secret while the job declares no environment:.
Permission denied (publickey)
- run: ssh -vvv -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" "echo ok"Checklist:
| Check | How |
|---|---|
Public key in server's authorized_keys? | grep "github-actions" ~/.ssh/authorized_keys on the server |
| Right user? | SSH_USER should be deploy, not root |
| Private key complete? | Must include BEGIN/END lines and a trailing newline |
| Permissions on the runner? | chmod 600 ~/.ssh/id_ed25519 |
| Server-side permissions? | 700 on ~/.ssh, 600 on authorized_keys, 755 on /home/deploy |
Blocked by AllowUsers? | sudo sshd -T | grep -i allowusers |
| Blocked by fail2ban? | sudo fail2ban-client status sshd |
Host key verification failed
SSH_KNOWN_HOSTS is missing or wrong. Regenerate:
ssh-keyscan -p 22 203.0.113.10If the server was rebuilt, its host key changed — update the secret.
error in libcrypto / invalid format
The private key secret is malformed — usually a missing trailing newline or copied without the BEGIN/END lines. Re-copy with pbcopy/xclip.
pnpm: command not found / node: command not found / pm2: command not found
The non-interactive shell problem. Source NVM in the remote script and symlink into /usr/local/bin (Level 6).
# LOCAL — reproduce it exactly
ssh deploy@203.0.113.10 "which node pnpm pm2"fatal: detected dubious ownership
The repository is owned by a different user than the one running Git — usually because something was once run with sudo.
# SERVER — the correct fix
sudo chown -R deploy:deploy /home/deploy/apps/myappNot git config --global --add safe.directory (Level 7).
ERR_PNPM_OUTDATED_LOCKFILE
pnpm-lock.yaml does not match package.json. Someone edited dependencies without running install, or committed a lockfile from a different pnpm version.
# LOCAL
pnpm install
git add pnpm-lock.yaml && git commit -m "update lockfile"Pin the pnpm version with packageManager in package.json to prevent recurrence.
Deploy succeeds but the site is unchanged
The single most confusing failure. Diagnose in this order:
# SERVER
cd /home/deploy/apps/myapp
git log -1 --oneline # did the code update?
find backend/src -newer backend/dist/main.js -name "*.ts" | head # did the build run?
pm2 list # did PM2 restart? check uptime
pm2 env 0 | grep GIT_COMMIT # is the new commit loaded?
curl -s http://127.0.0.1:3001/api/health | jq .commit| Finding | Cause |
|---|---|
git log shows the old commit | Wrong branch, or fetch failed |
Source newer than dist | Build did not run or failed silently — missing set -e |
| PM2 uptime is hours | pm2 reload did not run |
| Health shows the old SHA | Missing --update-env |
| Everything correct but browser shows old | Browser or CDN cache — hard refresh |
Workflow does not trigger
- Wrong branch name in
on: push: branches: - File is not in
.github/workflows/on the default branch for some event types - YAML syntax error — check the Actions tab for a parse failure
- Actions disabled for the repository, or the workflow was auto-disabled after 60 days of inactivity on a scheduled trigger
Timeout
timeout-minutes: 20Default is 360 minutes — a hung deploy would burn six hours of runner time and hold the concurrency lock. Set a realistic limit on every job.
Hardening the workflow
| Practice | Why |
|---|---|
permissions: contents: read | Least privilege for GITHUB_TOKEN |
| Pin third-party actions to a commit SHA | A compromised tag can be moved; a SHA cannot |
| Environment protection rules | Human approval for production |
Never echo a secret | Masking is best-effort |
timeout-minutes on every job | Bound the damage from a hang |
| Dedicated, restricted deploy key | Revocable, limited blast radius |
Enable Dependabot for github-actions | Actions have CVEs too |
PIN THIRD-PARTY ACTIONS BY SHA
- uses: some-org/some-action@v1 # ❌ tag can be moved
- uses: some-org/some-action@a1b2c3d4e5f6... # ✅ immutable@v1 is a mutable Git tag. If the action's repository is compromised, the attacker repoints v1 and every workflow using it runs their code — with your secrets. This has happened to widely-used actions more than once.
GitHub's own actions/* are lower risk; third-party actions should be SHA-pinned.
Production Checklist — Level 18
- [ ] Dedicated CI SSH key, not a personal key
- [ ] CI key restricted in
authorized_keys(command=and/orno-port-forwarding) - [ ]
SSH_KNOWN_HOSTSpopulated — noStrictHostKeyChecking=no - [ ] All secrets present and verified non-empty
- [ ]
ssh deploy@server "node -v && pnpm -v && pm2 -v"works from a laptop - [ ] NVM sourced explicitly in the remote script
- [ ]
set -euo pipefailin every remote script - [ ] Correct
$vs\$escaping in the heredoc - [ ] Triggers on
pushtoproduction; never onpull_request - [ ]
concurrencywithcancel-in-progress: falsefor deploys - [ ]
permissions: contents: read - [ ]
timeout-minuteson every job - [ ] Lint, typecheck, and test run in parallel; build gates on all three
- [ ]
environment: productionwith required reviewers if you want a gate - [ ] Health check on the server and a public smoke test
- [ ] Health check failure prints
pm2 logsand fails the workflow - [ ] Third-party actions pinned to commit SHAs
- [ ] Branch protection requires the CI checks to pass
- [ ] Deployed commit SHA verifiable via the health endpoint
- [ ] Rollback procedure documented and tested