Level 19 — GitLab CI/CD and GitLab Runner
The same deployment pipeline in GitLab, plus how to install and configure your own runner.
How GitLab CI works
GitLab looks for .gitlab-ci.yml in the repository root. Any push triggers a pipeline made of stages, each containing jobs, executed by runners.
| Term | Meaning |
|---|---|
| Pipeline | One complete run |
| Stage | An ordered phase. All jobs in a stage run in parallel; the next stage waits. |
| Job | One unit of work with a script: |
| Runner | The agent that executes jobs |
| Executor | How the runner executes — shell, Docker, Kubernetes |
| Artifact | Files a job produces, passed to later stages |
| Cache | Reused data between runs (dependencies) |
| Environment | A named deployment target with history and rollback UI |
The complete .gitlab-ci.yml
# .gitlab-ci.yml
stages:
- test
- build
- deploy
default:
image: node:22-slim
interruptible: true
before_script:
- corepack enable
- corepack prepare pnpm@9.12.0 --activate
- pnpm install --frozen-lockfile
variables:
PNPM_STORE: .pnpm-store
GIT_DEPTH: "20"
FF_USE_FASTZIP: "true"
cache:
key:
files:
- pnpm-lock.yaml
paths:
- .pnpm-store
policy: pull-push
workflow:
rules:
- if: $CI_PIPELINE_SOURCE == "merge_request_event"
- if: $CI_COMMIT_BRANCH == "main"
- if: $CI_COMMIT_BRANCH == "production"
- when: never
# ---------------------------------------------------------------
# Stage: test
# ---------------------------------------------------------------
lint:
stage: test
script:
- pnpm lint
typecheck:
stage: test
script:
- pnpm --filter backend prisma generate
- pnpm typecheck
test:
stage: test
services:
- name: postgres:16-alpine
alias: postgres
- name: redis:7-alpine
alias: redis
variables:
POSTGRES_USER: test
POSTGRES_PASSWORD: test
POSTGRES_DB: test_db
POSTGRES_HOST_AUTH_METHOD: trust
DATABASE_URL: "postgresql://test:test@postgres:5432/test_db"
REDIS_URL: "redis://redis:6379/0"
JWT_SECRET: "test-secret-at-least-32-characters-long-xx"
NODE_ENV: test
script:
- pnpm --filter backend prisma generate
- pnpm --filter backend prisma migrate deploy
- pnpm test
coverage: '/All files[^|]*\|[^|]*\s+([\d\.]+)/'
artifacts:
when: always
reports:
junit: backend/junit.xml
expire_in: 1 week
# ---------------------------------------------------------------
# Stage: build
# ---------------------------------------------------------------
build:
stage: build
variables:
NUXT_PUBLIC_API_BASE: "https://api.example.com"
NUXT_PUBLIC_SITE_URL: "https://app.example.com"
script:
- pnpm --filter backend prisma generate
- pnpm build
artifacts:
name: "build-$CI_COMMIT_SHORT_SHA"
paths:
- backend/dist/
- frontend/.output/
expire_in: 1 week
rules:
- if: $CI_COMMIT_BRANCH == "production"
- if: $CI_COMMIT_BRANCH == "main"
# ---------------------------------------------------------------
# Stage: deploy
# ---------------------------------------------------------------
.deploy_template: &deploy_template
stage: deploy
image: alpine:3.20
before_script:
- apk add --no-cache openssh-client curl
- mkdir -p ~/.ssh && chmod 700 ~/.ssh
- cp "$SSH_PRIVATE_KEY" ~/.ssh/id_ed25519
- chmod 600 ~/.ssh/id_ed25519
- cp "$SSH_KNOWN_HOSTS_FILE" ~/.ssh/known_hosts
- chmod 644 ~/.ssh/known_hosts
deploy:production:
<<: *deploy_template
environment:
name: production
url: https://app.example.com
rules:
- if: $CI_COMMIT_BRANCH == "production"
when: manual
allow_failure: false
script:
- |
ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<EOF
set -euo pipefail
export NVM_DIR="\$HOME/.nvm"
[ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"
cd "$DEPLOY_PATH"
echo "==> Previous: \$(git rev-parse --short HEAD)"
git fetch origin --prune
git checkout production
git reset --hard origin/production
git clean -fd
echo "==> New: \$(git rev-parse --short HEAD)"
pnpm install --frozen-lockfile
pnpm --filter backend prisma generate
pnpm --filter backend prisma migrate deploy
pnpm --filter backend build
pnpm --filter frontend build
export GIT_COMMIT="$CI_COMMIT_SHORT_SHA"
pm2 reload ecosystem.config.cjs --update-env
pm2 save
EOF
- |
ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<'EOF'
set -euo pipefail
for i in $(seq 1 15); do
if curl -fsS http://127.0.0.1:3001/api/health > /dev/null; then
echo "✅ API healthy"; exit 0
fi
sleep 3
done
echo "❌ Health check failed"
pm2 logs --err --lines 60 --nostream
exit 1
EOF
- curl -fsS -o /dev/null -w "app: %{http_code}\n" https://app.example.comExplaining the key parts
default: and before_script
default:
image: node:22-slim
before_script:
- corepack enable
- pnpm install --frozen-lockfileApplies to every job unless overridden. The image: is the Docker image the job runs inside (with the Docker executor).
before_script RUNS FOR EVERY JOB, INCLUDING DEPLOY
The deploy job uses alpine:3.20 and does not need Node, so it overrides before_script entirely. Forgetting to override means a pointless pnpm install in a container without pnpm — and a confusing failure.
workflow: rules — controlling when pipelines run
workflow:
rules:
- if: $CI_PIPELINE_SOURCE == "merge_request_event"
- if: $CI_COMMIT_BRANCH == "main"
- if: $CI_COMMIT_BRANCH == "production"
- when: neverTHIS PREVENTS DUPLICATE PIPELINES
Without it, pushing to a branch with an open merge request creates two pipelines — one for the branch, one for the MR. You burn double the runner minutes and get confusing duplicate statuses. This block is the standard fix.
cache vs artifacts
cache | artifacts | |
|---|---|---|
| Purpose | Speed up future runs | Pass files between stages |
| Guaranteed present | ❌ No — best effort | ✅ Yes |
| Scope | Shared across pipelines | This pipeline |
| Use for | .pnpm-store, node_modules | dist/, .output/, test reports |
NEVER CACHE SOMETHING YOU CANNOT REGENERATE
Cache is explicitly best-effort — the runner may evict it, and on a different runner it may not exist at all. If your build depends on cached output, it will fail unpredictably. Use artifacts for anything a later stage requires.
cache:
key:
files:
- pnpm-lock.yamlKeying on the lockfile means the cache invalidates exactly when dependencies change — not on every commit, and not never.
rules — when a job runs
rules:
- if: $CI_COMMIT_BRANCH == "production"
when: manual
allow_failure: false| Clause | Effect |
|---|---|
if: | Condition using CI variables |
when: manual | Shows a ▶ button — a human must click |
when: on_success | Default — runs if earlier stages passed |
when: always | Runs even if earlier stages failed |
allow_failure: false | A manual job that is skipped blocks the pipeline |
rules REPLACED only/except
only: and except: still work but are legacy and cannot express everything rules can. Do not mix them in one job — GitLab rejects that. Use rules throughout.
when: manual IS GITLAB'S DEPLOYMENT GATE
This is the equivalent of GitHub's environment approval. Everything runs automatically up to the deploy, which waits for a click. Combined with protected environments (Settings → CI/CD → Protected environments), you can restrict who is allowed to click.
YAML anchors
.deploy_template: &deploy_template
stage: deploy
before_script: [...]
deploy:production:
<<: *deploy_templateA key starting with . is a hidden job — a template, never executed. &name defines an anchor, <<: *name merges it in. This is how you share configuration between deploy:staging and deploy:production without duplication.
environment:
environment:
name: production
url: https://app.example.comGitLab tracks deployments per environment, giving you Operations → Environments with full history and a Rollback button that re-runs the deploy job for a previous commit.
THE ROLLBACK BUTTON RE-RUNS THE OLD PIPELINE'S DEPLOY JOB
It does not undo database migrations. If the release you are rolling back from included a destructive migration, the old code will fail against the new schema. Additive-only migrations are what make this button safe (Level 17).
CI/CD variables
Settings → CI/CD → Variables.
| Option | Meaning | Use |
|---|---|---|
| Protected | Only exposed on protected branches/tags | ✅ All deploy secrets |
| Masked | Replaced with [MASKED] in logs | ✅ Where the value qualifies |
| File | Value written to a temp file; the variable holds the path | ✅ SSH keys, certificates |
| Expanded | $OTHER_VAR inside the value is expanded | Usually leave off for secrets |
| Environment scope | Restrict to production, staging, etc. | ✅ Isolate environments |
Variables for this pipeline:
| Variable | Type | Protected | Value |
|---|---|---|---|
SSH_PRIVATE_KEY | File | ✅ | Contents of the deploy private key |
SSH_KNOWN_HOSTS_FILE | File | ✅ | ssh-keyscan output |
SSH_HOST | Variable | ✅ | 203.0.113.10 |
SSH_USER | Variable | ✅ | deploy |
SSH_PORT | Variable | ✅ | 22 |
DEPLOY_PATH | Variable | ✅ | /home/deploy/apps/myapp |
USE "FILE" TYPE FOR SSH KEYS
Multi-line values cannot be masked, and writing one with echo "$KEY" > id_ed25519 mangles newlines — producing error in libcrypto.
File-type variables solve both: GitLab writes the value to a temp file and sets the variable to its path.
- cp "$SSH_PRIVATE_KEY" ~/.ssh/id_ed25519
- chmod 600 ~/.ssh/id_ed25519ALWAYS TICK "PROTECTED" FOR DEPLOY SECRETS
Without it, any branch — including one pushed by a contributor with limited access — can read your production SSH key by adding a job that prints it. Protected variables are only exposed on protected branches, which only maintainers can push to.
Then protect the branch: Settings → Repository → Protected branches → production.
MASKING REQUIREMENTS ARE STRICT
A value can only be masked if it is at least 8 characters, single-line, and uses only base64 characters. Values that fail these rules are silently not masked — GitLab shows a warning when you save, which is easy to miss. Check after saving.
GitLab Runner
Shared vs self-hosted
| GitLab.com shared runners | Self-hosted runner | |
|---|---|---|
| Setup | None | Install and register |
| Cost | Free minutes, then paid | Your server's cost |
| Speed | Cold start each job | Warm caches, faster |
| Network access to your VPS | Public internet only | Can be on the same private network |
| Maintenance | None | Yours |
| Security | GitLab's isolation | Your responsibility |
DO NOT RUN THE CI RUNNER ON YOUR PRODUCTION SERVER
It is tempting — the runner is already where you want to deploy, so no SSH keys are needed.
Why not:
- CI jobs run untrusted code — every dependency's install script, every branch anyone pushes
- With the shell executor, that code runs as the runner's user on your production box
- Builds consume the CPU and RAM your application needs, causing latency spikes or OOM kills
- A compromised dependency in CI becomes a compromised production server
If you must self-host, use a separate, cheap VPS for the runner and have it SSH into production like any other client.
Installing on Ubuntu 24.04
# ON THE RUNNER SERVER (not production)
curl -L "https://packages.gitlab.com/install/repositories/runner/gitlab-runner/script.deb.sh" | sudo bash
sudo apt install -y gitlab-runner
sudo systemctl status gitlab-runnerRegistering
Get a runner authentication token: project → Settings → CI/CD → Runners → New project runner.
# RUNNER SERVER
sudo gitlab-runner register \
--non-interactive \
--url "https://gitlab.com/" \
--token "glrt-xxxxxxxxxxxxxxxx" \
--executor "docker" \
--docker-image "node:22-slim" \
--docker-privileged=false \
--docker-volumes "/cache" \
--description "prod-deploy-runner"| Flag | Meaning |
|---|---|
--url | Your GitLab instance |
--token | Runner authentication token (the old --registration-token is deprecated) |
--executor | How jobs run — see below |
--docker-image | Default image for jobs that do not specify one |
--docker-privileged | Keep false unless you need Docker-in-Docker |
Configuration lands in /etc/gitlab-runner/config.toml:
concurrent = 4
check_interval = 3
[[runners]]
name = "prod-deploy-runner"
url = "https://gitlab.com/"
token = "glrt-..."
executor = "docker"
[runners.docker]
image = "node:22-slim"
privileged = false
volumes = ["/cache"]
shm_size = 268435456
[runners.cache]
Type = "local"sudo systemctl restart gitlab-runner
sudo gitlab-runner verify
sudo gitlab-runner listExecutors
| Executor | How it runs jobs | Isolation | Use when |
|---|---|---|---|
| shell | Directly on the host, as the gitlab-runner user | ❌ None | Never on a shared or production machine |
| docker | Each job in a fresh container | ✅ Good | Default choice |
| docker+machine | Provisions a VM per job | ✅ Excellent | Autoscaling |
| kubernetes | A pod per job | ✅ Excellent | You already run Kubernetes |
THE SHELL EXECUTOR RUNS UNTRUSTED CODE AS A REAL USER
Every npm install executes arbitrary postinstall scripts from your dependency tree, as the gitlab-runner user, on that host, with no isolation. Files persist between jobs, so one job can plant something the next job runs.
Use the Docker executor. Each job gets a fresh container that is destroyed afterwards.
privileged = true IS ROOT ON THE HOST
Required for Docker-in-Docker builds, and it grants the container effectively full host access. If you need to build images, prefer Kaniko or BuildKit in rootless mode over privileged DinD.
Tags
deploy:production:
tags:
- production-deploysudo gitlab-runner register ... --tag-list "production-deploy"A job with tags: runs only on a runner carrying all of them.
A JOB WITH TAGS AND NO MATCHING RUNNER HANGS FOREVER
It sits in "pending" with no error, waiting for a runner that will never appear. Check Settings → CI/CD → Runners that a runner with that tag exists and is online (green).
Conversely, if your runner is registered to only run tagged jobs (the default for project runners in some configurations), untagged jobs never get picked up.
Runner user and permissions
# RUNNER SERVER — the shell executor's user, if you use it
sudo usermod -aG docker gitlab-runner # ⚠️ equivalent to rootADDING gitlab-runner TO THE docker GROUP IS ROOT ACCESS
Any CI job can then run docker run -v /:/host and own the machine (Level 11). Combined with the shell executor and untrusted branches, this is a full compromise path.
Acceptable only on a disposable, isolated runner host that touches nothing else.
GitHub Actions vs GitLab CI
| GitHub Actions | GitLab CI | |
|---|---|---|
| Config file | .github/workflows/*.yml (multiple) | .gitlab-ci.yml (one, with include:) |
| Model | Jobs with needs: (a DAG) | Stages, with needs: for a DAG |
| Reuse | Marketplace actions (huge ecosystem) | include:, templates, YAML anchors |
| Manual gate | Environment protection rules | when: manual + protected environments |
| Secrets | Repository / environment / org secrets | Variables with protected/masked/file |
| File-type secrets | ❌ Not built in | ✅ Yes — much better for SSH keys |
| Self-hosted runners | Yes | Yes — historically stronger |
| Built-in registry | GHCR | Built-in container registry per project |
| Deployment history + rollback UI | Basic | ✅ Environments view with rollback |
| Free tier (public repos) | Generous | Generous |
| Free tier (private) | 2,000 min/month | 400 min/month (free tier) |
| Ecosystem | ✅ Far larger | Smaller |
| Self-hostable platform | GitHub Enterprise (paid) | ✅ GitLab CE (free) |
WHICH TO CHOOSE
GitHub Actions if your code is on GitHub. The marketplace is a real advantage — almost any integration already exists as an action.
GitLab CI if you self-host GitLab, want everything (repo, CI, registry, issues) in one tool, or need the environments/rollback UI. File-type variables alone make SSH-key handling noticeably cleaner.
Both are entirely capable. The right answer is almost always "whichever platform hosts your code" — running CI on a different platform than your repository adds a webhook integration and a second set of credentials for no benefit.
Useful predefined variables
| Variable | Value |
|---|---|
$CI_COMMIT_SHA | Full commit SHA |
$CI_COMMIT_SHORT_SHA | First 8 characters |
$CI_COMMIT_BRANCH | Branch name |
$CI_COMMIT_TAG | Tag, if the pipeline was triggered by one |
$CI_PIPELINE_SOURCE | push, merge_request_event, schedule, web |
$CI_PROJECT_DIR | Checkout path |
$CI_JOB_ID | Unique job ID |
$CI_REGISTRY_IMAGE | This project's container registry path |
$CI_ENVIRONMENT_NAME | The environment: name: value |
Troubleshooting
| Problem | Cause | Fix |
|---|---|---|
| Pipeline stuck "pending" | No runner with the required tags, or runner offline | Settings → CI/CD → Runners; check tags and status |
This job is stuck | Runner only accepts tagged jobs, job is untagged | Add tags, or enable "run untagged jobs" on the runner |
Permission denied (publickey) | Key not installed, wrong user, bad permissions | Same checklist as Level 18 |
error in libcrypto | SSH key pasted as a normal variable | Use a File-type variable |
| Secret is empty in the job | Variable is Protected, branch is not | Protect the branch, or untick Protected (less safe) |
Host key verification failed | known_hosts missing | Add SSH_KNOWN_HOSTS_FILE as a file variable |
| Two pipelines per push | Branch + MR pipelines both triggered | Add the workflow: rules block |
| Cache never hits | Different runners, or a changing cache key | Use distributed cache (S3), or key on the lockfile |
pnpm: command not found on the server | Non-interactive shell, NVM not loaded | Source NVM; symlink to /usr/local/bin (Level 6) |
| Services unreachable in tests | Wrong hostname | Use the service alias, not 127.0.0.1, with the Docker executor |
| Job times out | Default 1 hour | timeout: 20m on the job |
| Runner runs out of disk | Docker images and caches accumulate | docker system prune on a schedule; gitlab-runner cleanup |
SERVICE HOSTNAMES DIFFER BETWEEN GITHUB AND GITLAB
GitHub Actions service containers are reachable at 127.0.0.1 (ports are mapped to the runner). GitLab's Docker executor puts services on a shared network reachable by their alias — postgres:5432, not 127.0.0.1:5432.
Copying a workflow between the two without changing this gives ECONNREFUSED in tests.
Production Checklist — Level 19
- [ ]
.gitlab-ci.ymlwithtest→build→deploystages - [ ]
workflow: rulesprevents duplicate pipelines - [ ] Runner uses the Docker executor, never
shellon a shared host - [ ] Runner is not on the production server
- [ ]
privileged = falseunless Docker-in-Docker is genuinely required - [ ] SSH key stored as a File-type variable
- [ ] All deploy variables marked Protected
- [ ]
productionbranch is protected in repository settings - [ ] Masked variables verified as actually masked
- [ ]
SSH_KNOWN_HOSTS_FILEset — noStrictHostKeyChecking=no - [ ] Deploy job is
when: manualif you want a human gate - [ ]
environment:set so GitLab tracks deployments and offers rollback - [ ] Artifacts (not cache) used for anything a later stage requires
- [ ]
set -euo pipefailin every remote script - [ ] NVM sourced explicitly on the server
- [ ] Health check after deploy, with logs printed on failure
- [ ]
timeout:set on jobs - [ ] Runner host disk cleaned on a schedule