Skip to content

Level 19 — GitLab CI/CD and GitLab Runner ​

The same deployment pipeline in GitLab, plus how to install and configure your own runner.

How GitLab CI works ​

GitLab looks for .gitlab-ci.yml in the repository root. Any push triggers a pipeline made of stages, each containing jobs, executed by runners.

TermMeaning
PipelineOne complete run
StageAn ordered phase. All jobs in a stage run in parallel; the next stage waits.
JobOne unit of work with a script:
RunnerThe agent that executes jobs
ExecutorHow the runner executes — shell, Docker, Kubernetes
ArtifactFiles a job produces, passed to later stages
CacheReused data between runs (dependencies)
EnvironmentA named deployment target with history and rollback UI

The complete .gitlab-ci.yml ​

yaml
# .gitlab-ci.yml

stages:
  - test
  - build
  - deploy

default:
  image: node:22-slim
  interruptible: true
  before_script:
    - corepack enable
    - corepack prepare pnpm@9.12.0 --activate
    - pnpm install --frozen-lockfile

variables:
  PNPM_STORE: .pnpm-store
  GIT_DEPTH: "20"
  FF_USE_FASTZIP: "true"

cache:
  key:
    files:
      - pnpm-lock.yaml
  paths:
    - .pnpm-store
  policy: pull-push

workflow:
  rules:
    - if: $CI_PIPELINE_SOURCE == "merge_request_event"
    - if: $CI_COMMIT_BRANCH == "main"
    - if: $CI_COMMIT_BRANCH == "production"
    - when: never

# ---------------------------------------------------------------
# Stage: test
# ---------------------------------------------------------------
lint:
  stage: test
  script:
    - pnpm lint

typecheck:
  stage: test
  script:
    - pnpm --filter backend prisma generate
    - pnpm typecheck

test:
  stage: test
  services:
    - name: postgres:16-alpine
      alias: postgres
    - name: redis:7-alpine
      alias: redis
  variables:
    POSTGRES_USER: test
    POSTGRES_PASSWORD: test
    POSTGRES_DB: test_db
    POSTGRES_HOST_AUTH_METHOD: trust
    DATABASE_URL: "postgresql://test:test@postgres:5432/test_db"
    REDIS_URL: "redis://redis:6379/0"
    JWT_SECRET: "test-secret-at-least-32-characters-long-xx"
    NODE_ENV: test
  script:
    - pnpm --filter backend prisma generate
    - pnpm --filter backend prisma migrate deploy
    - pnpm test
  coverage: '/All files[^|]*\|[^|]*\s+([\d\.]+)/'
  artifacts:
    when: always
    reports:
      junit: backend/junit.xml
    expire_in: 1 week

# ---------------------------------------------------------------
# Stage: build
# ---------------------------------------------------------------
build:
  stage: build
  variables:
    NUXT_PUBLIC_API_BASE: "https://api.example.com"
    NUXT_PUBLIC_SITE_URL: "https://app.example.com"
  script:
    - pnpm --filter backend prisma generate
    - pnpm build
  artifacts:
    name: "build-$CI_COMMIT_SHORT_SHA"
    paths:
      - backend/dist/
      - frontend/.output/
    expire_in: 1 week
  rules:
    - if: $CI_COMMIT_BRANCH == "production"
    - if: $CI_COMMIT_BRANCH == "main"

# ---------------------------------------------------------------
# Stage: deploy
# ---------------------------------------------------------------
.deploy_template: &deploy_template
  stage: deploy
  image: alpine:3.20
  before_script:
    - apk add --no-cache openssh-client curl
    - mkdir -p ~/.ssh && chmod 700 ~/.ssh
    - cp "$SSH_PRIVATE_KEY" ~/.ssh/id_ed25519
    - chmod 600 ~/.ssh/id_ed25519
    - cp "$SSH_KNOWN_HOSTS_FILE" ~/.ssh/known_hosts
    - chmod 644 ~/.ssh/known_hosts

deploy:production:
  <<: *deploy_template
  environment:
    name: production
    url: https://app.example.com
  rules:
    - if: $CI_COMMIT_BRANCH == "production"
      when: manual
      allow_failure: false
  script:
    - |
      ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<EOF
        set -euo pipefail
        export NVM_DIR="\$HOME/.nvm"
        [ -s "\$NVM_DIR/nvm.sh" ] && . "\$NVM_DIR/nvm.sh"

        cd "$DEPLOY_PATH"
        echo "==> Previous: \$(git rev-parse --short HEAD)"

        git fetch origin --prune
        git checkout production
        git reset --hard origin/production
        git clean -fd

        echo "==> New: \$(git rev-parse --short HEAD)"

        pnpm install --frozen-lockfile
        pnpm --filter backend prisma generate
        pnpm --filter backend prisma migrate deploy
        pnpm --filter backend build
        pnpm --filter frontend build

        export GIT_COMMIT="$CI_COMMIT_SHORT_SHA"
        pm2 reload ecosystem.config.cjs --update-env
        pm2 save
      EOF
    - |
      ssh -p "$SSH_PORT" "$SSH_USER@$SSH_HOST" bash -s <<'EOF'
        set -euo pipefail
        for i in $(seq 1 15); do
          if curl -fsS http://127.0.0.1:3001/api/health > /dev/null; then
            echo "✅ API healthy"; exit 0
          fi
          sleep 3
        done
        echo "❌ Health check failed"
        pm2 logs --err --lines 60 --nostream
        exit 1
      EOF
    - curl -fsS -o /dev/null -w "app: %{http_code}\n" https://app.example.com

Explaining the key parts ​

default: and before_script ​

yaml
default:
  image: node:22-slim
  before_script:
    - corepack enable
    - pnpm install --frozen-lockfile

Applies to every job unless overridden. The image: is the Docker image the job runs inside (with the Docker executor).

before_script RUNS FOR EVERY JOB, INCLUDING DEPLOY

The deploy job uses alpine:3.20 and does not need Node, so it overrides before_script entirely. Forgetting to override means a pointless pnpm install in a container without pnpm — and a confusing failure.

workflow: rules — controlling when pipelines run ​

yaml
workflow:
  rules:
    - if: $CI_PIPELINE_SOURCE == "merge_request_event"
    - if: $CI_COMMIT_BRANCH == "main"
    - if: $CI_COMMIT_BRANCH == "production"
    - when: never

THIS PREVENTS DUPLICATE PIPELINES

Without it, pushing to a branch with an open merge request creates two pipelines — one for the branch, one for the MR. You burn double the runner minutes and get confusing duplicate statuses. This block is the standard fix.

cache vs artifacts ​

cacheartifacts
PurposeSpeed up future runsPass files between stages
Guaranteed present❌ No — best effort✅ Yes
ScopeShared across pipelinesThis pipeline
Use for.pnpm-store, node_modulesdist/, .output/, test reports

NEVER CACHE SOMETHING YOU CANNOT REGENERATE

Cache is explicitly best-effort — the runner may evict it, and on a different runner it may not exist at all. If your build depends on cached output, it will fail unpredictably. Use artifacts for anything a later stage requires.

yaml
cache:
  key:
    files:
      - pnpm-lock.yaml

Keying on the lockfile means the cache invalidates exactly when dependencies change — not on every commit, and not never.

rules — when a job runs ​

yaml
  rules:
    - if: $CI_COMMIT_BRANCH == "production"
      when: manual
      allow_failure: false
ClauseEffect
if:Condition using CI variables
when: manualShows a ▶ button — a human must click
when: on_successDefault — runs if earlier stages passed
when: alwaysRuns even if earlier stages failed
allow_failure: falseA manual job that is skipped blocks the pipeline

rules REPLACED only/except

only: and except: still work but are legacy and cannot express everything rules can. Do not mix them in one job — GitLab rejects that. Use rules throughout.

when: manual IS GITLAB'S DEPLOYMENT GATE

This is the equivalent of GitHub's environment approval. Everything runs automatically up to the deploy, which waits for a click. Combined with protected environments (Settings → CI/CD → Protected environments), you can restrict who is allowed to click.

YAML anchors ​

yaml
.deploy_template: &deploy_template
  stage: deploy
  before_script: [...]

deploy:production:
  <<: *deploy_template

A key starting with . is a hidden job — a template, never executed. &name defines an anchor, <<: *name merges it in. This is how you share configuration between deploy:staging and deploy:production without duplication.

environment: ​

yaml
  environment:
    name: production
    url: https://app.example.com

GitLab tracks deployments per environment, giving you Operations → Environments with full history and a Rollback button that re-runs the deploy job for a previous commit.

THE ROLLBACK BUTTON RE-RUNS THE OLD PIPELINE'S DEPLOY JOB

It does not undo database migrations. If the release you are rolling back from included a destructive migration, the old code will fail against the new schema. Additive-only migrations are what make this button safe (Level 17).

CI/CD variables ​

Settings → CI/CD → Variables.

OptionMeaningUse
ProtectedOnly exposed on protected branches/tags✅ All deploy secrets
MaskedReplaced with [MASKED] in logs✅ Where the value qualifies
FileValue written to a temp file; the variable holds the path✅ SSH keys, certificates
Expanded$OTHER_VAR inside the value is expandedUsually leave off for secrets
Environment scopeRestrict to production, staging, etc.✅ Isolate environments

Variables for this pipeline:

VariableTypeProtectedValue
SSH_PRIVATE_KEYFile✅Contents of the deploy private key
SSH_KNOWN_HOSTS_FILEFile✅ssh-keyscan output
SSH_HOSTVariable✅203.0.113.10
SSH_USERVariable✅deploy
SSH_PORTVariable✅22
DEPLOY_PATHVariable✅/home/deploy/apps/myapp

USE "FILE" TYPE FOR SSH KEYS

Multi-line values cannot be masked, and writing one with echo "$KEY" > id_ed25519 mangles newlines — producing error in libcrypto.

File-type variables solve both: GitLab writes the value to a temp file and sets the variable to its path.

yaml
- cp "$SSH_PRIVATE_KEY" ~/.ssh/id_ed25519
- chmod 600 ~/.ssh/id_ed25519

ALWAYS TICK "PROTECTED" FOR DEPLOY SECRETS

Without it, any branch — including one pushed by a contributor with limited access — can read your production SSH key by adding a job that prints it. Protected variables are only exposed on protected branches, which only maintainers can push to.

Then protect the branch: Settings → Repository → Protected branches → production.

MASKING REQUIREMENTS ARE STRICT

A value can only be masked if it is at least 8 characters, single-line, and uses only base64 characters. Values that fail these rules are silently not masked — GitLab shows a warning when you save, which is easy to miss. Check after saving.

GitLab Runner ​

Shared vs self-hosted ​

GitLab.com shared runnersSelf-hosted runner
SetupNoneInstall and register
CostFree minutes, then paidYour server's cost
SpeedCold start each jobWarm caches, faster
Network access to your VPSPublic internet onlyCan be on the same private network
MaintenanceNoneYours
SecurityGitLab's isolationYour responsibility

DO NOT RUN THE CI RUNNER ON YOUR PRODUCTION SERVER

It is tempting — the runner is already where you want to deploy, so no SSH keys are needed.

Why not:

  • CI jobs run untrusted code — every dependency's install script, every branch anyone pushes
  • With the shell executor, that code runs as the runner's user on your production box
  • Builds consume the CPU and RAM your application needs, causing latency spikes or OOM kills
  • A compromised dependency in CI becomes a compromised production server

If you must self-host, use a separate, cheap VPS for the runner and have it SSH into production like any other client.

Installing on Ubuntu 24.04 ​

bash
# ON THE RUNNER SERVER (not production)
curl -L "https://packages.gitlab.com/install/repositories/runner/gitlab-runner/script.deb.sh" | sudo bash
sudo apt install -y gitlab-runner
sudo systemctl status gitlab-runner

Registering ​

Get a runner authentication token: project → Settings → CI/CD → Runners → New project runner.

bash
# RUNNER SERVER
sudo gitlab-runner register \
  --non-interactive \
  --url "https://gitlab.com/" \
  --token "glrt-xxxxxxxxxxxxxxxx" \
  --executor "docker" \
  --docker-image "node:22-slim" \
  --docker-privileged=false \
  --docker-volumes "/cache" \
  --description "prod-deploy-runner"
FlagMeaning
--urlYour GitLab instance
--tokenRunner authentication token (the old --registration-token is deprecated)
--executorHow jobs run — see below
--docker-imageDefault image for jobs that do not specify one
--docker-privilegedKeep false unless you need Docker-in-Docker

Configuration lands in /etc/gitlab-runner/config.toml:

toml
concurrent = 4
check_interval = 3

[[runners]]
  name = "prod-deploy-runner"
  url = "https://gitlab.com/"
  token = "glrt-..."
  executor = "docker"
  [runners.docker]
    image = "node:22-slim"
    privileged = false
    volumes = ["/cache"]
    shm_size = 268435456
  [runners.cache]
    Type = "local"
bash
sudo systemctl restart gitlab-runner
sudo gitlab-runner verify
sudo gitlab-runner list

Executors ​

ExecutorHow it runs jobsIsolationUse when
shellDirectly on the host, as the gitlab-runner user❌ NoneNever on a shared or production machine
dockerEach job in a fresh container✅ GoodDefault choice
docker+machineProvisions a VM per job✅ ExcellentAutoscaling
kubernetesA pod per job✅ ExcellentYou already run Kubernetes

THE SHELL EXECUTOR RUNS UNTRUSTED CODE AS A REAL USER

Every npm install executes arbitrary postinstall scripts from your dependency tree, as the gitlab-runner user, on that host, with no isolation. Files persist between jobs, so one job can plant something the next job runs.

Use the Docker executor. Each job gets a fresh container that is destroyed afterwards.

privileged = true IS ROOT ON THE HOST

Required for Docker-in-Docker builds, and it grants the container effectively full host access. If you need to build images, prefer Kaniko or BuildKit in rootless mode over privileged DinD.

Tags ​

yaml
deploy:production:
  tags:
    - production-deploy
bash
sudo gitlab-runner register ... --tag-list "production-deploy"

A job with tags: runs only on a runner carrying all of them.

A JOB WITH TAGS AND NO MATCHING RUNNER HANGS FOREVER

It sits in "pending" with no error, waiting for a runner that will never appear. Check Settings → CI/CD → Runners that a runner with that tag exists and is online (green).

Conversely, if your runner is registered to only run tagged jobs (the default for project runners in some configurations), untagged jobs never get picked up.

Runner user and permissions ​

bash
# RUNNER SERVER — the shell executor's user, if you use it
sudo usermod -aG docker gitlab-runner    # ⚠️ equivalent to root

ADDING gitlab-runner TO THE docker GROUP IS ROOT ACCESS

Any CI job can then run docker run -v /:/host and own the machine (Level 11). Combined with the shell executor and untrusted branches, this is a full compromise path.

Acceptable only on a disposable, isolated runner host that touches nothing else.

GitHub Actions vs GitLab CI ​

GitHub ActionsGitLab CI
Config file.github/workflows/*.yml (multiple).gitlab-ci.yml (one, with include:)
ModelJobs with needs: (a DAG)Stages, with needs: for a DAG
ReuseMarketplace actions (huge ecosystem)include:, templates, YAML anchors
Manual gateEnvironment protection ruleswhen: manual + protected environments
SecretsRepository / environment / org secretsVariables with protected/masked/file
File-type secrets❌ Not built in✅ Yes — much better for SSH keys
Self-hosted runnersYesYes — historically stronger
Built-in registryGHCRBuilt-in container registry per project
Deployment history + rollback UIBasic✅ Environments view with rollback
Free tier (public repos)GenerousGenerous
Free tier (private)2,000 min/month400 min/month (free tier)
Ecosystem✅ Far largerSmaller
Self-hostable platformGitHub Enterprise (paid)✅ GitLab CE (free)

WHICH TO CHOOSE

GitHub Actions if your code is on GitHub. The marketplace is a real advantage — almost any integration already exists as an action.

GitLab CI if you self-host GitLab, want everything (repo, CI, registry, issues) in one tool, or need the environments/rollback UI. File-type variables alone make SSH-key handling noticeably cleaner.

Both are entirely capable. The right answer is almost always "whichever platform hosts your code" — running CI on a different platform than your repository adds a webhook integration and a second set of credentials for no benefit.

Useful predefined variables ​

VariableValue
$CI_COMMIT_SHAFull commit SHA
$CI_COMMIT_SHORT_SHAFirst 8 characters
$CI_COMMIT_BRANCHBranch name
$CI_COMMIT_TAGTag, if the pipeline was triggered by one
$CI_PIPELINE_SOURCEpush, merge_request_event, schedule, web
$CI_PROJECT_DIRCheckout path
$CI_JOB_IDUnique job ID
$CI_REGISTRY_IMAGEThis project's container registry path
$CI_ENVIRONMENT_NAMEThe environment: name: value

Troubleshooting ​

ProblemCauseFix
Pipeline stuck "pending"No runner with the required tags, or runner offlineSettings → CI/CD → Runners; check tags and status
This job is stuckRunner only accepts tagged jobs, job is untaggedAdd tags, or enable "run untagged jobs" on the runner
Permission denied (publickey)Key not installed, wrong user, bad permissionsSame checklist as Level 18
error in libcryptoSSH key pasted as a normal variableUse a File-type variable
Secret is empty in the jobVariable is Protected, branch is notProtect the branch, or untick Protected (less safe)
Host key verification failedknown_hosts missingAdd SSH_KNOWN_HOSTS_FILE as a file variable
Two pipelines per pushBranch + MR pipelines both triggeredAdd the workflow: rules block
Cache never hitsDifferent runners, or a changing cache keyUse distributed cache (S3), or key on the lockfile
pnpm: command not found on the serverNon-interactive shell, NVM not loadedSource NVM; symlink to /usr/local/bin (Level 6)
Services unreachable in testsWrong hostnameUse the service alias, not 127.0.0.1, with the Docker executor
Job times outDefault 1 hourtimeout: 20m on the job
Runner runs out of diskDocker images and caches accumulatedocker system prune on a schedule; gitlab-runner cleanup

SERVICE HOSTNAMES DIFFER BETWEEN GITHUB AND GITLAB

GitHub Actions service containers are reachable at 127.0.0.1 (ports are mapped to the runner). GitLab's Docker executor puts services on a shared network reachable by their alias — postgres:5432, not 127.0.0.1:5432.

Copying a workflow between the two without changing this gives ECONNREFUSED in tests.

Production Checklist — Level 19 ​

  • [ ] .gitlab-ci.yml with test → build → deploy stages
  • [ ] workflow: rules prevents duplicate pipelines
  • [ ] Runner uses the Docker executor, never shell on a shared host
  • [ ] Runner is not on the production server
  • [ ] privileged = false unless Docker-in-Docker is genuinely required
  • [ ] SSH key stored as a File-type variable
  • [ ] All deploy variables marked Protected
  • [ ] production branch is protected in repository settings
  • [ ] Masked variables verified as actually masked
  • [ ] SSH_KNOWN_HOSTS_FILE set — no StrictHostKeyChecking=no
  • [ ] Deploy job is when: manual if you want a human gate
  • [ ] environment: set so GitLab tracks deployments and offers rollback
  • [ ] Artifacts (not cache) used for anything a later stage requires
  • [ ] set -euo pipefail in every remote script
  • [ ] NVM sourced explicitly on the server
  • [ ] Health check after deploy, with logs printed on failure
  • [ ] timeout: set on jobs
  • [ ] Runner host disk cleaned on a schedule

Next: Level 20 — Deployment Strategies →