Level 17 — CI/CD Concepts
Automating the path from git push to running code. This chapter is the vocabulary and the thinking; the next two implement it in GitHub Actions and GitLab CI.
CI and CD
| Term | Stands for | Means |
|---|---|---|
| CI | Continuous Integration | Every push is automatically built and tested against the shared branch |
| CD | Continuous Delivery | Every passing build is ready to deploy; a human presses the button |
| CD | Continuous Deployment | Every passing build deploys automatically, no human step |
The two CDs are genuinely different decisions.
WHICH CD SHOULD YOU CHOOSE?
Continuous Delivery (manual approval) fits a small project without a comprehensive test suite. You keep a human gate, which is honest about the fact that your tests do not prove correctness.
Continuous Deployment (fully automatic) requires real confidence: good test coverage, feature flags for risky changes, monitoring that catches regressions in minutes, and a rollback you have actually practised.
Start with Delivery. Move to Deployment when the manual approval has stopped adding value — when you find yourself clicking "approve" without looking, that gate is theatre and you should either automate it or make it meaningful.
Why bother
Manual deployment is a sequence of ten commands you run from memory over SSH. It works. Then:
| Manual failure mode | What CI/CD does instead |
|---|---|
You forget prisma migrate deploy | The pipeline runs the same steps every time |
| You deploy from a dirty local branch | The pipeline builds from a specific commit in the remote |
| "Works on my machine" | The build runs in a clean, defined environment |
| A colleague deploys differently | One definition, in Git, reviewed like code |
| Deploying is scary, so you batch changes | Small, frequent deploys — each one low-risk |
| No record of what was deployed when | Every run is logged with its commit SHA |
| Broken code reaches production | Lint, typecheck, and tests gate the deploy |
| You are the only person who can deploy | Anyone with merge rights can |
THE REAL BENEFIT IS SMALLER DEPLOYS
When deploying is a 20-minute manual ritual, you deploy once a week and each release contains 40 changes. When something breaks, you have 40 suspects.
When deploying is automatic, you deploy 5 times a day and each release contains 1–2 changes. When something breaks, you know exactly what did it, and reverting costs nothing. Deployment frequency and incident recovery time are strongly correlated — that is the finding behind the DORA metrics.
Anatomy of a pipeline
| Concept | Meaning |
|---|---|
| Pipeline / workflow | The whole automated process |
| Trigger / event | What starts it — a push, a tag, a schedule, a manual click |
| Job | A group of steps running on one runner |
| Step | A single command or action |
| Stage | An ordered group of jobs (GitLab); expressed with needs: in GitHub |
| Runner | The machine executing the job |
| Artifact | A file produced by a job and passed to later jobs |
| Cache | Reusable data (like node_modules) that speeds up future runs |
| Environment | A named deployment target with its own secrets and rules |
| Matrix | Running the same job across several configurations |
The quality gates
Each gate is cheap and catches a different class of problem. Order them fastest-first so failures surface quickly.
Lint
pnpm lintCatches style inconsistency and a surprising number of real bugs: unused variables, unreachable code, missing await on a promise, == where === was meant.
THE MISSING-await RULE ALONE JUSTIFIES LINTING
@typescript-eslint/no-floating-promises catches prisma.user.create(...) without await — code that appears to work in development and silently loses writes under load. That single rule has saved more production incidents than most test suites.
Typecheck
pnpm typecheck # tsc --noEmitTHIS IS A SEPARATE STEP FROM THE BUILD
Many bundlers — esbuild, SWC, Vite — strip types without checking them. nest build and nuxt build can succeed with type errors present. A build that compiles is not a build that is type-correct.
Run tsc --noEmit as its own step. It is the only thing that actually validates your types.
Tests
pnpm test # unit tests
pnpm test:e2e # integration/e2eWHAT TO TEST IF YOU HAVE NO TESTS
Do not aim for coverage percentages. Write tests for:
- Authentication and authorisation — a bug here is a breach
- Payment and billing — a bug here costs money
- Anything you have broken before — a regression test per past incident
- Complex pure functions — pricing, permissions, date arithmetic
Ten tests covering those beat 400 tests asserting that getters return values.
Build
pnpm buildProves the code compiles and bundles. In CI it also produces the artifact you deploy.
Security scanning
pnpm audit --audit-level=highpnpm audit PRODUCES A LOT OF NOISE
Most advisories are in dev dependencies, or in code paths you never execute (a ReDoS in a CLI argument parser is not exploitable by your users). Failing the build on every advisory teaches people to bypass the check.
Fail on high/critical in runtime dependencies. Review the rest weekly. Consider socket.dev or Snyk, which weight advisories by actual exploitability. (Level 23)
Build once, deploy many
REBUILDING PER ENVIRONMENT MEANS YOU DID NOT TEST WHAT YOU SHIPPED
If staging builds artifact A and production builds artifact B from the same commit, they can differ: a floating dependency version, a different Node patch release, a different build-time environment variable baked in.
Build once, from one commit, producing one artifact. Promote that same artifact through environments. Configuration comes from the environment at runtime, never baked into the build.
This is why runtimeConfig in Nuxt matters (Level 8) — it lets one build serve every environment.
Environments
| Environment | Purpose | Data | Who deploys |
|---|---|---|---|
| Local | Development | Fake/seed | You, constantly |
| CI | Automated verification | Ephemeral test DB | The pipeline |
| Staging | Pre-production validation | Anonymised copy of production | Automatic on merge to main |
| Production | Real users | Real | Automatic or gated on merge to production |
STAGING MUST NOT SHARE ANYTHING WITH PRODUCTION
Common and dangerous shortcuts:
- Staging pointed at the production database — one bad migration destroys real data
- Staging using production API keys — test payments become real charges
- Staging sending real email to real user addresses from a data copy
Staging gets its own database, its own API keys (test mode), and an email sink (Mailtrap) or a hard allowlist of recipients. If you copy production data to staging, anonymise it — that is also a legal requirement in most jurisdictions.
Deployment triggers
| Trigger | Config | Use for |
|---|---|---|
Push to main | on: push: branches: [main] | Staging |
Push to production | on: push: branches: [production] | Production |
Tag v* | on: push: tags: ['v*'] | Versioned releases |
| Manual | workflow_dispatch | Hotfixes, re-runs |
| Schedule | on: schedule | Nightly builds, dependency updates |
| Pull request | on: pull_request | Tests only — never deploy |
NEVER DEPLOY FROM pull_request EVENTS
Anyone can open a pull request from a fork. If that workflow could deploy — or even just read your secrets — an attacker's first PR would exfiltrate your entire secret store.
GitHub deliberately withholds secrets from fork PRs. Do not work around this. Never combine pull_request_target with a checkout of the PR's code; that combination runs untrusted code with secrets available and is a well-documented compromise path.
Deploy on push to a protected branch. Test on pull_request.
Deployment architecture
Two deployment models
Model A — build on the server (simplest)
CI validates, then SSHes in and runs the same script you would run by hand.
- ✅ Simple; one script that also works manually
- ✅ No artifact storage needed
- ❌ Server needs build-time RAM (Level 12)
- ❌ Longer deploy window
- ❌ Rollback means rebuilding an old commit
Model B — build in CI, ship the artifact
CI builds, uploads a tarball or Docker image, the server unpacks and switches to it.
- ✅ Server needs no devtools or build RAM
- ✅ Deploy is fast — just unpack and reload
- ✅ Instant rollback: the previous artifact is still on disk
- ✅ Exactly the tested artifact reaches production
- ❌ More moving parts; needs artifact transfer or a registry
START WITH A, MOVE TO B WHEN IT HURTS
Model A is genuinely fine for a small project and is what Level 18 shows first. Move to B when: your builds OOM, deploy time becomes annoying, or you need rollback measured in seconds rather than minutes.
Rollback
A PIPELINE WITHOUT A TESTED ROLLBACK IS HALF A PIPELINE
The question is not "will a bad deploy happen" but "how long until we are back". If the answer involves reading documentation while users are affected, you do not have a rollback plan.
| Strategy | Speed | Requires |
|---|---|---|
| Re-deploy the previous commit | Minutes (full rebuild) | Nothing extra |
| Switch a symlink to the previous release | Seconds | Release-directory layout (Level 20) |
| Re-tag a previous Docker image | Seconds | Registry |
git revert and let CI deploy | Minutes | Nothing extra |
DATABASE MIGRATIONS DO NOT ROLL BACK
This is the constraint that shapes everything. Reverting code to before a migration leaves the schema ahead of the code.
Therefore: make every migration backward-compatible with the previous code version. Additive only — new tables, new nullable columns, new indexes. Anything destructive (drop, rename, NOT NULL on existing data) must be split across two deploys with the expand/contract pattern (Level 9).
Follow this rule and rollback is always just "redeploy the old code". Break it and rollback means restoring a backup — with data loss for everything written since.
PREFER ROLL-FORWARD FOR SMALL FIXES
If the fix is obvious and small, deploying a fix is often faster and less disruptive than rolling back — especially with a fast pipeline. Reserve rollback for "we do not yet understand what broke".
Secrets in the pipeline
Covered in Level 8; the pipeline-specific points:
| Rule | Why |
|---|---|
| Store in the platform's encrypted store | Never in the workflow file, never in the repo |
| Use a dedicated SSH key for CI | Revocable independently of your personal key |
Restrict the CI key in authorized_keys | from=, command=, no-port-forwarding (Level 1) |
| Scope secrets to an environment | So a feature branch cannot reach production credentials |
Never echo a secret | Masking is best-effort and defeated by transformation |
| Rotate on any team change | An ex-colleague may have copied them |
THE STRONGEST CI KEY RESTRICTION
In /home/deploy/.ssh/authorized_keys:
command="/home/deploy/apps/myapp/deploy.sh",no-agent-forwarding,no-port-forwarding,no-pty ssh-ed25519 AAAA... ci-deployThis key can run only that script, no matter what command the client sends. A stolen CI key can trigger a deploy — it cannot get a shell, read files, or install anything.
Notifications
A pipeline nobody watches is a pipeline that fails silently.
| Event | Notify | Where |
|---|---|---|
| Deploy succeeded | Low priority | Team chat |
| Deploy failed | High | Chat + the person who pushed |
| Health check failed after deploy | Critical | Page someone |
| Nightly dependency scan found a critical CVE | High | Chat + issue |
ALERT FATIGUE IS A REAL FAILURE MODE
If every successful deploy pings the channel, people mute the channel, and then miss the failure. Notify loudly on failure, quietly (or not at all) on success.
Pipeline speed
A pipeline slower than ~10 minutes stops being a feedback loop and becomes an interruption.
| Technique | Saving |
|---|---|
| Cache the pnpm store | 30–60s per run |
| Run lint/typecheck/test in parallel jobs | Wall time = the slowest, not the sum |
Shallow clone (fetch-depth: 1) | Seconds on a large repo |
Skip CI for docs-only changes (paths-ignore) | Whole runs |
Cancel superseded runs (concurrency) | Frees runners |
| Reuse the Docker build cache | Minutes |
THE SINGLE BIGGEST WIN IS PARALLELISM
Lint, typecheck, and unit tests are independent. Run them as three concurrent jobs and your gate takes as long as the slowest one, not all three added together. Only the deploy job needs needs: [lint, typecheck, test].
What CI/CD does not solve
Being honest about the limits:
- It does not make bad code good. It runs the checks you wrote. Weak tests give false confidence.
- It does not prevent bad migrations. Only design discipline does.
- It does not replace monitoring. A deploy can pass every check and still break under real traffic (Level 21).
- It adds a new attack surface. Your CI system can deploy to production; treat its secrets and permissions as production-critical.
- It can become a bottleneck. A flaky test that fails 20% of the time trains people to re-run until green — at which point the gate is worthless. Fix or delete flaky tests immediately.
Production Checklist — Level 17
- [ ] Pipeline defined in code, committed to the repository
- [ ] Triggers on
pushto a protected branch; never deploys frompull_request - [ ] Gates: install → lint → typecheck → test → build
- [ ] Typecheck is a separate step from the build
- [ ] Independent gates run in parallel
- [ ] Build once; promote the same artifact through environments
- [ ] Staging has its own database, its own API keys, and no real email delivery
- [ ] Secrets in the platform's encrypted store, scoped to an environment
- [ ] Dedicated, restricted CI SSH key
- [ ] Migrations are additive and backward-compatible with the previous release
- [ ] Health check after deploy; failure is loud
- [ ] Rollback procedure documented and practised
- [ ] Failure notifications reach a human; success notifications are quiet
- [ ] Pipeline completes in under ~10 minutes
- [ ] No flaky tests tolerated
- [ ] Deployed commit SHA visible from a health endpoint