Skip to content

Level 24 — Production Architecture ​

The complete design, with an explicit statement of what is public, what is private, and why.

The architecture ​

Exposure classification ​

This table is the security model in one place.

ComponentBind addressReachable fromFirewall
SSH0.0.0.0:22Internet (key auth only)Allowed, rate-limited
Nginx HTTP0.0.0.0:80InternetAllowed
Nginx HTTPS0.0.0.0:443InternetAllowed
Nuxt SSR127.0.0.1:3000Nginx onlyNot in firewall
NestJS API127.0.0.1:3001Nginx and Nuxt onlyNot in firewall
PostgreSQL127.0.0.1:5432NestJS onlyNever
Redis127.0.0.1:6379NestJS and Nuxt onlyNever
Netdata (if used)127.0.0.1:19999SSH tunnel onlyNever
bash
# SERVER — the one command that verifies this whole table
sudo ss -tulpn | grep -vE "127\.0\.0\.1|\[::1\]"

IF THAT COMMAND SHOWS ANYTHING BEYOND 22, 80, AND 443, INVESTIGATE IMMEDIATELY

Run it after every deployment, every package installation, and every Docker Compose change. It is the fastest audit in this guide and it catches the mistakes that matter most.

Why each boundary exists ​

Two firewalls ​

LayerRunsProtects against
Cloud firewallProvider network, outside your VMMisconfigured UFW; also stops junk traffic before it costs you CPU
UFWYour kernelA misconfigured cloud firewall; per-rule granularity

Neither is redundant: the cloud firewall is the one you can still fix from your phone after locking yourself out of SSH.

Nginx as the only public entrypoint ​

Every request from the internet passes through one process that handles TLS, applies rate limits and security headers, writes access logs, serves static files, and returns a controlled error when a backend is down. Removing it would mean solving all of that in each application, separately, worse.

127.0.0.1 binding ​

The most important single control in the architecture. A service bound to loopback is unreachable from the network even if every firewall rule is wrong. It is the layer that does not depend on configuration you might get wrong.

PM2 between the app and the OS ​

Crash restarts, boot startup, zero-downtime reloads, cluster mode across cores, and log management. Without it, "the site is down" is a manual, human-response problem.

Data flow — a page load ​

Two details worth noting:

  • Nuxt calls the API over loopback, not through the public URL — no DNS, no TLS handshake, no round trip through Nginx. Typically saves 10–50 ms per SSR render (Level 12).
  • Static assets never reach Node. On a page with 30 assets, that is 30 requests Nginx handles from disk at a fraction of the cost.

Improvements on the original sketch ​

The architecture in the brief was close. What this design adds and why:

AdditionReason
Cloud firewall as a distinct layerSurvives UFW misconfiguration; fixable without SSH
fail2banBehavioural blocking that firewall rules cannot express
Static files served by NginxRemoves most requests from Node entirely
Nuxt → API over loopbackRemoves DNS + TLS + proxy from the SSR hot path
Redis explicitly shown for the Socket.IO adapterCluster mode silently breaks WebSockets without it (Level 10)
Off-site backups on a different providerA local backup does not survive losing the VPS
External monitoringAn on-server monitor cannot detect its own server dying
PM2 shown as supervising, not in the request pathIt is a supervisor, not a proxy — a common misreading of the original diagram

PM2 IS NOT IN THE REQUEST PATH

The original sketch placed PM2 below the app, as though requests flow through it. They do not. PM2 starts the processes, restarts them when they die, and rolls them during a reload. Traffic goes Nginx → Node directly. Getting this right matters when debugging: a PM2 problem shows up as a process that is not running, never as a slow request.

Resource allocation on a 4 GB / 2 vCPU server ​

ComponentRAMNotes
Nuxt SSR ×2~300 MB150 MB each
NestJS API ×2~250 MB125 MB each
PostgreSQL~1 GBshared_buffers = 1GB (25% of RAM)
Redis~512 MBmaxmemory 512mb
Nginx~30 MBVery light
System + PM2~400 MB
Total~2.5 GBLeaving ~1.5 GB for page cache and headroom
Swap2 GBAirbag for build spikes (Level 4)

DO NOT BUILD ON THIS SERVER WITHOUT SWAP

nuxt build peaks above 2 GB. With ~1.5 GB free, the OOM killer terminates something — usually PostgreSQL, being the largest process. Either add swap, or build in CI and ship the artifact (Level 18).

Filesystem layout ​

/
├── etc/
│   ├── nginx/
│   │   ├── nginx.conf
│   │   ├── sites-available/{app,api}.example.com
│   │   ├── sites-enabled/          → symlinks
│   │   └── snippets/{proxy-params,security-headers,ssl-hardening}.conf
│   ├── letsencrypt/live/app.example.com/{fullchain,privkey}.pem
│   ├── postgresql/16/main/{postgresql.conf,pg_hba.conf}
│   ├── redis/redis.conf
│   ├── ssh/sshd_config.d/99-hardening.conf
│   ├── fail2ban/jail.local
│   └── sudoers.d/deploy-services
│
├── var/
│   ├── lib/postgresql/16/main/     ← database files
│   ├── lib/redis/                  ← RDB/AOF
│   ├── log/{nginx,postgresql,redis}/
│   └── backups/myapp/              ← local backups (also pushed off-site)
│
└── home/deploy/
    ├── .ssh/{authorized_keys,github_deploy}
    ├── .nvm/
    ├── logs/                       ← PM2 output
    ├── scripts/{backup,health-check,security-audit}.sh
    └── apps/myapp/
        ├── current → releases/20260811-143022-a1b2c3d
        ├── releases/               ← last 5 kept
        ├── shared/{.env,uploads,logs}
        └── ecosystem.config.cjs

shared/ IS WHAT SURVIVES DEPLOYS

.env, user uploads, and logs live outside any release directory and are symlinked in. A deploy replaces current; it never touches shared. This separation is what makes the release-directory pattern safe (Level 20).

When to grow beyond one server ​

ONE VPS GOES FURTHER THAN PEOPLE EXPECT

A 4 GB / 2 vCPU server running this architecture comfortably handles a few hundred requests per second and tens of thousands of daily users. Most projects never outgrow it. Resist the urge to distribute before you have a measurement telling you to.

Stage 1 — vertical scaling ​

Resize the VPS. Minutes of downtime, no architectural change. Take this as far as it goes; it is nearly always cheaper than the complexity of the alternatives.

Stage 2 — separate the database ​

Do this when the database and application compete for RAM. The database moves to its own machine, connected over the provider's private network.

THE DATABASE MOVES TO A PRIVATE IP, NOT A PUBLIC ONE

conf
listen_addresses = '10.0.0.5'        # private interface only — never 0.0.0.0
# pg_hba.conf
hostssl  myapp_production  myapp  10.0.0.10/32  scram-sha-256
bash
sudo ufw allow from 10.0.0.10 to any port 5432 proto tcp

Require SSL for the connection (sslmode=require in DATABASE_URL) — the private network is not encrypted, and other tenants may share it depending on the provider.

Stage 3 — multiple application servers ​

This is where the design decisions from earlier chapters pay off:

RequirementWhyWhere covered
Sessions in RedisServers do not share memoryLevel 10
Socket.IO Redis adapterBroadcasts must cross serversLevel 10
Uploads on object storageLocal disk is not shared—
Rate limiting in RedisPer-server counters multiply your limitLevel 10
Sticky sessions or WS-only transportSocket.IO handshake affinityLevel 14
Migrations run once, not per serverConcurrent migrations conflictLevel 12

IF YOU BUILT FOR CLUSTER MODE, YOU ARE ALREADY MOSTLY THERE

PM2 cluster mode forces you to solve shared-state problems on day one. A single-server app that runs correctly with 4 workers usually runs correctly across 4 servers. That is the real reason to use cluster mode early — not the CPU utilisation.

Stage 4 — managed services ​

At some point, running PostgreSQL yourself stops being worth it. Managed databases handle backups, failover, patching, and point-in-time recovery. The cost is real (often 3–5× a VPS) and so is the time saved.

Consider it when: you need high availability with automatic failover, you have compliance requirements around backup retention, or database operations are consuming meaningful engineering time.

Single points of failure ​

Honest assessment of this architecture:

SPOFImpactMitigation
The one serverTotal outageSnapshots + off-site backups + a documented rebuild (Level 28)
PostgreSQLTotal outageBackups; a read replica later
NginxTotal outagesystemd restarts it; nginx -t before every reload
DNS providerTotal outageUse a reliable provider; know how to switch
Domain expiryTotal outageAuto-renew + calendar reminder
TLS expiryTotal outage (worse with HSTS)Auto-renewal + independent monitoring
Certificate AuthorityCannot issue new certsCertificates remain valid for 90 days — time to react

A SINGLE SERVER MEANS ACCEPTING SOME DOWNTIME

That is a legitimate engineering decision, not a failure. The question is whether you have decided it deliberately and know your recovery time.

What is not acceptable is a single server with no tested backup and no rebuild runbook. That is not "accepting downtime" — it is accepting data loss. (Level 22)

Production Checklist — Level 24 ​

  • [ ] sudo ss -tulpn shows only 22, 80, 443 on public interfaces
  • [ ] Cloud firewall and UFW both configured
  • [ ] Nginx is the only public entrypoint
  • [ ] All application and data services bound to 127.0.0.1
  • [ ] PM2 cluster mode with ≥2 instances per app
  • [ ] Static assets served by Nginx from disk
  • [ ] Nuxt calls the API over loopback during SSR
  • [ ] Redis backs sessions, cache, rate limits, and the Socket.IO adapter
  • [ ] Resource allocation fits in RAM with headroom; swap present if ≤4 GB
  • [ ] shared/ holds .env, uploads, and logs, surviving deploys
  • [ ] Off-site backups on a different provider
  • [ ] External monitoring configured
  • [ ] All shared state is in Redis or PostgreSQL, never in process memory
  • [ ] Single points of failure identified and consciously accepted
  • [ ] Rebuild runbook exists and has been tested

Next: Level 25 — Complete From-Zero Deployment →