Level 24 — Production Architecture
The complete design, with an explicit statement of what is public, what is private, and why.
The architecture
Exposure classification
This table is the security model in one place.
| Component | Bind address | Reachable from | Firewall |
|---|---|---|---|
| SSH | 0.0.0.0:22 | Internet (key auth only) | Allowed, rate-limited |
| Nginx HTTP | 0.0.0.0:80 | Internet | Allowed |
| Nginx HTTPS | 0.0.0.0:443 | Internet | Allowed |
| Nuxt SSR | 127.0.0.1:3000 | Nginx only | Not in firewall |
| NestJS API | 127.0.0.1:3001 | Nginx and Nuxt only | Not in firewall |
| PostgreSQL | 127.0.0.1:5432 | NestJS only | Never |
| Redis | 127.0.0.1:6379 | NestJS and Nuxt only | Never |
| Netdata (if used) | 127.0.0.1:19999 | SSH tunnel only | Never |
# SERVER — the one command that verifies this whole table
sudo ss -tulpn | grep -vE "127\.0\.0\.1|\[::1\]"IF THAT COMMAND SHOWS ANYTHING BEYOND 22, 80, AND 443, INVESTIGATE IMMEDIATELY
Run it after every deployment, every package installation, and every Docker Compose change. It is the fastest audit in this guide and it catches the mistakes that matter most.
Why each boundary exists
Two firewalls
| Layer | Runs | Protects against |
|---|---|---|
| Cloud firewall | Provider network, outside your VM | Misconfigured UFW; also stops junk traffic before it costs you CPU |
| UFW | Your kernel | A misconfigured cloud firewall; per-rule granularity |
Neither is redundant: the cloud firewall is the one you can still fix from your phone after locking yourself out of SSH.
Nginx as the only public entrypoint
Every request from the internet passes through one process that handles TLS, applies rate limits and security headers, writes access logs, serves static files, and returns a controlled error when a backend is down. Removing it would mean solving all of that in each application, separately, worse.
127.0.0.1 binding
The most important single control in the architecture. A service bound to loopback is unreachable from the network even if every firewall rule is wrong. It is the layer that does not depend on configuration you might get wrong.
PM2 between the app and the OS
Crash restarts, boot startup, zero-downtime reloads, cluster mode across cores, and log management. Without it, "the site is down" is a manual, human-response problem.
Data flow — a page load
Two details worth noting:
- Nuxt calls the API over loopback, not through the public URL — no DNS, no TLS handshake, no round trip through Nginx. Typically saves 10–50 ms per SSR render (Level 12).
- Static assets never reach Node. On a page with 30 assets, that is 30 requests Nginx handles from disk at a fraction of the cost.
Improvements on the original sketch
The architecture in the brief was close. What this design adds and why:
| Addition | Reason |
|---|---|
| Cloud firewall as a distinct layer | Survives UFW misconfiguration; fixable without SSH |
| fail2ban | Behavioural blocking that firewall rules cannot express |
| Static files served by Nginx | Removes most requests from Node entirely |
| Nuxt → API over loopback | Removes DNS + TLS + proxy from the SSR hot path |
| Redis explicitly shown for the Socket.IO adapter | Cluster mode silently breaks WebSockets without it (Level 10) |
| Off-site backups on a different provider | A local backup does not survive losing the VPS |
| External monitoring | An on-server monitor cannot detect its own server dying |
| PM2 shown as supervising, not in the request path | It is a supervisor, not a proxy — a common misreading of the original diagram |
PM2 IS NOT IN THE REQUEST PATH
The original sketch placed PM2 below the app, as though requests flow through it. They do not. PM2 starts the processes, restarts them when they die, and rolls them during a reload. Traffic goes Nginx → Node directly. Getting this right matters when debugging: a PM2 problem shows up as a process that is not running, never as a slow request.
Resource allocation on a 4 GB / 2 vCPU server
| Component | RAM | Notes |
|---|---|---|
| Nuxt SSR ×2 | ~300 MB | 150 MB each |
| NestJS API ×2 | ~250 MB | 125 MB each |
| PostgreSQL | ~1 GB | shared_buffers = 1GB (25% of RAM) |
| Redis | ~512 MB | maxmemory 512mb |
| Nginx | ~30 MB | Very light |
| System + PM2 | ~400 MB | |
| Total | ~2.5 GB | Leaving ~1.5 GB for page cache and headroom |
| Swap | 2 GB | Airbag for build spikes (Level 4) |
DO NOT BUILD ON THIS SERVER WITHOUT SWAP
nuxt build peaks above 2 GB. With ~1.5 GB free, the OOM killer terminates something — usually PostgreSQL, being the largest process. Either add swap, or build in CI and ship the artifact (Level 18).
Filesystem layout
/
├── etc/
│ ├── nginx/
│ │ ├── nginx.conf
│ │ ├── sites-available/{app,api}.example.com
│ │ ├── sites-enabled/ → symlinks
│ │ └── snippets/{proxy-params,security-headers,ssl-hardening}.conf
│ ├── letsencrypt/live/app.example.com/{fullchain,privkey}.pem
│ ├── postgresql/16/main/{postgresql.conf,pg_hba.conf}
│ ├── redis/redis.conf
│ ├── ssh/sshd_config.d/99-hardening.conf
│ ├── fail2ban/jail.local
│ └── sudoers.d/deploy-services
│
├── var/
│ ├── lib/postgresql/16/main/ ← database files
│ ├── lib/redis/ ← RDB/AOF
│ ├── log/{nginx,postgresql,redis}/
│ └── backups/myapp/ ← local backups (also pushed off-site)
│
└── home/deploy/
├── .ssh/{authorized_keys,github_deploy}
├── .nvm/
├── logs/ ← PM2 output
├── scripts/{backup,health-check,security-audit}.sh
└── apps/myapp/
├── current → releases/20260811-143022-a1b2c3d
├── releases/ ← last 5 kept
├── shared/{.env,uploads,logs}
└── ecosystem.config.cjsshared/ IS WHAT SURVIVES DEPLOYS
.env, user uploads, and logs live outside any release directory and are symlinked in. A deploy replaces current; it never touches shared. This separation is what makes the release-directory pattern safe (Level 20).
When to grow beyond one server
ONE VPS GOES FURTHER THAN PEOPLE EXPECT
A 4 GB / 2 vCPU server running this architecture comfortably handles a few hundred requests per second and tens of thousands of daily users. Most projects never outgrow it. Resist the urge to distribute before you have a measurement telling you to.
Stage 1 — vertical scaling
Resize the VPS. Minutes of downtime, no architectural change. Take this as far as it goes; it is nearly always cheaper than the complexity of the alternatives.
Stage 2 — separate the database
Do this when the database and application compete for RAM. The database moves to its own machine, connected over the provider's private network.
THE DATABASE MOVES TO A PRIVATE IP, NOT A PUBLIC ONE
listen_addresses = '10.0.0.5' # private interface only — never 0.0.0.0# pg_hba.conf
hostssl myapp_production myapp 10.0.0.10/32 scram-sha-256sudo ufw allow from 10.0.0.10 to any port 5432 proto tcpRequire SSL for the connection (sslmode=require in DATABASE_URL) — the private network is not encrypted, and other tenants may share it depending on the provider.
Stage 3 — multiple application servers
This is where the design decisions from earlier chapters pay off:
| Requirement | Why | Where covered |
|---|---|---|
| Sessions in Redis | Servers do not share memory | Level 10 |
| Socket.IO Redis adapter | Broadcasts must cross servers | Level 10 |
| Uploads on object storage | Local disk is not shared | — |
| Rate limiting in Redis | Per-server counters multiply your limit | Level 10 |
| Sticky sessions or WS-only transport | Socket.IO handshake affinity | Level 14 |
| Migrations run once, not per server | Concurrent migrations conflict | Level 12 |
IF YOU BUILT FOR CLUSTER MODE, YOU ARE ALREADY MOSTLY THERE
PM2 cluster mode forces you to solve shared-state problems on day one. A single-server app that runs correctly with 4 workers usually runs correctly across 4 servers. That is the real reason to use cluster mode early — not the CPU utilisation.
Stage 4 — managed services
At some point, running PostgreSQL yourself stops being worth it. Managed databases handle backups, failover, patching, and point-in-time recovery. The cost is real (often 3–5× a VPS) and so is the time saved.
Consider it when: you need high availability with automatic failover, you have compliance requirements around backup retention, or database operations are consuming meaningful engineering time.
Single points of failure
Honest assessment of this architecture:
| SPOF | Impact | Mitigation |
|---|---|---|
| The one server | Total outage | Snapshots + off-site backups + a documented rebuild (Level 28) |
| PostgreSQL | Total outage | Backups; a read replica later |
| Nginx | Total outage | systemd restarts it; nginx -t before every reload |
| DNS provider | Total outage | Use a reliable provider; know how to switch |
| Domain expiry | Total outage | Auto-renew + calendar reminder |
| TLS expiry | Total outage (worse with HSTS) | Auto-renewal + independent monitoring |
| Certificate Authority | Cannot issue new certs | Certificates remain valid for 90 days — time to react |
A SINGLE SERVER MEANS ACCEPTING SOME DOWNTIME
That is a legitimate engineering decision, not a failure. The question is whether you have decided it deliberately and know your recovery time.
What is not acceptable is a single server with no tested backup and no rebuild runbook. That is not "accepting downtime" — it is accepting data loss. (Level 22)
Production Checklist — Level 24
- [ ]
sudo ss -tulpnshows only 22, 80, 443 on public interfaces - [ ] Cloud firewall and UFW both configured
- [ ] Nginx is the only public entrypoint
- [ ] All application and data services bound to
127.0.0.1 - [ ] PM2 cluster mode with ≥2 instances per app
- [ ] Static assets served by Nginx from disk
- [ ] Nuxt calls the API over loopback during SSR
- [ ] Redis backs sessions, cache, rate limits, and the Socket.IO adapter
- [ ] Resource allocation fits in RAM with headroom; swap present if ≤4 GB
- [ ]
shared/holds.env, uploads, and logs, surviving deploys - [ ] Off-site backups on a different provider
- [ ] External monitoring configured
- [ ] All shared state is in Redis or PostgreSQL, never in process memory
- [ ] Single points of failure identified and consciously accepted
- [ ] Rebuild runbook exists and has been tested