Level 21 — Logging and Monitoring
The difference between "working" and "production-ready" is largely this chapter. Without monitoring, you find out about outages from users.
Where every log lives
| Source | Location | View with |
|---|---|---|
| systemd services | journald (binary) | journalctl -u <service> |
| Nginx access | /var/log/nginx/access.log | tail -f |
| Nginx error | /var/log/nginx/error.log | tail -f |
| PM2 apps | /home/deploy/logs/*.log, ~/.pm2/logs/ | pm2 logs |
| PM2 daemon | ~/.pm2/pm2.log | tail |
| PostgreSQL | /var/log/postgresql/postgresql-16-main.log | tail -f |
| Redis | /var/log/redis/redis-server.log | tail -f |
| Docker containers | Docker's JSON log driver | docker logs |
| Authentication | /var/log/auth.log + journald | journalctl -u ssh |
| Kernel / OOM | dmesg, journald | journalctl -k |
| fail2ban | /var/log/fail2ban.log | tail -f |
| Unattended upgrades | /var/log/unattended-upgrades/ | tail |
| UFW blocks | /var/log/syslog (UFW BLOCK) | grep |
journalctl
systemd's journal is the central log for every service it manages.
# SERVER
journalctl -u nginx # one service
journalctl -u nginx -f # follow
journalctl -u nginx -n 100 # last 100 lines
journalctl -u nginx --since "1 hour ago"
journalctl -u nginx --since "2026-08-11 09:00" --until "2026-08-11 10:00"
journalctl -u nginx -p err # errors and worse
journalctl -u nginx --no-pager # no interactive pager (scripts)
journalctl -u nginx -o json-pretty # structured output
journalctl -k # kernel messages
journalctl -b # this boot
journalctl -b -1 # previous boot — what happened before the crash?
journalctl --list-boots
journalctl -xe # recent entries with explanations
journalctl --disk-usage
journalctl --vacuum-size=500M # trim the journal
journalctl --vacuum-time=14d| Priority | -p value |
|---|---|
| emerg | 0 |
| alert | 1 |
| crit | 2 |
| err | 3 |
| warning | 4 |
| notice | 5 |
| info | 6 |
| debug | 7 |
journalctl -b -1 -p err AFTER AN UNEXPLAINED REBOOT
Shows the errors from the boot before this one. If the server restarted on its own — OOM, kernel panic, provider maintenance, unattended-upgrades reboot — the evidence is there.
THE JOURNAL IS VOLATILE BY DEFAULT ON SOME IMAGES
If /var/log/journal/ does not exist, logs live in /run and are lost on reboot — exactly when you need them.
sudo mkdir -p /var/log/journal
sudo systemd-tmpfiles --create --prefix /var/log/journal
sudo systemctl restart systemd-journald
journalctl --disk-usageThen cap it in /etc/systemd/journald.conf:
[Journal]
Storage=persistent
SystemMaxUse=1G
MaxRetentionSec=1monthApplication logs
# SERVER
pm2 logs # all apps, live
pm2 logs api --lines 200
pm2 logs api --err # errors only
pm2 logs api --nostream --lines 100 # print and exit — for scripts
pm2 flush # truncate all
pm2 logs api --raw # no PM2 prefix — needed for jqStructured logging
JSON LOGS TURN GREP INTO QUERIES
// NestJS with nestjs-pino
import { LoggerModule } from 'nestjs-pino';
LoggerModule.forRoot({
pinoHttp: {
level: process.env.LOG_LEVEL ?? 'info',
genReqId: (req) => req.headers['x-request-id'] ?? randomUUID(),
redact: {
paths: ['req.headers.authorization', 'req.headers.cookie', 'req.body.password'],
remove: true,
},
serializers: {
req: (req) => ({ method: req.method, url: req.url, id: req.id }),
},
},
});Now:
pm2 logs api --raw --nostream | jq 'select(.level >= 50)' # errors
pm2 logs api --raw --nostream | jq -r 'select(.reqId=="abc-123")' # one request
pm2 logs api --raw --nostream | jq -r '.msg' | sort | uniq -c | sort -rn # most commonThe request ID is the highest-value field: it lets you reconstruct everything that happened during one request, across every log line it produced.
NEVER LOG SECRETS OR PERSONAL DATA
Passwords, tokens, full credit card numbers, and API keys must never reach a log file. Logs are backed up, shipped to third-party services, and read by more people than you think.
The redact config above is not optional. Also check that your error reporter (Sentry) scrubs Authorization headers, cookies, and request bodies.
Log levels
| Level | Use for | Production |
|---|---|---|
error | Something broke and needs attention | ✅ |
warn | Recoverable problem, degraded behaviour | ✅ |
info | Significant events (startup, deploy, job completed) | ✅ |
debug | Detailed flow | ❌ Off |
trace | Everything | ❌ Off |
debug IN PRODUCTION FILLS YOUR DISK AND SLOWS THE APP
Every log line is a synchronous-ish write and a disk operation. Debug-level logging on a busy API can produce gigabytes per day and measurably increase latency. Keep it at info, and make the level an environment variable so you can raise it temporarily during an incident.
Nginx logs
# SERVER
sudo tail -f /var/log/nginx/access.log
sudo tail -f /var/log/nginx/error.log
sudo tail -f /var/log/nginx/api.example.com.error.logWith the log_format from Level 14, each line ends with rt= and urt=:
# Slowest requests
sudo awk '{print $(NF-1), $7}' /var/log/nginx/access.log | sort -rn | head -20
# Status code distribution
sudo awk '{print $9}' /var/log/nginx/access.log | sort | uniq -c | sort -rn
# 5xx in the last hour
sudo grep " $(date -u '+%d/%b/%Y:%H')" /var/log/nginx/access.log | awk '$9 ~ /^5/' | tail -30
# Top clients
sudo awk '{print $1}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -20
# Requests per minute — traffic spike detection
sudo awk '{print substr($4,2,17)}' /var/log/nginx/access.log | uniq -c | tail -30rt vs urt TELLS YOU WHERE THE TIME WENT
rt($request_time) — total, from first byte received to last byte senturt($upstream_response_time) — how long your app took
rt=5.2 urt=0.05 → your app was fast; the client's connection was slow (mobile, large upload). Not your problem.
rt=5.2 urt=5.1 → your app was slow. This is your problem. Go look at the database.
Docker logs
docker logs -f --tail 100 myapp-postgres
docker compose logs -f api
docker compose logs --since 30mDOCKER LOGS ARE UNBOUNDED BY DEFAULT
A chatty container fills /var/lib/docker/containers/ until the disk is full, at which point PostgreSQL cannot write, Nginx cannot log, and everything fails at once (Level 11).
// /etc/docker/daemon.json
{ "log-driver": "json-file", "log-opts": { "max-size": "10m", "max-file": "3" } }Applies to newly created containers only. Recreate existing ones.
Log rotation
Without rotation, disk fills. Check every log producer:
# SERVER
sudo du -sh /var/log/* | sort -rh | head -20
du -sh ~/.pm2/logs /home/deploy/logs
docker system df
journalctl --disk-usage| Producer | Rotation |
|---|---|
| Nginx | /etc/logrotate.d/nginx (package default: daily, 14 days) |
| PostgreSQL | /etc/logrotate.d/postgresql-common |
| systemd journal | SystemMaxUse in journald.conf |
| Docker | log-opts in daemon.json |
| PM2 | pm2 install pm2-logrotate — not configured by default |
# SERVER
pm2 install pm2-logrotate
pm2 set pm2-logrotate:max_size 10M
pm2 set pm2-logrotate:retain 14
pm2 set pm2-logrotate:compress true
pm2 set pm2-logrotate:rotateInterval '0 0 * * *'Custom rotation for your own log files:
# SERVER
sudo nano /etc/logrotate.d/myapp/home/deploy/logs/*.log {
daily
rotate 14
compress
delaycompress
missingok
notifempty
copytruncate
su deploy deploy
}copytruncate vs create
create renames the file and makes a new one — but a running process holding the old file descriptor keeps writing to the renamed file, which then gets deleted. Logs vanish.
copytruncate copies the content and truncates the original in place, so the process's file descriptor stays valid. Use it for any process that does not reopen its log file on SIGHUP. There is a tiny race where lines written during the copy can be lost; that is an acceptable trade.
sudo logrotate -d /etc/logrotate.d/myapp # dry run
sudo logrotate -f /etc/logrotate.d/myapp # force nowResource monitoring
CPU and load
# SERVER
uptime 14:32:01 up 12 days, 3:14, 1 user, load average: 0.52, 0.48, 0.41INTERPRETING LOAD AVERAGE
Load is the number of processes running or waiting (for CPU or disk), averaged over 1, 5, and 15 minutes.
Compare it to your core count: nproc.
- 2 cores, load 1.0 → 50% utilised. Fine.
- 2 cores, load 2.0 → fully utilised. At capacity.
- 2 cores, load 8.0 → heavily overloaded; requests are queueing.
Read the trend, not the instant. 0.5 2.0 4.0 means load is falling (1-min lowest). 4.0 2.0 0.5 means it is rising fast — investigate now.
High load with low CPU usage means I/O wait — check disk with iostat.
htop # sudo apt install htop
top
ps aux --sort=-%cpu | head -10
mpstat 1 5 # sudo apt install sysstatMemory
free -h total used free shared buff/cache available
Mem: 3.8Gi 1.9Gi 234Mi 12Mi 1.7Gi 1.7Gi
Swap: 2.0Gi 12Mi 2.0GiREAD available, NOT free
Linux uses spare RAM for disk cache, which is why free looks alarming. That cache is released instantly when a program needs memory.
available is the real number. Below ~10% of total, you are in trouble.
Watch swap too: a small amount used is fine, but active swapping (si/so columns in vmstat 1) means severe pressure and terrible latency.
ps aux --sort=-%mem | head -10
vmstat 1 5
sudo smem -tk 2>/dev/null | tail -5 # more accurate per-process accountingThe OOM killer
WHEN THE SERVER RUNS OUT OF MEMORY, THE KERNEL KILLS SOMETHING
It picks the process with the highest "badness" score — usually the largest, which is Node or PostgreSQL. Your app disappears with no application-level error.
Detect it:
sudo dmesg -T | grep -i "killed process"
sudo journalctl -k --since "1 hour ago" | grep -i "out of memory"Out of memory: Killed process 12345 (node) total-vm:2841232kB, anon-rss:1834920kBThat is definitive. Fixes: add swap (Level 4), reduce PM2 instances, set max_memory_restart lower, move builds to CI, or resize the server.
Protect PostgreSQL specifically:
sudo systemctl edit postgresql[Service]
OOMScoreAdjust=-500This makes the kernel prefer killing something else. Your database is the hardest thing to recover.
Disk
df -h
df -i # inodes — can fill with bytes free
du -h --max-depth=1 / 2>/dev/null | sort -rh | head -20
du -sh /var/log /var/lib/docker /home/deploy 2>/dev/null
ncdu / # interactive — sudo apt install ncduA FULL DISK BREAKS EVERYTHING AT ONCE, CONFUSINGLY
PostgreSQL stops accepting writes, Nginx cannot log (and may stop serving), PM2 cannot write logs, builds fail, apt fails, and even SSH login can fail because it cannot write to ~/.bash_history or create a session file.
Alert at 80%. Act at 85%. At 100% you may not be able to log in to fix it.
Emergency triage:
sudo journalctl --vacuum-size=200M
pm2 flush
docker system prune -af # ⚠️ never add --volumes
sudo apt clean
sudo find /var/log -name "*.gz" -mtime +7 -deleteCommon space consumers, in rough order:
| Path | Usual cause |
|---|---|
/var/lib/docker | Images, build cache, container logs |
/var/log | Unrotated logs |
~/.pm2/logs | PM2 logs without logrotate |
/var/lib/postgresql | Data plus WAL |
~/apps/*/releases | Old releases not pruned |
~/.cache, ~/.pnpm-store | Package manager caches |
/boot | Old kernels — sudo apt autoremove |
Network
sudo ss -tulpn # listening sockets — the security audit
ss -s # summary
ss -tan state established | wc -l # open connections
sudo iftop # live bandwidth by connection
vnstat -d # daily traffic totalsss -tulpn IS THE MOST IMPORTANT SECURITY COMMAND
Run it after every deployment and every package install. Anything bound to 0.0.0.0 or [::] other than 22, 80, and 443 needs immediate investigation (Level 5).
Process monitoring
pm2 monit # live dashboard
pm2 list
pm2 describe api
systemctl status nginx postgresql redis-server docker
systemctl --failed # anything that failed to start# SERVER — quick health of everything
systemctl is-active nginx postgresql redis-server docker pm2-deployA monitoring script
#!/usr/bin/env bash
# /home/deploy/scripts/health-check.sh
set -uo pipefail
ALERT_EMAIL="you@example.com"
DISK_THRESHOLD=85
MEM_THRESHOLD=90
ISSUES=()
# --- Disk ---
DISK=$(df / --output=pcent | tail -1 | tr -dc '0-9')
[ "$DISK" -gt "$DISK_THRESHOLD" ] && ISSUES+=("Disk usage ${DISK}%")
# --- Memory ---
MEM=$(free | awk '/^Mem:/ {printf "%.0f", (($2-$7)/$2)*100}')
[ "$MEM" -gt "$MEM_THRESHOLD" ] && ISSUES+=("Memory usage ${MEM}%")
# --- Services ---
for svc in nginx postgresql redis-server; do
systemctl is-active --quiet "$svc" || ISSUES+=("Service $svc is DOWN")
done
# --- PM2 apps ---
export NVM_DIR="$HOME/.nvm"; [ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"
if ! pm2 jlist 2>/dev/null | jq -e 'all(.[]; .pm2_env.status == "online")' > /dev/null; then
ISSUES+=("One or more PM2 processes are not online")
fi
# --- Endpoints ---
curl -fsS --max-time 10 http://127.0.0.1:3001/api/health > /dev/null || ISSUES+=("API health check failed")
curl -fsS --max-time 10 http://127.0.0.1:3000/ > /dev/null || ISSUES+=("Web health check failed")
# --- TLS expiry ---
DAYS=$(( ( $(date -d "$(echo | openssl s_client -connect app.example.com:443 -servername app.example.com 2>/dev/null \
| openssl x509 -noout -enddate | cut -d= -f2)" +%s) - $(date +%s) ) / 86400 ))
[ "$DAYS" -lt 20 ] && ISSUES+=("TLS certificate expires in ${DAYS} days")
# --- Recent OOM kills ---
sudo dmesg -T 2>/dev/null | grep -qi "killed process" && ISSUES+=("OOM killer activity detected")
# --- Report ---
if [ ${#ISSUES[@]} -gt 0 ]; then
printf 'ALERT on %s\n\n' "$(hostname)"
printf ' - %s\n' "${ISSUES[@]}"
printf ' - %s\n' "${ISSUES[@]}" | mail -s "[ALERT] $(hostname)" "$ALERT_EMAIL" 2>/dev/null || true
exit 1
fi
echo "OK — disk ${DISK}%, mem ${MEM}%, cert ${DAYS}d"# SERVER
chmod +x /home/deploy/scripts/health-check.sh
crontab -e*/10 * * * * /home/deploy/scripts/health-check.sh >> /home/deploy/logs/health.log 2>&1A MONITORING SCRIPT ON THE MONITORED SERVER CANNOT DETECT ITS OWN DEATH
If the server is down, off the network, or out of disk, the cron job does not run and you get no alert. Silence looks identical to health.
You need at least one external check. Free options: UptimeRobot, Better Stack, Hetzner's own monitoring, or a cron job on a different machine.
External monitoring
| Service | Free tier | Checks |
|---|---|---|
| UptimeRobot | 50 monitors, 5-min interval | HTTP, port, keyword, ping |
| Better Stack | 10 monitors, 3-min | HTTP + incident management + on-call |
| Healthchecks.io | 20 checks | Dead-man's switch for cron jobs |
| Hetzner Cloud | Included | Basic resource metrics |
DEAD-MAN'S SWITCHES FOR CRON JOBS
The failure mode of a backup cron job is that it silently stops running — a changed path, a full disk, a permissions change. You discover it when you need a restore.
Healthchecks.io inverts this: your job pings a URL on success, and the service alerts you when the ping does not arrive.
# At the end of your backup script
curl -fsS -m 10 --retry 3 https://hc-ping.com/your-uuid > /dev/nullApply it to: database backups, certificate renewal, log rotation, cleanup jobs. (Level 22)
What to monitor externally
| Check | Interval | Alert when |
|---|---|---|
https://app.example.com returns 200 | 1–5 min | 2 consecutive failures |
https://api.example.com/api/health returns 200 | 1–5 min | 2 consecutive failures |
| TLS certificate expiry | Daily | < 20 days |
| Response time | 5 min | > 3 s sustained |
| Backup dead-man's switch | Daily | No ping in 26 hours |
| Domain expiry | Weekly | < 30 days |
ALERT ON TWO CONSECUTIVE FAILURES, NOT ONE
A single failed check catches every transient network blip and trains you to ignore alerts. Two consecutive failures at 1-minute intervals still means you know within two minutes, with far fewer false positives.
Error tracking
pnpm add @sentry/node @sentry/nuxt// backend/src/main.ts — before everything else
import * as Sentry from '@sentry/node';
Sentry.init({
dsn: process.env.SENTRY_DSN,
environment: process.env.NODE_ENV,
release: process.env.GIT_COMMIT,
tracesSampleRate: 0.1,
beforeSend(event) {
// Never ship secrets to a third party
delete event.request?.headers?.authorization;
delete event.request?.headers?.cookie;
if (event.request?.data && typeof event.request.data === 'object') {
delete (event.request.data as Record<string, unknown>).password;
}
return event;
},
});ERROR TRACKING BEATS LOG GREPPING
Logs tell you an error happened. Sentry tells you: it happened 847 times, started 20 minutes ago, affects 3% of users, correlates with release a1b2c3d, and here is the exact line with the variable values.
release: process.env.GIT_COMMIT is the field that matters most — it directly ties a spike in errors to the deploy that caused it.
REVIEW WHAT YOUR ERROR TRACKER SENDS
By default, error reporters capture request headers, bodies, and local variables — which can include passwords, tokens, and personal data, shipped to a third party and retained for months. The beforeSend scrubbing above is the minimum. Check what actually arrives by triggering a test error and reading the event.
Metrics dashboards
For a single VPS, full observability stacks are usually more than you need.
| Option | Setup | Cost | Good for |
|---|---|---|---|
| Netdata | One command | Free | Best value — instant per-second dashboards |
| Prometheus + Grafana | Substantial | Free (self-hosted) | Custom metrics, long retention |
| Grafana Cloud | Agent install | Free tier | Managed, no maintenance |
| Provider metrics | None | Included | Basic CPU/disk/network |
# SERVER — Netdata
wget -O /tmp/netdata-kickstart.sh https://get.netdata.cloud/kickstart.sh
sh /tmp/netdata-kickstart.sh --stable-channel --disable-telemetryNETDATA BINDS 0.0.0.0:19999 BY DEFAULT
That dashboard exposes detailed information about your server — running processes, open ports, resource usage — to anyone on the internet. It is a reconnaissance goldmine.
# SERVER — /etc/netdata/netdata.conf
sudo tee -a /etc/netdata/netdata.conf <<'EOF'
[web]
bind to = 127.0.0.1
EOF
sudo systemctl restart netdataAccess it via an SSH tunnel:
# LOCAL
ssh -L 19999:127.0.0.1:19999 deploy@203.0.113.10
# open http://localhost:19999Verify: sudo ss -tulpn | grep 19999 must show 127.0.0.1.
The four golden signals
If you monitor only four things:
| Signal | Measure | Alert when |
|---|---|---|
| Latency | p95 response time | p95 > 1 s for 5 minutes |
| Traffic | Requests per second | Sudden 10× spike or drop to zero |
| Errors | 5xx rate | > 1% of requests |
| Saturation | CPU, memory, disk | Disk > 85%, memory > 90%, load > cores × 2 |
A DROP TO ZERO TRAFFIC IS AN OUTAGE
Alerts are usually configured for "too much". A sudden drop to zero requests means DNS broke, the certificate expired, or Nginx is down — and nobody can reach you. Alert on both directions.
Investigating an incident
# SERVER — the 60-second triage
uptime && free -h && df -h / # resources
systemctl --failed # anything down?
pm2 list # app status + restart counts
sudo tail -50 /var/log/nginx/error.log # proxy-level errors
pm2 logs --err --lines 50 --nostream # app errors
sudo dmesg -T | tail -20 # OOM? disk errors?
sudo journalctl -p err --since "30 min ago" | tail -50
curl -i http://127.0.0.1:3001/api/health # is the app itself alive?
sudo -u postgres psql -c "SELECT count(*) FROM pg_stat_activity;"THE ORDER MATTERS
Work outward from the resource layer: is the machine healthy → are services running → is the app responding → is the database responding.
The alternative — starting with application logs — often means reading a flood of connection errors that are symptoms of a full disk you would have spotted in the first command.
Production Checklist — Level 21
- [ ] journald storage is persistent with a size cap
- [ ]
pm2-logrotateinstalled and configured - [ ] Docker log limits set in
daemon.json - [ ] Logrotate configured for any custom log files
- [ ] Application uses structured JSON logging with request IDs
- [ ] Secrets and personal data redacted from logs
- [ ] Log level is
infoin production, and changeable via an env var - [ ] Nginx
log_formatincludes$request_timeand$upstream_response_time - [ ] Per-site Nginx access and error logs
- [ ] Disk usage monitored, alerting at 85%
- [ ] Memory monitored via
available, notfree - [ ] I check
dmesgfor OOM kills when a process disappears - [ ] PostgreSQL protected with
OOMScoreAdjust=-500 - [ ] Local health-check script running on cron
- [ ] External uptime monitoring configured — the local script cannot detect its own death
- [ ] Alerts fire on two consecutive failures, not one
- [ ] Dead-man's switch on backup and renewal cron jobs
- [ ] TLS expiry monitored independently of Certbot
- [ ] Error tracking (Sentry) with
releaseset to the commit SHA - [ ] Error tracker configured to scrub auth headers and passwords
- [ ] Netdata (if installed) bound to
127.0.0.1 - [ ] Alerts on both traffic spikes and traffic dropping to zero
- [ ]
sudo ss -tulpnreviewed after every deployment
Next: Level 22 — Backups →