Skip to content

Level 21 — Logging and Monitoring ​

The difference between "working" and "production-ready" is largely this chapter. Without monitoring, you find out about outages from users.

Where every log lives ​

SourceLocationView with
systemd servicesjournald (binary)journalctl -u <service>
Nginx access/var/log/nginx/access.logtail -f
Nginx error/var/log/nginx/error.logtail -f
PM2 apps/home/deploy/logs/*.log, ~/.pm2/logs/pm2 logs
PM2 daemon~/.pm2/pm2.logtail
PostgreSQL/var/log/postgresql/postgresql-16-main.logtail -f
Redis/var/log/redis/redis-server.logtail -f
Docker containersDocker's JSON log driverdocker logs
Authentication/var/log/auth.log + journaldjournalctl -u ssh
Kernel / OOMdmesg, journaldjournalctl -k
fail2ban/var/log/fail2ban.logtail -f
Unattended upgrades/var/log/unattended-upgrades/tail
UFW blocks/var/log/syslog (UFW BLOCK)grep

journalctl ​

systemd's journal is the central log for every service it manages.

bash
# SERVER
journalctl -u nginx                      # one service
journalctl -u nginx -f                   # follow
journalctl -u nginx -n 100               # last 100 lines
journalctl -u nginx --since "1 hour ago"
journalctl -u nginx --since "2026-08-11 09:00" --until "2026-08-11 10:00"
journalctl -u nginx -p err               # errors and worse
journalctl -u nginx --no-pager           # no interactive pager (scripts)
journalctl -u nginx -o json-pretty       # structured output

journalctl -k                            # kernel messages
journalctl -b                            # this boot
journalctl -b -1                         # previous boot — what happened before the crash?
journalctl --list-boots

journalctl -xe                           # recent entries with explanations
journalctl --disk-usage
journalctl --vacuum-size=500M            # trim the journal
journalctl --vacuum-time=14d
Priority-p value
emerg0
alert1
crit2
err3
warning4
notice5
info6
debug7

journalctl -b -1 -p err AFTER AN UNEXPLAINED REBOOT

Shows the errors from the boot before this one. If the server restarted on its own — OOM, kernel panic, provider maintenance, unattended-upgrades reboot — the evidence is there.

THE JOURNAL IS VOLATILE BY DEFAULT ON SOME IMAGES

If /var/log/journal/ does not exist, logs live in /run and are lost on reboot — exactly when you need them.

bash
sudo mkdir -p /var/log/journal
sudo systemd-tmpfiles --create --prefix /var/log/journal
sudo systemctl restart systemd-journald
journalctl --disk-usage

Then cap it in /etc/systemd/journald.conf:

ini
[Journal]
Storage=persistent
SystemMaxUse=1G
MaxRetentionSec=1month

Application logs ​

bash
# SERVER
pm2 logs                               # all apps, live
pm2 logs api --lines 200
pm2 logs api --err                     # errors only
pm2 logs api --nostream --lines 100    # print and exit — for scripts
pm2 flush                              # truncate all
pm2 logs api --raw                     # no PM2 prefix — needed for jq

Structured logging ​

JSON LOGS TURN GREP INTO QUERIES

ts
// NestJS with nestjs-pino
import { LoggerModule } from 'nestjs-pino';

LoggerModule.forRoot({
  pinoHttp: {
    level: process.env.LOG_LEVEL ?? 'info',
    genReqId: (req) => req.headers['x-request-id'] ?? randomUUID(),
    redact: {
      paths: ['req.headers.authorization', 'req.headers.cookie', 'req.body.password'],
      remove: true,
    },
    serializers: {
      req: (req) => ({ method: req.method, url: req.url, id: req.id }),
    },
  },
});

Now:

bash
pm2 logs api --raw --nostream | jq 'select(.level >= 50)'                  # errors
pm2 logs api --raw --nostream | jq -r 'select(.reqId=="abc-123")'          # one request
pm2 logs api --raw --nostream | jq -r '.msg' | sort | uniq -c | sort -rn   # most common

The request ID is the highest-value field: it lets you reconstruct everything that happened during one request, across every log line it produced.

NEVER LOG SECRETS OR PERSONAL DATA

Passwords, tokens, full credit card numbers, and API keys must never reach a log file. Logs are backed up, shipped to third-party services, and read by more people than you think.

The redact config above is not optional. Also check that your error reporter (Sentry) scrubs Authorization headers, cookies, and request bodies.

Log levels ​

LevelUse forProduction
errorSomething broke and needs attention✅
warnRecoverable problem, degraded behaviour✅
infoSignificant events (startup, deploy, job completed)✅
debugDetailed flow❌ Off
traceEverything❌ Off

debug IN PRODUCTION FILLS YOUR DISK AND SLOWS THE APP

Every log line is a synchronous-ish write and a disk operation. Debug-level logging on a busy API can produce gigabytes per day and measurably increase latency. Keep it at info, and make the level an environment variable so you can raise it temporarily during an incident.

Nginx logs ​

bash
# SERVER
sudo tail -f /var/log/nginx/access.log
sudo tail -f /var/log/nginx/error.log
sudo tail -f /var/log/nginx/api.example.com.error.log

With the log_format from Level 14, each line ends with rt= and urt=:

bash
# Slowest requests
sudo awk '{print $(NF-1), $7}' /var/log/nginx/access.log | sort -rn | head -20

# Status code distribution
sudo awk '{print $9}' /var/log/nginx/access.log | sort | uniq -c | sort -rn

# 5xx in the last hour
sudo grep " $(date -u '+%d/%b/%Y:%H')" /var/log/nginx/access.log | awk '$9 ~ /^5/' | tail -30

# Top clients
sudo awk '{print $1}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -20

# Requests per minute — traffic spike detection
sudo awk '{print substr($4,2,17)}' /var/log/nginx/access.log | uniq -c | tail -30

rt vs urt TELLS YOU WHERE THE TIME WENT

  • rt ($request_time) — total, from first byte received to last byte sent
  • urt ($upstream_response_time) — how long your app took

rt=5.2 urt=0.05 → your app was fast; the client's connection was slow (mobile, large upload). Not your problem.

rt=5.2 urt=5.1 → your app was slow. This is your problem. Go look at the database.

Docker logs ​

bash
docker logs -f --tail 100 myapp-postgres
docker compose logs -f api
docker compose logs --since 30m

DOCKER LOGS ARE UNBOUNDED BY DEFAULT

A chatty container fills /var/lib/docker/containers/ until the disk is full, at which point PostgreSQL cannot write, Nginx cannot log, and everything fails at once (Level 11).

json
// /etc/docker/daemon.json
{ "log-driver": "json-file", "log-opts": { "max-size": "10m", "max-file": "3" } }

Applies to newly created containers only. Recreate existing ones.

Log rotation ​

Without rotation, disk fills. Check every log producer:

bash
# SERVER
sudo du -sh /var/log/* | sort -rh | head -20
du -sh ~/.pm2/logs /home/deploy/logs
docker system df
journalctl --disk-usage
ProducerRotation
Nginx/etc/logrotate.d/nginx (package default: daily, 14 days)
PostgreSQL/etc/logrotate.d/postgresql-common
systemd journalSystemMaxUse in journald.conf
Dockerlog-opts in daemon.json
PM2pm2 install pm2-logrotate — not configured by default
bash
# SERVER
pm2 install pm2-logrotate
pm2 set pm2-logrotate:max_size 10M
pm2 set pm2-logrotate:retain 14
pm2 set pm2-logrotate:compress true
pm2 set pm2-logrotate:rotateInterval '0 0 * * *'

Custom rotation for your own log files:

bash
# SERVER
sudo nano /etc/logrotate.d/myapp
/home/deploy/logs/*.log {
    daily
    rotate 14
    compress
    delaycompress
    missingok
    notifempty
    copytruncate
    su deploy deploy
}

copytruncate vs create

create renames the file and makes a new one — but a running process holding the old file descriptor keeps writing to the renamed file, which then gets deleted. Logs vanish.

copytruncate copies the content and truncates the original in place, so the process's file descriptor stays valid. Use it for any process that does not reopen its log file on SIGHUP. There is a tiny race where lines written during the copy can be lost; that is an acceptable trade.

bash
sudo logrotate -d /etc/logrotate.d/myapp     # dry run
sudo logrotate -f /etc/logrotate.d/myapp     # force now

Resource monitoring ​

CPU and load ​

bash
# SERVER
uptime
 14:32:01 up 12 days,  3:14,  1 user,  load average: 0.52, 0.48, 0.41

INTERPRETING LOAD AVERAGE

Load is the number of processes running or waiting (for CPU or disk), averaged over 1, 5, and 15 minutes.

Compare it to your core count: nproc.

  • 2 cores, load 1.0 → 50% utilised. Fine.
  • 2 cores, load 2.0 → fully utilised. At capacity.
  • 2 cores, load 8.0 → heavily overloaded; requests are queueing.

Read the trend, not the instant. 0.5 2.0 4.0 means load is falling (1-min lowest). 4.0 2.0 0.5 means it is rising fast — investigate now.

High load with low CPU usage means I/O wait — check disk with iostat.

bash
htop                    # sudo apt install htop
top
ps aux --sort=-%cpu | head -10
mpstat 1 5              # sudo apt install sysstat

Memory ​

bash
free -h
               total        used        free      shared  buff/cache   available
Mem:           3.8Gi       1.9Gi       234Mi        12Mi       1.7Gi       1.7Gi
Swap:          2.0Gi        12Mi       2.0Gi

READ available, NOT free

Linux uses spare RAM for disk cache, which is why free looks alarming. That cache is released instantly when a program needs memory.

available is the real number. Below ~10% of total, you are in trouble.

Watch swap too: a small amount used is fine, but active swapping (si/so columns in vmstat 1) means severe pressure and terrible latency.

bash
ps aux --sort=-%mem | head -10
vmstat 1 5
sudo smem -tk 2>/dev/null | tail -5    # more accurate per-process accounting

The OOM killer ​

WHEN THE SERVER RUNS OUT OF MEMORY, THE KERNEL KILLS SOMETHING

It picks the process with the highest "badness" score — usually the largest, which is Node or PostgreSQL. Your app disappears with no application-level error.

Detect it:

bash
sudo dmesg -T | grep -i "killed process"
sudo journalctl -k --since "1 hour ago" | grep -i "out of memory"
Out of memory: Killed process 12345 (node) total-vm:2841232kB, anon-rss:1834920kB

That is definitive. Fixes: add swap (Level 4), reduce PM2 instances, set max_memory_restart lower, move builds to CI, or resize the server.

Protect PostgreSQL specifically:

bash
sudo systemctl edit postgresql
ini
[Service]
OOMScoreAdjust=-500

This makes the kernel prefer killing something else. Your database is the hardest thing to recover.

Disk ​

bash
df -h
df -i                                    # inodes — can fill with bytes free
du -h --max-depth=1 / 2>/dev/null | sort -rh | head -20
du -sh /var/log /var/lib/docker /home/deploy 2>/dev/null
ncdu /                                   # interactive — sudo apt install ncdu

A FULL DISK BREAKS EVERYTHING AT ONCE, CONFUSINGLY

PostgreSQL stops accepting writes, Nginx cannot log (and may stop serving), PM2 cannot write logs, builds fail, apt fails, and even SSH login can fail because it cannot write to ~/.bash_history or create a session file.

Alert at 80%. Act at 85%. At 100% you may not be able to log in to fix it.

Emergency triage:

bash
sudo journalctl --vacuum-size=200M
pm2 flush
docker system prune -af          # ⚠️ never add --volumes
sudo apt clean
sudo find /var/log -name "*.gz" -mtime +7 -delete

Common space consumers, in rough order:

PathUsual cause
/var/lib/dockerImages, build cache, container logs
/var/logUnrotated logs
~/.pm2/logsPM2 logs without logrotate
/var/lib/postgresqlData plus WAL
~/apps/*/releasesOld releases not pruned
~/.cache, ~/.pnpm-storePackage manager caches
/bootOld kernels — sudo apt autoremove

Network ​

bash
sudo ss -tulpn                          # listening sockets — the security audit
ss -s                                    # summary
ss -tan state established | wc -l       # open connections
sudo iftop                               # live bandwidth by connection
vnstat -d                                # daily traffic totals

ss -tulpn IS THE MOST IMPORTANT SECURITY COMMAND

Run it after every deployment and every package install. Anything bound to 0.0.0.0 or [::] other than 22, 80, and 443 needs immediate investigation (Level 5).

Process monitoring ​

bash
pm2 monit                       # live dashboard
pm2 list
pm2 describe api
systemctl status nginx postgresql redis-server docker
systemctl --failed              # anything that failed to start
bash
# SERVER — quick health of everything
systemctl is-active nginx postgresql redis-server docker pm2-deploy

A monitoring script ​

bash
#!/usr/bin/env bash
# /home/deploy/scripts/health-check.sh
set -uo pipefail

ALERT_EMAIL="you@example.com"
DISK_THRESHOLD=85
MEM_THRESHOLD=90
ISSUES=()

# --- Disk ---
DISK=$(df / --output=pcent | tail -1 | tr -dc '0-9')
[ "$DISK" -gt "$DISK_THRESHOLD" ] && ISSUES+=("Disk usage ${DISK}%")

# --- Memory ---
MEM=$(free | awk '/^Mem:/ {printf "%.0f", (($2-$7)/$2)*100}')
[ "$MEM" -gt "$MEM_THRESHOLD" ] && ISSUES+=("Memory usage ${MEM}%")

# --- Services ---
for svc in nginx postgresql redis-server; do
  systemctl is-active --quiet "$svc" || ISSUES+=("Service $svc is DOWN")
done

# --- PM2 apps ---
export NVM_DIR="$HOME/.nvm"; [ -s "$NVM_DIR/nvm.sh" ] && . "$NVM_DIR/nvm.sh"
if ! pm2 jlist 2>/dev/null | jq -e 'all(.[]; .pm2_env.status == "online")' > /dev/null; then
  ISSUES+=("One or more PM2 processes are not online")
fi

# --- Endpoints ---
curl -fsS --max-time 10 http://127.0.0.1:3001/api/health > /dev/null || ISSUES+=("API health check failed")
curl -fsS --max-time 10 http://127.0.0.1:3000/          > /dev/null || ISSUES+=("Web health check failed")

# --- TLS expiry ---
DAYS=$(( ( $(date -d "$(echo | openssl s_client -connect app.example.com:443 -servername app.example.com 2>/dev/null \
  | openssl x509 -noout -enddate | cut -d= -f2)" +%s) - $(date +%s) ) / 86400 ))
[ "$DAYS" -lt 20 ] && ISSUES+=("TLS certificate expires in ${DAYS} days")

# --- Recent OOM kills ---
sudo dmesg -T 2>/dev/null | grep -qi "killed process" && ISSUES+=("OOM killer activity detected")

# --- Report ---
if [ ${#ISSUES[@]} -gt 0 ]; then
  printf 'ALERT on %s\n\n' "$(hostname)"
  printf ' - %s\n' "${ISSUES[@]}"
  printf ' - %s\n' "${ISSUES[@]}" | mail -s "[ALERT] $(hostname)" "$ALERT_EMAIL" 2>/dev/null || true
  exit 1
fi

echo "OK — disk ${DISK}%, mem ${MEM}%, cert ${DAYS}d"
bash
# SERVER
chmod +x /home/deploy/scripts/health-check.sh
crontab -e
cron
*/10 * * * * /home/deploy/scripts/health-check.sh >> /home/deploy/logs/health.log 2>&1

A MONITORING SCRIPT ON THE MONITORED SERVER CANNOT DETECT ITS OWN DEATH

If the server is down, off the network, or out of disk, the cron job does not run and you get no alert. Silence looks identical to health.

You need at least one external check. Free options: UptimeRobot, Better Stack, Hetzner's own monitoring, or a cron job on a different machine.

External monitoring ​

ServiceFree tierChecks
UptimeRobot50 monitors, 5-min intervalHTTP, port, keyword, ping
Better Stack10 monitors, 3-minHTTP + incident management + on-call
Healthchecks.io20 checksDead-man's switch for cron jobs
Hetzner CloudIncludedBasic resource metrics

DEAD-MAN'S SWITCHES FOR CRON JOBS

The failure mode of a backup cron job is that it silently stops running — a changed path, a full disk, a permissions change. You discover it when you need a restore.

Healthchecks.io inverts this: your job pings a URL on success, and the service alerts you when the ping does not arrive.

bash
# At the end of your backup script
curl -fsS -m 10 --retry 3 https://hc-ping.com/your-uuid > /dev/null

Apply it to: database backups, certificate renewal, log rotation, cleanup jobs. (Level 22)

What to monitor externally ​

CheckIntervalAlert when
https://app.example.com returns 2001–5 min2 consecutive failures
https://api.example.com/api/health returns 2001–5 min2 consecutive failures
TLS certificate expiryDaily< 20 days
Response time5 min> 3 s sustained
Backup dead-man's switchDailyNo ping in 26 hours
Domain expiryWeekly< 30 days

ALERT ON TWO CONSECUTIVE FAILURES, NOT ONE

A single failed check catches every transient network blip and trains you to ignore alerts. Two consecutive failures at 1-minute intervals still means you know within two minutes, with far fewer false positives.

Error tracking ​

bash
pnpm add @sentry/node @sentry/nuxt
ts
// backend/src/main.ts — before everything else
import * as Sentry from '@sentry/node';

Sentry.init({
  dsn: process.env.SENTRY_DSN,
  environment: process.env.NODE_ENV,
  release: process.env.GIT_COMMIT,
  tracesSampleRate: 0.1,
  beforeSend(event) {
    // Never ship secrets to a third party
    delete event.request?.headers?.authorization;
    delete event.request?.headers?.cookie;
    if (event.request?.data && typeof event.request.data === 'object') {
      delete (event.request.data as Record<string, unknown>).password;
    }
    return event;
  },
});

ERROR TRACKING BEATS LOG GREPPING

Logs tell you an error happened. Sentry tells you: it happened 847 times, started 20 minutes ago, affects 3% of users, correlates with release a1b2c3d, and here is the exact line with the variable values.

release: process.env.GIT_COMMIT is the field that matters most — it directly ties a spike in errors to the deploy that caused it.

REVIEW WHAT YOUR ERROR TRACKER SENDS

By default, error reporters capture request headers, bodies, and local variables — which can include passwords, tokens, and personal data, shipped to a third party and retained for months. The beforeSend scrubbing above is the minimum. Check what actually arrives by triggering a test error and reading the event.

Metrics dashboards ​

For a single VPS, full observability stacks are usually more than you need.

OptionSetupCostGood for
NetdataOne commandFreeBest value — instant per-second dashboards
Prometheus + GrafanaSubstantialFree (self-hosted)Custom metrics, long retention
Grafana CloudAgent installFree tierManaged, no maintenance
Provider metricsNoneIncludedBasic CPU/disk/network
bash
# SERVER — Netdata
wget -O /tmp/netdata-kickstart.sh https://get.netdata.cloud/kickstart.sh
sh /tmp/netdata-kickstart.sh --stable-channel --disable-telemetry

NETDATA BINDS 0.0.0.0:19999 BY DEFAULT

That dashboard exposes detailed information about your server — running processes, open ports, resource usage — to anyone on the internet. It is a reconnaissance goldmine.

bash
# SERVER — /etc/netdata/netdata.conf
sudo tee -a /etc/netdata/netdata.conf <<'EOF'
[web]
    bind to = 127.0.0.1
EOF
sudo systemctl restart netdata

Access it via an SSH tunnel:

bash
# LOCAL
ssh -L 19999:127.0.0.1:19999 deploy@203.0.113.10
# open http://localhost:19999

Verify: sudo ss -tulpn | grep 19999 must show 127.0.0.1.

The four golden signals ​

If you monitor only four things:

SignalMeasureAlert when
Latencyp95 response timep95 > 1 s for 5 minutes
TrafficRequests per secondSudden 10× spike or drop to zero
Errors5xx rate> 1% of requests
SaturationCPU, memory, diskDisk > 85%, memory > 90%, load > cores × 2

A DROP TO ZERO TRAFFIC IS AN OUTAGE

Alerts are usually configured for "too much". A sudden drop to zero requests means DNS broke, the certificate expired, or Nginx is down — and nobody can reach you. Alert on both directions.

Investigating an incident ​

bash
# SERVER — the 60-second triage
uptime && free -h && df -h /                        # resources
systemctl --failed                                   # anything down?
pm2 list                                             # app status + restart counts
sudo tail -50 /var/log/nginx/error.log               # proxy-level errors
pm2 logs --err --lines 50 --nostream                 # app errors
sudo dmesg -T | tail -20                             # OOM? disk errors?
sudo journalctl -p err --since "30 min ago" | tail -50
curl -i http://127.0.0.1:3001/api/health             # is the app itself alive?
sudo -u postgres psql -c "SELECT count(*) FROM pg_stat_activity;"

THE ORDER MATTERS

Work outward from the resource layer: is the machine healthy → are services running → is the app responding → is the database responding.

The alternative — starting with application logs — often means reading a flood of connection errors that are symptoms of a full disk you would have spotted in the first command.

Production Checklist — Level 21 ​

  • [ ] journald storage is persistent with a size cap
  • [ ] pm2-logrotate installed and configured
  • [ ] Docker log limits set in daemon.json
  • [ ] Logrotate configured for any custom log files
  • [ ] Application uses structured JSON logging with request IDs
  • [ ] Secrets and personal data redacted from logs
  • [ ] Log level is info in production, and changeable via an env var
  • [ ] Nginx log_format includes $request_time and $upstream_response_time
  • [ ] Per-site Nginx access and error logs
  • [ ] Disk usage monitored, alerting at 85%
  • [ ] Memory monitored via available, not free
  • [ ] I check dmesg for OOM kills when a process disappears
  • [ ] PostgreSQL protected with OOMScoreAdjust=-500
  • [ ] Local health-check script running on cron
  • [ ] External uptime monitoring configured — the local script cannot detect its own death
  • [ ] Alerts fire on two consecutive failures, not one
  • [ ] Dead-man's switch on backup and renewal cron jobs
  • [ ] TLS expiry monitored independently of Certbot
  • [ ] Error tracking (Sentry) with release set to the commit SHA
  • [ ] Error tracker configured to scrub auth headers and passwords
  • [ ] Netdata (if installed) bound to 127.0.0.1
  • [ ] Alerts on both traffic spikes and traffic dropping to zero
  • [ ] sudo ss -tulpn reviewed after every deployment

Next: Level 22 — Backups →