Level 10 — Redis
An in-memory data store used for caching, sessions, queues, rate limiting, and — in this stack — as the Socket.IO adapter that lets multiple app instances share WebSocket events.
What Redis is
Redis keeps its entire dataset in RAM, which makes operations sub-millisecond. It is single-threaded for command execution (so commands are atomic with no locking), and optionally persists to disk so data survives a restart.
| Property | Implication |
|---|---|
| In-memory | Extremely fast; dataset must fit in RAM |
| Single-threaded | Commands are atomic. One slow command blocks every client. |
| Rich data types | Strings, hashes, lists, sets, sorted sets, streams, bitmaps, HyperLogLog |
| Optional persistence | Can be a pure cache or a durable store — you choose |
| Pub/Sub | Message broadcasting between processes |
| TTL on any key | Automatic expiry, which is what makes caching trivial |
KEYS * WILL FREEZE YOUR PRODUCTION SITE
Because Redis is single-threaded, KEYS * scans every key while every other client waits. On a few million keys that is seconds of total outage.
Use SCAN instead — it iterates in small batches:
redis-cli --scan --pattern 'session:*' | head -20Same applies to FLUSHALL on a large dataset, and to SMEMBERS on a huge set. Redis 6.2+ lets you rename or disable dangerous commands (shown below).
Why applications use Redis
Caching
The most common use. Store the result of an expensive query; serve subsequent requests from memory.
async function getProducts() {
const cached = await redis.get('products:all');
if (cached) return JSON.parse(cached);
const products = await prisma.product.findMany();
await redis.setex('products:all', 60, JSON.stringify(products)); // 60s TTL
return products;
}A 50 ms database query becomes a 0.5 ms cache hit. At 100 requests/second that is the difference between a struggling database and an idle one.
CACHE INVALIDATION IS THE HARD PART
Two workable strategies:
- TTL-only — data can be up to N seconds stale, and you accept that. Simple, robust, and correct for most read-heavy data. Start here.
- Explicit invalidation — delete the key when the underlying data changes. Accurate but easy to get wrong: every write path must remember to invalidate, and missing one produces stale data that persists indefinitely.
Combine them: explicit invalidation plus a TTL as a safety net, so a missed invalidation self-corrects.
Sessions
// NestJS with express-session
import RedisStore from 'connect-redis';
app.use(session({
store: new RedisStore({ client: redisClient, prefix: 'sess:' }),
secret: process.env.SESSION_SECRET,
resave: false,
saveUninitialized: false,
cookie: { secure: true, httpOnly: true, sameSite: 'lax', maxAge: 86400000 },
}));Sessions in Redis rather than process memory means: they survive a deploy, and all PM2 cluster instances see the same session. In-memory sessions break the moment you run more than one process.
Queues
Background jobs — sending email, generating PDFs, processing uploads — should not block an HTTP request. BullMQ uses Redis as the queue backend:
import { Queue, Worker } from 'bullmq';
const connection = { host: '127.0.0.1', port: 6379, password: process.env.REDIS_PASSWORD };
const emailQueue = new Queue('emails', { connection });
await emailQueue.add('welcome', { userId }, {
attempts: 3,
backoff: { type: 'exponential', delay: 2000 },
});
new Worker('emails', async (job) => { await sendEmail(job.data); }, { connection });BULLMQ REQUIRES maxRetriesPerRequest: null
BullMQ's blocking commands conflict with ioredis's default retry behaviour. Without this setting you get intermittent MaxRetriesPerRequestError under load. It is documented but easy to miss.
Rate limiting
const key = `ratelimit:${ip}`;
const count = await redis.incr(key);
if (count === 1) await redis.expire(key, 60);
if (count > 100) throw new HttpException('Too many requests', 429);Atomic and shared across all instances — unlike in-memory rate limiting, which each PM2 worker would track separately, effectively multiplying your limit by the instance count.
The Socket.IO adapter — essential for this stack
WITHOUT THE REDIS ADAPTER, WEBSOCKETS BREAK IN CLUSTER MODE
With PM2 running 4 instances, a client connects to instance 2. When instance 1 handles an HTTP request and calls io.emit('notification', ...), only clients connected to instance 1 receive it. Your user on instance 2 sees nothing.
Symptoms: messages arrive "sometimes", or work perfectly in development (one process) and fail intermittently in production. This confuses people for days.
The Redis adapter publishes every emit over Redis pub/sub so all instances broadcast to their own clients.
// main.ts
import { createAdapter } from '@socket.io/redis-adapter';
import { createClient } from 'redis';
const pubClient = createClient({ url: process.env.REDIS_URL });
const subClient = pubClient.duplicate();
await Promise.all([pubClient.connect(), subClient.connect()]);
io.adapter(createAdapter(pubClient, subClient));Two clients are required: a Redis connection in subscribe mode cannot issue other commands.
You also need sticky sessions at the Nginx layer, because Socket.IO's HTTP long-polling handshake sends multiple requests that must reach the same instance (Level 14).
Option A — native Redis
# SERVER
sudo apt update
sudo apt install -y redis-server
sudo systemctl status redis-server
redis-cli ping # PONGUbuntu 24.04 ships Redis 7.x. The Debian package defaults to binding 127.0.0.1 and enabling supervised systemd — a safe starting point.
Option B — Redis in Docker
services:
redis:
image: redis:7-alpine
container_name: myapp-redis
restart: unless-stopped
command: >
redis-server
--requirepass ${REDIS_PASSWORD}
--appendonly yes
--appendfsync everysec
--maxmemory 512mb
--maxmemory-policy allkeys-lru
volumes:
- redis_data:/data
ports:
- "127.0.0.1:6379:6379" # ⚠️ localhost only
healthcheck:
test: ["CMD", "redis-cli", "-a", "${REDIS_PASSWORD}", "ping"]
interval: 10s
timeout: 3s
retries: 5
volumes:
redis_data:"6379:6379" IS A CATASTROPHIC TYPO
Docker publishes to 0.0.0.0 and bypasses UFW (Level 5). An exposed Redis is compromised in minutes — see the attack below. Always prefix 127.0.0.1:, or omit ports: entirely if only other containers need it.
Security — Redis needs deliberate configuration
THE EXPOSED-REDIS ATTACK, STEP BY STEP
Redis has no authentication by default and executes any command from anyone who can reach the port. The fully automated attack:
CONFIG SET dir /home/deploy/.ssh
CONFIG SET dbfilename authorized_keys
SET payload "\n\nssh-rsa AAAAB3Nza... attacker@evil\n\n"
SAVERedis has now written the attacker's public key into your authorized_keys. They SSH in as deploy. From there: sudo, crypto-miner, data exfiltration.
Variants write cron jobs to /var/spool/cron/crontabs/root or use the Lua sandbox escape. Shodan continuously indexes tens of thousands of open Redis instances. This is not theoretical — it is the single most common way small deployments get owned.
Three defences, all of them:
bind 127.0.0.1— not reachable from the networkrequirepass <strong>— even local access needs credentials- Never open 6379 in any firewall
The configuration file
/etc/redis/redis.conf# SERVER
sudo cp /etc/redis/redis.conf /etc/redis/redis.conf.bak
sudo nano /etc/redis/redis.conf# ---- Network ----
bind 127.0.0.1 -::1
port 6379
protected-mode yes
tcp-keepalive 300
timeout 0
# ---- Security ----
requirepass CHANGE-ME-openssl-rand-base64-32
# Disable or rename dangerous commands
rename-command CONFIG ""
rename-command FLUSHALL ""
rename-command FLUSHDB ""
rename-command KEYS ""
rename-command SHUTDOWN "SHUTDOWN_a8f3k2"
# ---- Memory ----
maxmemory 512mb
maxmemory-policy allkeys-lru
# ---- Persistence ----
appendonly yes
appendfsync everysec
save 900 1
save 300 10
save 60 10000
# ---- Logging ----
loglevel notice
logfile /var/log/redis/redis-server.log
# ---- Misc ----
supervised systemd
dir /var/lib/redis| Setting | Why |
|---|---|
bind 127.0.0.1 -::1 | Loopback only. The most important line. The - prefix means "do not fail if this address is unavailable". |
protected-mode yes | Refuses external connections when no password is set — a safety net, not a substitute for the above |
requirepass | Requires AUTH before any command. Generate with openssl rand -base64 32. |
rename-command CONFIG "" | Disables CONFIG, which is what the attack above uses to redirect the dump file. Note: this breaks redis-cli CONFIG GET for your own diagnostics too. |
rename-command KEYS "" | Prevents an accidental production freeze |
maxmemory | Hard cap. Without it, Redis grows until the OOM killer intervenes. |
maxmemory-policy | What to do when full — see below |
appendonly yes | AOF persistence — see below |
requirepass MUST BE LONG
Redis can process a very high volume of AUTH attempts per second — it is not designed to resist brute force. A short password is not meaningfully protective. Use 32+ random bytes.
ACLs ARE BETTER THAN requirepass (REDIS 6+)
requirepass gives full access. ACLs let you create restricted users:
user default off
user myapp on >StrongPasswordHere ~myapp:* +@read +@write +@keyspace -@dangerousThis user can only touch keys matching myapp:* and cannot run administrative commands. Multiple applications sharing one Redis should definitely use ACLs.
Apply:
# SERVER
sudo chmod 640 /etc/redis/redis.conf # it contains the password
sudo chown redis:redis /etc/redis/redis.conf
sudo systemctl restart redis-server
redis-cli -a "$REDIS_PASSWORD" ping-a LEAKS THE PASSWORD TO ps AND SHELL HISTORY
redis-cli -a secret puts the password in the process list, visible to any user via ps aux, and in ~/.bash_history.
Safer:
redis-cli
127.0.0.1:6379> AUTH your-password
# or
REDISCLI_AUTH="$REDIS_PASSWORD" redis-cli pingThe REDISCLI_AUTH environment variable is read automatically and does not appear in ps.
The connection URL
redis://:password@127.0.0.1:6379/0
↑
note the leading colon — Redis has no username with requirepassWith ACLs: redis://myapp:password@127.0.0.1:6379/0
The trailing /0 selects the database number (0–15). Use different numbers to separate concerns:
redis://:pass@127.0.0.1:6379/0 # cache
redis://:pass@127.0.0.1:6379/1 # sessions
redis://:pass@127.0.0.1:6379/2 # BullMQREDIS "DATABASES" ARE NOT ISOLATION
They share memory, the same password, and the same FLUSHALL. They are a namespace convenience only, and Redis Cluster does not support them at all. Prefer key prefixes (cache:, sess:, bull:) which work everywhere.
Persistence
Two mechanisms, usable together:
| RDB (snapshot) | AOF (append-only file) | |
|---|---|---|
| How | Periodic point-in-time dump | Every write command appended to a log |
| Worst-case data loss | Minutes | 1 second (with everysec) |
| Restart speed | Fast | Slower (replays the log) |
| File size | Compact | Larger, needs periodic rewrite |
| Fork cost | Forks the process — brief memory spike | Background rewrite also forks |
# RDB — save if ≥1 key changed in 900s, ≥10 in 300s, ≥10000 in 60s
save 900 1
save 300 10
save 60 10000
# AOF
appendonly yes
appendfsync everysec # fsync once per second — the right balanceappendfsync options: always (safe, slow), everysec (recommended), no (fastest, up to 30s loss).
DOES YOUR REDIS DATA NEED TO SURVIVE A RESTART?
| Use | Persistence needed? |
|---|---|
| Cache only | No. Losing it means a cold start, nothing more. Disable both for speed. |
| Sessions | Yes — otherwise every user is logged out on restart |
| BullMQ queues | Yes — otherwise queued jobs are lost |
| Socket.IO adapter | No — it is pure pub/sub, nothing is stored |
| Rate limit counters | No — losing them briefly resets limits, which is acceptable |
Since sessions and queues usually share the instance, enable AOF in this stack. Persistence is not free — the fork during rewrite briefly doubles memory usage — but the alternative is losing user sessions on every restart.
Manual snapshot:
redis-cli BGSAVE # background — non-blocking. Use this.
redis-cli LASTSAVE # timestamp of the last successful save
redis-cli BGREWRITEAOF # compact the AOFNEVER RUN SAVE ON PRODUCTION
SAVE (without BG) blocks the entire server until the dump completes. BGSAVE forks a child process instead.
Memory management
maxmemory 512mb
maxmemory-policy allkeys-lru| Policy | Behaviour | Use when |
|---|---|---|
noeviction | Return errors on write when full | Redis is a durable store (queues) — you want to notice, not lose data |
allkeys-lru | Evict least-recently-used key | Pure cache |
allkeys-lfu | Evict least-frequently-used | Cache with a stable hot set |
volatile-lru | Evict LRU among keys with a TTL | Mixed cache and durable data in one instance |
volatile-ttl | Evict the key expiring soonest | Mixed workload |
THE MIXED-WORKLOAD TRAP
If one Redis instance holds both cache entries and BullMQ jobs, allkeys-lru will happily evict your queued jobs when memory fills. Jobs silently vanish.
Options, best first:
- Separate instances — one cache (
allkeys-lru, no persistence), one durable (noeviction, AOF on). Different ports or different containers. volatile-lruand make sure only cache keys have a TTL and durable keys have none.
Do not run allkeys-lru on an instance holding anything you cannot afford to lose.
redis-cli INFO memory
redis-cli --bigkeys # find the largest keys
redis-cli MEMORY USAGE mykey
redis-cli INFO stats | grep evicted # is eviction happening?evicted_keys climbing means maxmemory is too low for your working set.
redis-cli
redis-cli # interactive
redis-cli -n 1 # database 1
redis-cli INFO # everything
redis-cli INFO server | head -20
redis-cli DBSIZE # number of keys
redis-cli --stat # live stats, refreshed each second
redis-cli --latency # measure latency
redis-cli MONITOR # ⚠️ see warningMONITOR DEGRADES PERFORMANCE SIGNIFICANTLY
It streams every command executed by every client. On a busy server this costs a large fraction of throughput and floods your terminal. Use briefly, then Ctrl+C. Never leave it running.
For ongoing analysis use SLOWLOG instead:
redis-cli SLOWLOG GET 10
redis-cli CONFIG SET slowlog-log-slower-than 10000 # log commands over 10msCommon commands:
SET key value
SET key value EX 60 # with a 60s TTL
GET key
DEL key
EXISTS key
TTL key # -1 = no expiry, -2 = key does not exist
EXPIRE key 60
INCR counter
HSET user:1 name "Aziz" email "a@b.c"
HGETALL user:1
LPUSH queue "job1"
RPOP queue
SADD tags "node" "redis"
SMEMBERS tags
ZADD leaderboard 100 "user1"
ZREVRANGE leaderboard 0 9 WITHSCORES
--scan --pattern 'sess:*' # safe key iterationNode.js client configuration
// ioredis — the usual choice for NestJS/BullMQ
import Redis from 'ioredis';
export const redis = new Redis(process.env.REDIS_URL!, {
maxRetriesPerRequest: null, // required by BullMQ
enableReadyCheck: true,
retryStrategy: (times) => Math.min(times * 100, 3000),
reconnectOnError: (err) => err.message.includes('READONLY'),
lazyConnect: false,
});
redis.on('error', (err) => console.error('[redis] error', err.message));
redis.on('connect', () => console.log('[redis] connected'));
redis.on('reconnecting', (d: number) => console.warn(`[redis] reconnecting in ${d}ms`));AN UNHANDLED error EVENT CRASHES YOUR NODE PROCESS
Node's EventEmitter throws if an error event has no listener. A momentary Redis blip then takes down your entire API. Always attach an error handler.
DEGRADE GRACEFULLY WHEN THE CACHE IS DOWN
async function getCached<T>(key: string, fetcher: () => Promise<T>, ttl = 60): Promise<T> {
try {
const hit = await redis.get(key);
if (hit) return JSON.parse(hit);
} catch (err) {
console.warn('[cache] read failed, falling through to DB', err);
}
const fresh = await fetcher();
try {
await redis.setex(key, ttl, JSON.stringify(fresh));
} catch { /* cache write failure must never fail the request */ }
return fresh;
}A cache outage should mean a slower site, not a broken one. Sessions and queues are different — those failures are real and should surface.
Service management
# Native
sudo systemctl start|stop|restart|status redis-server
sudo systemctl enable redis-server
sudo tail -f /var/log/redis/redis-server.log
sudo journalctl -u redis-server -f
# Docker
docker compose up -d redis
docker compose restart redis
docker compose logs -f redis
docker compose exec redis redis-cli -a "$REDIS_PASSWORD"Verify security
# SERVER — must show 127.0.0.1, never 0.0.0.0
sudo ss -tulpn | grep 6379
# 6379 must NOT appear
sudo ufw status | grep 6379
# Docker exposure check
docker ps --format "table {{.Names}}\t{{.Ports}}" | grep redis
# Is a password actually required?
redis-cli ping
# Expected: (error) NOAUTH Authentication required.# LOCAL — from outside; must time out or refuse
nc -vz 203.0.113.10 6379
redis-cli -h 203.0.113.10 pingIF redis-cli ping RETURNS PONG WITHOUT AUTH
Your Redis has no password. If it is also reachable from the network, assume it is already compromised. Check immediately:
cat ~/.ssh/authorized_keys # any key you do not recognise?
sudo crontab -l && crontab -l # unexpected cron jobs?
ps aux --sort=-%cpu | head # a miner burning CPU?
sudo last -20 # unexpected logins?Then follow the recovery procedure in Level 23.
Troubleshooting
| Problem | Cause | Diagnose | Fix |
|---|---|---|---|
ECONNREFUSED 127.0.0.1:6379 | Not running | systemctl status redis-server | Start it |
NOAUTH Authentication required | Password not supplied | — | Add password to REDIS_URL |
WRONGPASS | Wrong password | grep requirepass /etc/redis/redis.conf | Fix .env, pm2 reload --update-env |
OOM command not allowed | maxmemory reached with noeviction | redis-cli INFO memory | Raise maxmemory, set an eviction policy, or add TTLs |
MISCONF Redis is configured to save RDB snapshots... | Background save failing — usually disk full or bad permissions | df -h, ls -la /var/lib/redis | Free disk; chown redis:redis /var/lib/redis |
| Redis restarts, data gone | Persistence disabled | redis-cli CONFIG GET appendonly | Enable AOF |
| Keys vanishing unexpectedly | LRU eviction | redis-cli INFO stats | grep evicted | Raise maxmemory or change policy |
| Socket.IO events reach some clients only | No Redis adapter in cluster mode | pm2 list — more than one instance? | Add @socket.io/redis-adapter + sticky sessions |
| Everything intermittently slow | A KEYS or big SMEMBERS blocking | redis-cli SLOWLOG GET 10 | Replace with SCAN; disable KEYS |
MaxRetriesPerRequestError with BullMQ | Missing maxRetriesPerRequest: null | — | Set it |
| Node process crashes on Redis blip | No error listener | — | redis.on('error', ...) |
Production Checklist — Level 10
- [ ]
bind 127.0.0.1 -::1inredis.conf - [ ]
requirepassset to 32+ random bytes - [ ]
protected-mode yes - [ ]
sudo ss -tulpn | grep 6379shows127.0.0.1only - [ ] Port 6379 not in
ufw status - [ ] Docker
ports:is127.0.0.1:6379:6379or absent - [ ]
redis-cli pingwithout auth returnsNOAUTH - [ ] Verified from outside that 6379 is unreachable
- [ ]
CONFIG,FLUSHALL,FLUSHDB,KEYSrenamed or disabled - [ ]
maxmemoryset with an eviction policy appropriate to the data - [ ] Queues/sessions are not on an
allkeys-lruinstance - [ ] AOF enabled with
appendfsync everysec(sessions/queues present) - [ ]
redis.confis mode640, owned byredis - [ ] Node client has an
errorevent handler - [ ]
maxRetriesPerRequest: nullif using BullMQ - [ ] Socket.IO Redis adapter configured (required with PM2 cluster mode)
- [ ] Cache reads degrade gracefully when Redis is unavailable
- [ ]
REDIS_PASSWORDin.envat mode 600, backed up in a password manager - [ ] Memory usage monitored; alert before
maxmemoryis reached