Logging, Monitoring and Troubleshooting
Troubleshooting starts with observation rather than configuration changes. Logs explain what happened, monitoring shows state and trends, and alerts should identify situations that require action.
Related topics: End-to-End Web Application Troubleshooting, systemd, cron and Schedulers, HTTP, HTTPS and TLS and Docker.
1. Rule
Do not guess. Gather facts.
Typical order:
symptom
↓
service status
↓
logs
↓
ports
↓
network
↓
configuration
↓
resources
2. systemctl
systemctl status nginx
systemctl status myapp
3. journalctl
journalctl -u myapp
journalctl -u myapp -b
journalctl -u myapp -f
journalctl -u myapp -n 100
4. Log files
/var/log/
/var/log/nginx/access.log
/var/log/nginx/error.log
5. tail
tail -f /var/log/nginx/error.log
6. grep
grep ERROR app.log
grep -i timeout app.log
grep -C 3 ERROR app.log
7. CPU and RAM
top
htop
free -h
8. Disk
df -h
du -sh *
sudo du -xhd1 / | sort -h
9. Inodes
A filesystem can have free gigabytes but no free inodes:
df -i
10. Processes
ps aux
pgrep nginx
pgrep -af myapp
11. Ports
ss -lntup
12. lsof
sudo lsof -i :443
sudo lsof /path/to/file
13. Testing localhost
curl -v http://127.0.0.1:8080/health
If localhost works, investigate proxy, firewall and DNS next.
14. Health endpoint
A backend can expose:
GET /health
returning:
{"status":"ok"}
15. Docker logs
docker logs CONTAINER
docker logs -f CONTAINER
docker logs --tail 100 CONTAINER
docker compose logs -f
16. dmesg
Kernel or hardware problems:
dmesg
dmesg -T
17. uptime and load
uptime
Load average is not the same as CPU percentage.
18. iostat
After installing sysstat:
iostat -xz 1
Useful for I/O troubleshooting.
19. vmstat
vmstat 1
20. Network toolkit
ping
dig
curl
nc
ss
traceroute
tcpdump
21. 502 troubleshooting
- Check nginx.
- Check the backend.
- Check listening ports.
- Test the backend locally.
- Inspect nginx error logs.
22. “The site is down” checklist
- DNS,
- port 443,
- certificate,
- nginx,
- backend,
- database,
- resources.
23. Monitoring minimum
Monitor uptime, CPU, RAM, disk, health endpoint, container/service status and important log errors.
24. Alerts
A useful alert says what failed, where, since when, the relevant metric and ideally links to logs or a dashboard.
25. Do not alert on everything
Alert fatigue makes all alerts less useful. Alerts should correspond to issues that need action.
26. Log rotation
Logs need rotation. Linux commonly uses logrotate or journald retention.
27. What you should know
You should be able to read journalctl, locate processes and ports, assess CPU/RAM/disk, diagnose 502 errors, separate application problems from proxy/network problems and create a basic health check.
Official references
- systemd journalctl: https://www.freedesktop.org/software/systemd/man/latest/journalctl.html
- OpenTelemetry documentation: https://opentelemetry.io/docs/