Monitoring must not go down with what it monitors
Around 94 containers on a NAS, with a Mac as the second host. What is supposed to run sits in a manifest; monitoring compares reality against it. The most important choice is what does not run in Docker — that choice comes from a crash we went through ourselves.
The manifest decides, reality deviates
Monitoring that only checks whether a container is running misses what is absent. By comparing reality against a recorded list, "something is supposed to be running here" becomes a measurable fact — and that monitoring itself runs outside the stack it checks.
What went wrong, and what changed afterwards
A wrong command hit a home directory that held an SSH key. During the recovery the file system gave out. The current setup is not a design on paper — it is what was left after this went wrong.
Then
- Everything ran in the same Docker stack, monitoring included
- A wrong command hit a home directory with an SSH key in it
- During recovery the file system gave out
- The layer that should have been watching was down itself
- What was supposed to be running was recorded nowhere
Now
- Critical services run natively on the host system (systemd)
- If the container layer goes down, monitoring, logging and execution keep working
- A manifest records which containers are supposed to run
- Monitoring compares reality against that manifest and reports the deviation
- Timestamped backup before every change; verify first, then act
Six choices that come from experience
Critical services don’t run in the stack
Three services have deliberately been taken out of Docker and run natively under systemd on the host system. They watch over the container layer, so they must not depend on it. If Docker goes down, monitoring, logging and the execution layer stay reachable.
Manifest as the source of truth
A "desired containers" manifest records which containers are supposed to run. Monitoring compares that against reality on both hosts. A missing container is a deviation, not an interpretation — and a container that is deliberately switched off is recorded as such in the manifest.
Working rules born from a real crash
No recursive deletions or permission changes on home directories or volume roots. That rule isn’t theoretical: this is exactly where it went wrong. The SSH key that provided access to the system sat in the directory that got hit.
Backup before every change
Every file that gets touched first gets a timestamped copy next to the original. No separate backup moment, no exceptions for "small" changes. Rolling back is therefore always a matter of renaming one file.
Verify first, then act
Read the file, check the port, read the log — before the command, not after. Most incidents in a homelab of this size don’t come from complex errors, but from assumptions about the state of the system.
Data does not belong on the system partition
The system partition of the NAS operating system is small (around 8 GB). Application data always writes to the data volume. If that partition fills beyond roughly 70%, that isn’t a capacity question but a signal that something is landing in the wrong place.
Does your monitoring run in the stack it watches over?
We manage around 94 containers across two hosts and changed the architecture after it went genuinely wrong once. That experience is usable for your environment.