Showcase · Docker infrastructure

Monitoring must not go down with what it monitors

Around 94 containers on a NAS, with a Mac as the second host. What is supposed to run sits in a manifest; monitoring compares reality against it. The most important choice is what does not run in Docker — that choice comes from a crash we went through ourselves.

~94
containers in production
2
hosts: NAS and Mac
3
services deliberately outside Docker
~8 GB
system partition — no room for data
Architecture

The manifest decides, reality deviates

Monitoring that only checks whether a container is running misses what is absent. By comparing reality against a recorded list, "something is supposed to be running here" becomes a measurable fact — and that monitoring itself runs outside the stack it checks.

Manifest desired containers what is supposed to run Diff manifest vs. reality deviation = signal Hosts NAS + Mac ~94 containers Health checks per container status & storage Native layer — systemd, outside the container layer three critical services deliberately outside Docker · if the stack goes down, monitoring, logging and execution keep running
Case · the crash

What went wrong, and what changed afterwards

A wrong command hit a home directory that held an SSH key. During the recovery the file system gave out. The current setup is not a design on paper — it is what was left after this went wrong.

Then

  1. Everything ran in the same Docker stack, monitoring included
  2. A wrong command hit a home directory with an SSH key in it
  3. During recovery the file system gave out
  4. The layer that should have been watching was down itself
  5. What was supposed to be running was recorded nowhere

Now

  1. Critical services run natively on the host system (systemd)
  2. If the container layer goes down, monitoring, logging and execution keep working
  3. A manifest records which containers are supposed to run
  4. Monitoring compares reality against that manifest and reports the deviation
  5. Timestamped backup before every change; verify first, then act
Tech in detail

Six choices that come from experience

Critical services don’t run in the stack

Three services have deliberately been taken out of Docker and run natively under systemd on the host system. They watch over the container layer, so they must not depend on it. If Docker goes down, monitoring, logging and the execution layer stay reachable.

Manifest as the source of truth

A "desired containers" manifest records which containers are supposed to run. Monitoring compares that against reality on both hosts. A missing container is a deviation, not an interpretation — and a container that is deliberately switched off is recorded as such in the manifest.

Working rules born from a real crash

No recursive deletions or permission changes on home directories or volume roots. That rule isn’t theoretical: this is exactly where it went wrong. The SSH key that provided access to the system sat in the directory that got hit.

Backup before every change

Every file that gets touched first gets a timestamped copy next to the original. No separate backup moment, no exceptions for "small" changes. Rolling back is therefore always a matter of renaming one file.

Verify first, then act

Read the file, check the port, read the log — before the command, not after. Most incidents in a homelab of this size don’t come from complex errors, but from assumptions about the state of the system.

Data does not belong on the system partition

The system partition of the NAS operating system is small (around 8 GB). Application data always writes to the data volume. If that partition fills beyond roughly 70%, that isn’t a capacity question but a signal that something is landing in the wrong place.

Built with Docker ComposesystemdPortainerNAS + MacHealth-checksManifest-diff
Container infrastructure

Does your monitoring run in the stack it watches over?

We manage around 94 containers across two hosts and changed the architecture after it went genuinely wrong once. That experience is usable for your environment.