Docker homelab pitfalls: disk space, logs, and config drift
A developer shares critical lessons from managing a Docker homelab, covering image cleanup risks, unbounded log growth, and configuration persistence issues.
A recent technical post details several operational hazards encountered while managing a self-hosted Docker environment on Proxmox and network-attached storage. The author describes how standard diagnostic commands can mislead administrators into deleting active services or overlooking critical configuration drift. These incidents highlight the gap between container health status and actual service functionality in complex home lab setups.
What happened
The investigation began when a Docker host reached 82 percent disk capacity. Initial diagnostics using docker system df suggested 6.67 GB of reclaimable space, but closer inspection revealed that summary views were misleading. The command docker ps displayed bare SHA identifiers instead of image names for several containers, making an active image named unclecode/crawl4ai appear orphaned. The container was healthy, had been running for two weeks, and was serving traffic. The administrator avoided deleting it only by cross-referencing listening ports with container inspections.
Further analysis showed that disk pressure came from images and build caches rather than user data. One guest held 16 GB of images, while another had 11.4 GB of images and 11.3 GB of build cache. Removing sixteen superseded tags of a custom application freed only 60 MB despite each tag listing as 240 MB, due to shared layers. A separate incident involved a different host reaching 100 percent disk usage because the Docker daemon lacked a configuration file to limit log sizes. A single Home Assistant log file had grown to 1.9 GB, causing five systemd units to fail without alerting anyone.
Configuration drift also caused extended outages. A file browser container remained down for six days because its restart policy had reverted to no despite the compose file specifying unless-stopped. This occurred because the container was created before the policy change and never recreated. Additionally, environment variable interpolation issues mangled an Argon2 password string, causing authentication failures, while hardcoded defaults in a tracing stack led to database connection errors. Network security gaps were found where published ports allowed direct access to private instances, bypassing reverse proxies.
Key details
- Disk cleanup tools like
docker system dfcan misrepresent space usage due to shared image layers and build caches. - Default Docker logging has no size cap, allowing single log files to fill entire filesystems if
daemon.jsonis not configured. - Changing
.envfiles or compose configurations requiresdocker compose up -d --force-recreateto take effect, not just a restart. - Containers created before a restart policy update retain their original policy until explicitly recreated or updated.
- Health checks indicate process status but do not verify functional connectivity, such as tunnel connections or GPU availability.
- Firewall rules for published ports must match original destination addresses using
--ctorigdstbecause of DNAT behavior.
Background
Docker containers are lightweight virtualized environments that share the host operating system kernel. When a container is created, Docker captures the configuration, including environment variables, restart policies, and logging drivers, into a specific runtime state. Subsequent changes to source files, such as compose YAML files or environment variable definitions, do not automatically update running containers. Administrators must explicitly recreate containers to apply these changes. This behavior often leads to "configuration drift," where the live system diverges from the declared infrastructure code.
Logging in Docker typically uses a JSON file driver by default, which writes all standard output and error streams to disk without rotation unless configured otherwise. In production or long-running home labs, this can lead to rapid disk consumption. Similarly, image management involves layers that are shared across multiple tags and versions. Deleting a specific tag does not necessarily remove the underlying data blocks if other images reference them, making disk space reclamation less predictable than simple file deletion.
Why it matters
For teams running self-hosted software, these pitfalls represent significant reliability risks. Relying on superficial health checks can mask service failures, such as a tunnel agent that is running but not connected, or a machine learning container falling back to CPU due to hardware incompatibility. Without robust monitoring of actual service functionality, outages may persist for days unnoticed. The incident where a file browser stayed down for six days illustrates how silent failures can disrupt workflows without triggering immediate alerts.
Disk space exhaustion is another critical threat to availability. Unbounded log growth can crash entire hosts, taking down all co-located services simultaneously. The lack of default log rotation means that every new deployment inherits this risk unless explicitly mitigated. Furthermore, configuration drift undermines the reproducibility benefits of infrastructure-as-code practices. If environment variables or restart policies are not synchronized between the code repository and the running containers, debugging becomes difficult and deployments become unpredictable.
Network security is also compromised when default publishing behaviors are not carefully managed. Exposing private services on all interfaces allows unauthorized lateral movement within the network. Proper firewall configuration requires understanding Docker’s network address translation mechanics, which differ from standard host-based routing. Misconfigured rules can leave sensitive internal services accessible to less trusted segments, violating least-privilege principles.
What you can do
- Verify image usage with
docker inspectand check listening ports before removing any images or containers. - Configure
daemon.jsonwithmax-sizeandmax-filelimits for logs, then recreate existing containers to apply the settings. - Use
docker compose up -d --force-recreateafter changing environment files to ensure new values are loaded. - Audit restart policies regularly using
docker inspectto confirm they match the intended compose configuration. - Implement firewall rules in the
DOCKER-USERchain using--ctorigdstto restrict access to published ports. - Test service functionality beyond health checks by verifying external connectivity and internal dependencies periodically.



