--- title: systemd unit restarted 2 million times unnoticed: "Failed to load environment files: No such file or directory" date: 2026-09-10 summary: Template units whose EnvironmentFile and user had been deleted restart-looped every 3 s for 73 days (1,958,017 restarts each). systemd fails before spawning a process, so CPU and ps show nothing, and RestartSec=3 stays under the start limit. Find them by NRestarts; vacuum the journal by size. tags: systemd, journald, linux, monitoring environment: Arch Linux (40-thread Xeon server); systemd and journald, versions not recorded author: shaun, written up with Claude status: published --- ## Symptoms Nothing visible: no alerts, no CPU load, no stray processes. A sweep for restart loops found three instances of a template unit, left enabled after the application they belonged to had been removed (its user, home directory and files were gone): ```ini [Unit] Description=Old agent bridge (%i) After=network-online.target [Service] User=olduser WorkingDirectory=/home/olduser/runtime EnvironmentFile=/home/olduser/agents/%i/agent.env ExecStart=/usr/bin/python3 /home/olduser/runtime/bridge.py Restart=always RestartSec=3 ``` The journal had this every 3 seconds per instance, for 73 days: ```text myapp@a.service: Scheduled restart job, restart counter is at 1958017. myapp@a.service: Failed to load environment files: No such file or directory myapp@a.service: Failed to spawn 'start' task: No such file or directory myapp@a.service: Failed with result 'resources'. ``` Each of the three instances had restarted 1,958,017 times (about 5.9 million in total). A related daemon those units had depended on had been in `failed` state for the same 73 days. The journal had grown to 4 G, almost all of it from this loop. ## Root cause Two things combined to keep it invisible: 1. **No process is ever started.** `EnvironmentFile=` points to a file that no longer exists, so systemd fails the start *before* `ExecStart=` runs (`Failed with result 'resources'`). The unit status showed `Mem peak: 0B, CPU: 0`. CPU or memory monitoring, `ps` and the load average all see nothing. 2. **`RestartSec=3` stays under the start rate limit.** With the defaults (`StartLimitBurst=5` starts within `StartLimitIntervalSec=10s`), a fast loop is stopped and left in `failed` state. One start every 3 seconds is about 3 to 4 starts per 10 seconds, which never reaches 5, so `Restart=always` keeps going forever. The units were orphans: the application was removed, but its unit files stayed enabled. ## Fix Stop and disable the units, then clear the failed state: ```bash systemctl disable --now myapp@a.service myapp@b.service myapp@c.service systemctl reset-failed ``` Then remove the leftover unit file (and any drop-in directories for related units) and reload: ```bash rm -f /etc/systemd/system/myapp@.service systemctl daemon-reload ``` Afterwards: no unit files left for the old application, and no units in `failed` or `activating` state. ### Journal cleanup: size, not time ```bash journalctl --vacuum-size=500M # 4G -> 467.8M, freed about 3.5G ``` `journalctl --vacuum-time=30d` freed **0 B** here. Time-based vacuuming only removes archived journal files whose newest entry is older than the cutoff. The loop rotated a new 80 MB file about every 24 minutes, so every retained file was recent. Vacuuming by size works. Check journal size **as root**. As an ordinary user, `journalctl --disk-usage` reported only that user's own journal (`8M`), while the system total was 467.8 M after cleanup. ### Removing the orphaned package The related daemon's package was removed with `pacman -Rns`. Preview the cascade first: `pacman -Rs --print ` (pacman rejects `-Rns --print`, because `--nosave` and `--print` cannot be combined). The preview listed a password-hashing library among the dependencies to be removed; `pacman -Qi ` showed `Required By:` only the package being removed, so it was safe. ## How it was found List every service with a non-trivial restart counter: ```bash systemctl list-units --type=service --all --no-legend --plain | awk '{print $1}' | while read u; do n=$(systemctl show "$u" -p NRestarts --value 2>/dev/null) case "$n" in ""|0|1|2|3) ;; *) echo "$u NRestarts=$n";; esac done ``` Also useful: ```bash systemctl list-units --state=failed,activating --all ``` A unit that keeps showing `activating (auto-restart)` is in a restart loop. The same sweep on the same day found a different crash loop that used 2.24 CPU cores. That one was easy to see once CPU per cgroup was measured, but it would not stand out in restart counts. A sweep needs both: CPU per cgroup **and** `NRestarts`. ## References - `systemd.unit(5)`: `StartLimitIntervalSec=`, `StartLimitBurst=`. - `systemd.exec(5)`: `EnvironmentFile=` (a missing file fails the start unless the path is prefixed with `-`). - `journalctl(1)`: `--vacuum-size=`, `--vacuum-time=`.