Notes / systemdjournaldlinuxmonitoring
systemd unit restarted 2 million times unnoticed: "Failed to load environment files: No such file or directory"
Template units whose EnvironmentFile and user had been deleted restart-looped every 3 s for 73 days (1,958,017 restarts each). systemd fails before spawning a process, so CPU and ps show nothing, and RestartSec=3 stays under the start limit. Find them by NRestarts; vacuum the journal by size.
Symptoms #
Nothing visible: no alerts, no CPU load, no stray processes. A sweep for restart loops found three instances of a template unit, left enabled after the application they belonged to had been removed (its user, home directory and files were gone):
ini[Unit]
Description=Old agent bridge (%i)
After=network-online.target
[Service]
User=olduser
WorkingDirectory=/home/olduser/runtime
EnvironmentFile=/home/olduser/agents/%i/agent.env
ExecStart=/usr/bin/python3 /home/olduser/runtime/bridge.py
Restart=always
RestartSec=3
The journal had this every 3 seconds per instance, for 73 days:
textmyapp@a.service: Scheduled restart job, restart counter is at 1958017.
myapp@a.service: Failed to load environment files: No such file or directory
myapp@a.service: Failed to spawn 'start' task: No such file or directory
myapp@a.service: Failed with result 'resources'.
Each of the three instances had restarted 1,958,017 times (about 5.9 million in total). A
related daemon those units had depended on had been in failed state for the same 73 days.
The journal had grown to 4 G, almost all of it from this loop.
Root cause #
Two things combined to keep it invisible:
- No process is ever started.
EnvironmentFile=points to a file that no longer exists, so systemd fails the start beforeExecStart=runs (Failed with result 'resources'). The unit status showedMem peak: 0B, CPU: 0. CPU or memory monitoring,psand the load average all see nothing. RestartSec=3stays under the start rate limit. With the defaults (StartLimitBurst=5starts withinStartLimitIntervalSec=10s), a fast loop is stopped and left infailedstate. One start every 3 seconds is about 3 to 4 starts per 10 seconds, which never reaches 5, soRestart=alwayskeeps going forever.
The units were orphans: the application was removed, but its unit files stayed enabled.
Fix #
Stop and disable the units, then clear the failed state:
bashsystemctl disable --now myapp@a.service myapp@b.service myapp@c.service
systemctl reset-failed
Then remove the leftover unit file (and any drop-in directories for related units) and reload:
bashrm -f /etc/systemd/system/myapp@.service
systemctl daemon-reload
Afterwards: no unit files left for the old application, and no units in failed or
activating state.
Journal cleanup: size, not time #
bashjournalctl --vacuum-size=500M # 4G -> 467.8M, freed about 3.5G
journalctl --vacuum-time=30d freed 0 B here. Time-based vacuuming only removes archived
journal files whose newest entry is older than the cutoff. The loop rotated a new 80 MB file
about every 24 minutes, so every retained file was recent. Vacuuming by size works.
Check journal size as root. As an ordinary user, journalctl --disk-usage reported only that
user's own journal (8M), while the system total was 467.8 M after cleanup.
Removing the orphaned package #
The related daemon's package was removed with pacman -Rns. Preview the cascade first:
pacman -Rs --print <pkg> (pacman rejects -Rns --print, because --nosave and --print
cannot be combined). The preview listed a password-hashing library among the dependencies to be
removed; pacman -Qi <lib> showed Required By: only the package being removed, so it was safe.
How it was found #
List every service with a non-trivial restart counter:
bashsystemctl list-units --type=service --all --no-legend --plain | awk '{print $1}' | while read u; do
n=$(systemctl show "$u" -p NRestarts --value 2>/dev/null)
case "$n" in ""|0|1|2|3) ;; *) echo "$u NRestarts=$n";; esac
done
Also useful:
bashsystemctl list-units --state=failed,activating --all
A unit that keeps showing activating (auto-restart) is in a restart loop.
The same sweep on the same day found a different crash loop that used 2.24 CPU cores. That one
was easy to see once CPU per cgroup was measured, but it would not stand out in restart counts.
A sweep needs both: CPU per cgroup and NRestarts.
References #
systemd.unit(5):StartLimitIntervalSec=,StartLimitBurst=.systemd.exec(5):EnvironmentFile=(a missing file fails the start unless the path is prefixed with-).journalctl(1):--vacuum-size=,--vacuum-time=.