known·good

Notes / systemdjournaldlinuxmonitoring

systemd unit restarted 2 million times unnoticed: "Failed to load environment files: No such file or directory"

· by shaun, written up with Claude

Template units whose EnvironmentFile and user had been deleted restart-looped every 3 s for 73 days (1,958,017 restarts each). systemd fails before spawning a process, so CPU and ps show nothing, and RestartSec=3 stays under the start limit. Find them by NRestarts; vacuum the journal by size.

Tested onArch Linux (40-thread Xeon server); systemd and journald, versions not recorded

Symptoms #

Nothing visible: no alerts, no CPU load, no stray processes. A sweep for restart loops found three instances of a template unit, left enabled after the application they belonged to had been removed (its user, home directory and files were gone):

ini[Unit]
Description=Old agent bridge (%i)
After=network-online.target

[Service]
User=olduser
WorkingDirectory=/home/olduser/runtime
EnvironmentFile=/home/olduser/agents/%i/agent.env
ExecStart=/usr/bin/python3 /home/olduser/runtime/bridge.py
Restart=always
RestartSec=3

The journal had this every 3 seconds per instance, for 73 days:

textmyapp@a.service: Scheduled restart job, restart counter is at 1958017.
myapp@a.service: Failed to load environment files: No such file or directory
myapp@a.service: Failed to spawn 'start' task: No such file or directory
myapp@a.service: Failed with result 'resources'.

Each of the three instances had restarted 1,958,017 times (about 5.9 million in total). A related daemon those units had depended on had been in failed state for the same 73 days.

The journal had grown to 4 G, almost all of it from this loop.

Root cause #

Two things combined to keep it invisible:

  1. No process is ever started. EnvironmentFile= points to a file that no longer exists, so systemd fails the start before ExecStart= runs (Failed with result 'resources'). The unit status showed Mem peak: 0B, CPU: 0. CPU or memory monitoring, ps and the load average all see nothing.
  2. RestartSec=3 stays under the start rate limit. With the defaults (StartLimitBurst=5 starts within StartLimitIntervalSec=10s), a fast loop is stopped and left in failed state. One start every 3 seconds is about 3 to 4 starts per 10 seconds, which never reaches 5, so Restart=always keeps going forever.

The units were orphans: the application was removed, but its unit files stayed enabled.

Fix #

Stop and disable the units, then clear the failed state:

bashsystemctl disable --now myapp@a.service myapp@b.service myapp@c.service
systemctl reset-failed

Then remove the leftover unit file (and any drop-in directories for related units) and reload:

bashrm -f /etc/systemd/system/myapp@.service
systemctl daemon-reload

Afterwards: no unit files left for the old application, and no units in failed or activating state.

Journal cleanup: size, not time #

bashjournalctl --vacuum-size=500M      # 4G -> 467.8M, freed about 3.5G

journalctl --vacuum-time=30d freed 0 B here. Time-based vacuuming only removes archived journal files whose newest entry is older than the cutoff. The loop rotated a new 80 MB file about every 24 minutes, so every retained file was recent. Vacuuming by size works.

Check journal size as root. As an ordinary user, journalctl --disk-usage reported only that user's own journal (8M), while the system total was 467.8 M after cleanup.

Removing the orphaned package #

The related daemon's package was removed with pacman -Rns. Preview the cascade first: pacman -Rs --print <pkg> (pacman rejects -Rns --print, because --nosave and --print cannot be combined). The preview listed a password-hashing library among the dependencies to be removed; pacman -Qi <lib> showed Required By: only the package being removed, so it was safe.

How it was found #

List every service with a non-trivial restart counter:

bashsystemctl list-units --type=service --all --no-legend --plain | awk '{print $1}' | while read u; do
    n=$(systemctl show "$u" -p NRestarts --value 2>/dev/null)
    case "$n" in ""|0|1|2|3) ;; *) echo "$u NRestarts=$n";; esac
done

Also useful:

bashsystemctl list-units --state=failed,activating --all

A unit that keeps showing activating (auto-restart) is in a restart loop.

The same sweep on the same day found a different crash loop that used 2.24 CPU cores. That one was easy to see once CPU per cgroup was measured, but it would not stand out in restart counts. A sweep needs both: CPU per cgroup and NRestarts.

References #

  • systemd.unit(5): StartLimitIntervalSec=, StartLimitBurst=.
  • systemd.exec(5): EnvironmentFile= (a missing file fails the start unless the path is prefixed with -).
  • journalctl(1): --vacuum-size=, --vacuum-time=.