known·good

Notes / nvidiaarchlinuxxorgdkmskernel

Black screen after restarting X: NVIDIA: Failed to initialize the NVIDIA kernel module (kernel upgraded, not rebooted)

· by shaun, written up with Claude

After a kernel and NVIDIA driver upgrade without a reboot, restarting X gave a black screen and "Failed to initialize the NVIDIA kernel module". The running kernel had only the old module loaded. Build the module for the running kernel with DKMS, swap the modules, recreate /dev/dri, then reboot properly.

Tested onArch Linux; running kernel 6.18.13-arch1-1, installed kernel 7.0.12-arch1-1; nvidia-open-dkms 610.43.02 (old loaded module 590.48.01); LightDM (SDDM also installed); Chrome Remote Desktop running its own Xorg

Symptoms #

A GPU application crashed X on the desktop. Restarting X then gave a permanent black screen. The display manager log showed X dying straight away:

textFailed to read display number from pipe

The real error was in the Xorg log (/var/log/Xorg.0.log):

text(EE) NVIDIA: Failed to initialize the NVIDIA kernel module

dmesg had an API mismatch message: the client (the Xorg driver) was 610.43.02 and the kernel module was 590.48.01.

Root cause #

The kernel and the NVIDIA driver had been upgraded with pacman, but the machine had not been rebooted:

bashuname -r                          # 6.18.13-arch1-1   (running)
pacman -Q linux                   # linux 7.0.12.arch1-1  (installed)
cat /proc/driver/nvidia/version   # 590.48.01  (module in memory)
dkms status                       # nvidia 610.43.02 built for 7.0.12 only
  1. DKMS built nvidia-open-dkms 610.43.02 only for the new kernel, 7.0.12. There was no 610 module for the running kernel, 6.18.13.
  2. The old 590.48.01 module was still loaded. The new 610.43.02 Xorg driver cannot talk to it, so X fails at startup.
  3. The old module could not simply be unloaded: a second Xorg server (Chrome Remote Desktop runs its own, here on display :20) still had it open. fuser /dev/nvidia* /dev/dri/* showed that process.
  4. Even after that server was stopped, nvidia_drm still had kernel-internal DRM references (cat /sys/module/nvidia_drm/refcnt showed 9), so a normal rmmod failed.

The simple fix is to reboot into the new kernel. The steps below got a working display back on the old kernel without a reboot, then prepared a clean reboot.

Fix #

Live recovery (no reboot) #

Run as root. Replace the versions with yours.

  1. Install the headers for the running kernel (they were still in the pacman cache):

    bashpacman -U /var/cache/pacman/pkg/linux-headers-<running-kernel-version>-x86_64.pkg.tar.zst
    
  2. Rebuild the module database for the running kernel (the upgrade had removed its kernel/ module tree):

    bashdepmod -a <running-kernel-version>
    
  3. Build the new driver for the running kernel. The DKMS source is named nvidia, even when the installed package is nvidia-open-dkms:

    bashdkms install nvidia/<driver-version> -k <running-kernel-version>
    
  4. Stop everything that holds the module open (here, Chrome Remote Desktop's Xorg):

    bashsystemctl stop chrome-remote-desktop@<user>.service
    
  5. Unload the old modules. nvidia_drm needed a forced removal because of the stuck references:

    bashrmmod -f nvidia_drm
    rmmod nvidia_modeset nvidia
    
  6. Load the new modules and check the version:

    bashmodprobe -a nvidia nvidia_modeset nvidia_drm nvidia_uvm
    cat /proc/driver/nvidia/version     # must show the new version
    
  7. The forced rmmod deleted /dev/dri/. Re-add the GPU's PCI device so udev recreates it (find the address with lspci -D | grep -i nvidia, e.g. 0000:02:00.0):

    bashudevadm trigger --action=add /sys/bus/pci/devices/<gpu-pci-address>
    udevadm settle
    
  8. Start the display manager, then Chrome Remote Desktop:

    bashsystemctl start lightdm             # or sddm
    systemctl start chrome-remote-desktop@<user>.service
    

Notes:

  • modprobe does nothing if a module of that name is already loaded, even an old version. Always check /proc/driver/nvidia/version after loading.
  • modprobe with several module names needs -a. Without it, the names after the first are passed as parameters to the first module. (Our original notes left out -a; the version check in step 6 is what showed the right module was loaded.)
  • rmmod -f is a last resort. It was used here only because X was already dead and nothing else would release nvidia_drm.
  • Start the main display manager before restarting Chrome Remote Desktop. Its Xorg holds the module and, if it comes up first, can take over the user session.

This recovery was only partly clean. X came back, but the desktop session started on the remote desktop's display instead of the physical one. Two settings caused that:

  • ~/.dmrc still had Session=xfce from an old desktop switch, so LightDM started a session that exited at once, leaving only a cursor on the physical screen.
  • ReuseSession=true in the SDDM config (SDDM was used during the recovery) made it attach to the existing remote session instead of starting a new one.

As a stopgap on the physical console: startx /usr/bin/startlxqt -- :0 vt2.

Before rebooting into the new kernel #

Step 1 installs the old kernel's headers, which replaces the new kernel's headers. DKMS then had the driver for the new kernel only "added", not "installed". Put the new headers back and check before rebooting:

bashpacman -U /var/cache/pacman/pkg/linux-headers-<new-kernel-version>-x86_64.pkg.tar.zst
dkms status -k <new-kernel-version>

The DKMS pacman hook rebuilt the module when the headers were installed. dkms status should then show nvidia/<driver-version>, <new-kernel-version>, x86_64: installed.

Also fix the default session in ~/.dmrc (here Session=xfce → Session=lxqt).

After the reboot into the new kernel, the journal had no NVIDIA errors and systemctl --failed listed no units.

How it was found #

  • journalctl -u lightdm -u sddm showed X exiting at once; /var/log/Xorg.0.log gave the real error.
  • uname -r against pacman -Q linux showed the pending reboot.
  • /proc/driver/nvidia/version and dkms status showed the loaded module was older than the userspace driver, with no matching build for the running kernel.
  • fuser /dev/nvidia* /dev/dri/* found the second Xorg holding the module; /sys/module/nvidia_drm/refcnt showed the leftover references after it stopped.
  • ls /tmp/.X*-lock and ss -xl | grep X11 listed the running X displays (:0 and :20).

The lesson: after a pacman upgrade that touches the kernel or the NVIDIA driver, reboot before restarting X.