Notes / nvidiaarchlinuxxorgdkmskernel
Black screen after restarting X: NVIDIA: Failed to initialize the NVIDIA kernel module (kernel upgraded, not rebooted)
After a kernel and NVIDIA driver upgrade without a reboot, restarting X gave a black screen and "Failed to initialize the NVIDIA kernel module". The running kernel had only the old module loaded. Build the module for the running kernel with DKMS, swap the modules, recreate /dev/dri, then reboot properly.
Symptoms #
A GPU application crashed X on the desktop. Restarting X then gave a permanent black screen. The display manager log showed X dying straight away:
textFailed to read display number from pipe
The real error was in the Xorg log (/var/log/Xorg.0.log):
text(EE) NVIDIA: Failed to initialize the NVIDIA kernel module
dmesg had an API mismatch message: the client (the Xorg driver) was 610.43.02 and the
kernel module was 590.48.01.
Root cause #
The kernel and the NVIDIA driver had been upgraded with pacman, but the machine had not been rebooted:
bashuname -r # 6.18.13-arch1-1 (running)
pacman -Q linux # linux 7.0.12.arch1-1 (installed)
cat /proc/driver/nvidia/version # 590.48.01 (module in memory)
dkms status # nvidia 610.43.02 built for 7.0.12 only
- DKMS built
nvidia-open-dkms 610.43.02only for the new kernel, 7.0.12. There was no 610 module for the running kernel, 6.18.13. - The old 590.48.01 module was still loaded. The new 610.43.02 Xorg driver cannot talk to it, so X fails at startup.
- The old module could not simply be unloaded: a second Xorg server (Chrome Remote Desktop
runs its own, here on display
:20) still had it open.fuser /dev/nvidia* /dev/dri/*showed that process. - Even after that server was stopped,
nvidia_drmstill had kernel-internal DRM references (cat /sys/module/nvidia_drm/refcntshowed9), so a normalrmmodfailed.
The simple fix is to reboot into the new kernel. The steps below got a working display back on the old kernel without a reboot, then prepared a clean reboot.
Fix #
Live recovery (no reboot) #
Run as root. Replace the versions with yours.
-
Install the headers for the running kernel (they were still in the pacman cache):
bash
pacman -U /var/cache/pacman/pkg/linux-headers-<running-kernel-version>-x86_64.pkg.tar.zst -
Rebuild the module database for the running kernel (the upgrade had removed its
kernel/module tree):bash
depmod -a <running-kernel-version> -
Build the new driver for the running kernel. The DKMS source is named
nvidia, even when the installed package isnvidia-open-dkms:bash
dkms install nvidia/<driver-version> -k <running-kernel-version> -
Stop everything that holds the module open (here, Chrome Remote Desktop's Xorg):
bash
systemctl stop chrome-remote-desktop@<user>.service -
Unload the old modules.
nvidia_drmneeded a forced removal because of the stuck references:bash
rmmod -f nvidia_drm rmmod nvidia_modeset nvidia -
Load the new modules and check the version:
bash
modprobe -a nvidia nvidia_modeset nvidia_drm nvidia_uvm cat /proc/driver/nvidia/version # must show the new version -
The forced
rmmoddeleted/dev/dri/. Re-add the GPU's PCI device so udev recreates it (find the address withlspci -D | grep -i nvidia, e.g.0000:02:00.0):bash
udevadm trigger --action=add /sys/bus/pci/devices/<gpu-pci-address> udevadm settle -
Start the display manager, then Chrome Remote Desktop:
bash
systemctl start lightdm # or sddm systemctl start chrome-remote-desktop@<user>.service
Notes:
modprobedoes nothing if a module of that name is already loaded, even an old version. Always check/proc/driver/nvidia/versionafter loading.modprobewith several module names needs-a. Without it, the names after the first are passed as parameters to the first module. (Our original notes left out-a; the version check in step 6 is what showed the right module was loaded.)rmmod -fis a last resort. It was used here only because X was already dead and nothing else would releasenvidia_drm.- Start the main display manager before restarting Chrome Remote Desktop. Its Xorg holds the module and, if it comes up first, can take over the user session.
This recovery was only partly clean. X came back, but the desktop session started on the remote desktop's display instead of the physical one. Two settings caused that:
~/.dmrcstill hadSession=xfcefrom an old desktop switch, so LightDM started a session that exited at once, leaving only a cursor on the physical screen.ReuseSession=truein the SDDM config (SDDM was used during the recovery) made it attach to the existing remote session instead of starting a new one.
As a stopgap on the physical console: startx /usr/bin/startlxqt -- :0 vt2.
Before rebooting into the new kernel #
Step 1 installs the old kernel's headers, which replaces the new kernel's headers. DKMS then had the driver for the new kernel only "added", not "installed". Put the new headers back and check before rebooting:
bashpacman -U /var/cache/pacman/pkg/linux-headers-<new-kernel-version>-x86_64.pkg.tar.zst
dkms status -k <new-kernel-version>
The DKMS pacman hook rebuilt the module when the headers were installed. dkms status should
then show nvidia/<driver-version>, <new-kernel-version>, x86_64: installed.
Also fix the default session in ~/.dmrc (here Session=xfce → Session=lxqt).
After the reboot into the new kernel, the journal had no NVIDIA errors and
systemctl --failed listed no units.
How it was found #
journalctl -u lightdm -u sddmshowed X exiting at once;/var/log/Xorg.0.loggave the real error.uname -ragainstpacman -Q linuxshowed the pending reboot./proc/driver/nvidia/versionanddkms statusshowed the loaded module was older than the userspace driver, with no matching build for the running kernel.fuser /dev/nvidia* /dev/dri/*found the second Xorg holding the module;/sys/module/nvidia_drm/refcntshowed the leftover references after it stopped.ls /tmp/.X*-lockandss -xl | grep X11listed the running X displays (:0and:20).
The lesson: after a pacman upgrade that touches the kernel or the NVIDIA driver, reboot before restarting X.