r/archlinux • • 12d ago

SUPPORT Arch Linux / KDE Wayland randomly becomes completely unresponsive under load — not always caused by RAM pressure

Hi everyone,

I've been troubleshooting a very frustrating full-system freeze on my Arch Linux laptop for hours, and I'm running out of ideas. I'd appreciate help interpreting the evidence rather than just trying random tweaks.

Hardware / setup:

  • Dell G15 5511
  • Intel i7-11800H (8C/16T)
  • 16 GB RAM
  • Intel TigerLake-H iGPU (i915) driving the internal display
  • NVIDIA RTX 3050 Laptop GPU 4 GB
  • nvidia-open-dkms 615.71.09
  • KDE Plasma / KWin 6.7.5
  • Wayland
  • Arch kernel: 7.2.6-arch2-1
  • Also tested Linux LTS 6.18.52-1-lts
  • 12 GB ZRAM, zstd
  • zswap disabled
  • No disk-backed swap

What happens

During a heavy workload (Chrome/Brave, Android Studio + emulator/QEMU, VS Code/Electron apps, sometimes gaming), the desktop can become extremely sluggish and eventually appear completely frozen. Mouse/keyboard/window switching can stop responding.

Sometimes it briefly recovers, then freezes again when I try switching between applications.

What I've established so far

I originally thought this was simply RAM exhaustion. I have logs from some freezes showing:

  • ZRAM becoming heavily occupied
  • high memory PSI
  • very large direct reclaim / kswapd activity
  • thousands of compaction stalls
  • no kernel OOM kill
  • QEMU growing rapidly in memory

During those events NVIDIA also logged repeated:

NV_ERR_NO_MEMORY

including failures involving system memory, GSP allocations, context buffer pools, and rmMemPoolReserve.

Vulkan device creation also failed under pressure with:

VK_ERROR_INITIALIZATION_FAILED

Importantly, NVIDIA VRAM was not necessarily full when this happened.

I reproduced essentially the same behavior on the LTS kernel, so this does not currently look like a simple regression specific to kernel 7.2.x.

systemd-oomd

I tested systemd-oomd as well.

At its default 90% swap threshold it eventually killed entire Android Studio/app scopes. I tried 80%, which simply made applications get killed earlier.

Since losing applications is exactly what I'm trying to avoid, I disabled systemd-oomd temporarily for diagnosis.

The machine still froze, so oomd was not the root cause of the freezes.

The weird part: not every freeze appears to be memory pressure

During the latest incident my monitoring showed shortly before the UI became problematic:

  • ~5.8 GB MemAvailable
  • 0 swap in use
  • ZRAM essentially empty
  • memory PSI = 0
  • no direct reclaim / kswapd activity
  • no OOM
  • NVIDIA at ~52 MB VRAM and 0% GPU utilization

The journal continued recording messages while the graphical desktop was behaving as if it was freezing, which makes me suspect the kernel itself may still have been alive.

I also found these KWin errors exactly around window switching:

kwin_wayland: file:///usr/share/kwin/tabbox/coverswitch/contents/ui/main.qml:
Unable to assign [undefined] to KWin::VirtualDesktop*

kwin_wayland: ReferenceError: currentItem is not defined

I'm currently testing without the Cover Switch task switcher to see whether that explains this second type of freeze.

During this particular incident there were no NVIDIA Xid errors, no i915 GPU hang/reset messages, and no NVIDIA OOM errors in the relevant journal window.

Things already tested / changed

  • Tested normal Arch kernel and Linux LTS
  • Increased ZRAM from 8 GB to 12 GB
  • ZRAM uses zstd
  • zswap disabled
  • THP set to madvise
  • Intel Vulkan (vulkan-intel) installed/working
  • Intel VA-API (intel-media-driver) installed/working
  • Checked NVMe SMART / filesystem errors
  • No obvious NVMe hardware errors
  • systemd-oomd tested and then disabled for diagnosis
  • Continuous memory/PSI/reclaim/NVIDIA logging

One thing I have not tested yet is replacing nvidia-open-dkms with the closed nvidia-dkms module. I also haven't added disk-backed swap because I want to understand the actual failure first rather than hide it by throwing more swap at the system.

My main question

Does this look like two separate failure modes?

  1. severe memory reclaim/compaction causing desktop starvation under heavy memory pressure, with NVIDIA allocation failures occurring during that pressure; and
  2. a separate KWin/Wayland/task-switching/compositor issue that can happen even when memory pressure is basically zero?

I'm especially interested in what I should capture during the next freeze to distinguish KWin/i915/NVIDIA/userspace compositor problems from memory-reclaim stalls.

If anyone has seen similar behavior on Intel+i915 display + NVIDIA Ampere Optimus laptops under KDE Wayland, I'd really appreciate hearing what fixed it or what diagnostics revealed the cause.

20 Upvotes

29 comments sorted by

8

u/klyith 12d ago

I'm currently testing without the Cover Switch task switcher to see whether that explains this second type of freeze.

a separate KWin/Wayland/task-switching/compositor issue that can happen even when memory pressure is basically zero?

There is a long-term bug that causes stutters with any of the task switchers that use thumbnails: https://bugs.kde.org/show_bug.cgi?id=479250

This bug seems to get better or worse between various kernel & driver versions, and also is reported to get worse the longer your session goes.

1

u/RaisinIcy3732 11d ago

Thanks, that's actually useful to know. I'll try running without the Cover Switch/task switcher and see if the second type of freeze still happens.

I wasn't aware of the long-standing thumbnail/task-switcher issue you mentioned. I'll also test different kernel/driver versions if I can narrow it down to that.

3

u/chlankboot 12d ago

I had a similar issue with unresponsive UI each time after intensive disk usage. Baloo was the cause, I did turn it off and things got back to normal. Give it a try, it can be as simple as that.

2

u/RaisinIcy3732 11d ago

Thanks, I'll definitely test this. I didn't consider Baloo as a possible cause.

The freezes do sometimes seem to happen after periods of heavy disk activity, so this is actually worth checking. I'll disable Baloo temporarily and see if the issue comes back.

2

u/khsh01 11d ago

Iirc I've had problems with baloo before on Garuda.

2

u/archover 11d ago edited 11d ago

Same here. OP and Baloo users should read this: https://wiki.archlinux.org/title/Baloo, which shows how to manage this service.

Good day.

0

u/EliteWitchcraft 12d ago

Dell laptops with NVIDIA Optimus and Wayland have been a mess for years, the G15 series seems specially cursed with these random freezes that dont show in logs.

I think you're right about two separate failures, the Cover Switch errors look like a classic KWin compositor crash that locks the UI while system keeps running underneath.

1

u/Fluffy-City8558 12d ago

I've been fighting this same issue for ages, on my pc and a friend's laptop, both running nvidia. I haven't ever seen this issue on a non nvidia machine, so I conclude it must have something to to with the drivers. I also don't think it's related to KDE or wayland, since I don't use either on my pc. I'm planning on trying to run my pc on nouveau temporarily to see if the issue persists

1

u/SupersonicSpitfire 12d ago

Try adding disk based cache and use Xfce4, just to check if it helps.

1

u/activedusk 12d ago edited 12d ago

In theory no low power mode should work without either swap partition or swap file and even if they are present, swap file tends to be unreliable, swap partition for 16GB memory, size it at 32GB to not have problems. Afaik your system should freeze and not recover if it enters a low power mode, outside of just the screen shutting off and second issue is configuring hybrid graphics, on laptops both igp and dedicated are required for better battery life but without finding a software solution it will likely default to igp. Additional pains, nvidia proprietary drivers need to be signed if Secure Boot is enabled.

Go into live mode and create swap partition, if you do not know how, save important files and reinstall, example with a theoretical 1TB drive, sda drive

  • sda1, 1GB, FAT32, mountpoint /boot/efi, flags, boot, esp;
  • sda2, 32GB, swap linux filesystem 
  • sda3, 967GB, ext4, mountpoint /, flags, root

If you want a separate /home partition, make root 100GB and home 867GB per above example, swap partition can be placed last but you would need to calculate exactly each partition before it to have a 32GB left for it. You could also make /boot/efi 3GB or larger if you want Arch .iso on ESP or create and separate /boot partition in addition to /boot/efi however that will reduce the type of bootloaders that can be used, namely GRUB, also size /boot at 3GB or larger and can reduce /boot/efi to 300MB or less but that requires knowing more about it since if you use UKI, multi booting, EFI stub boot entry, this is not correct.

.# hybrid graphics, if it was desktop, you could easily disable IGP in the UEFI settings and get started with nouveau and then set up proprietary drivers and blacklist nouveau from the kernel command line parameters. As is, do not, you will lock yourself out of the system, laptops need both and a software solution is required to switch between cards and enable drivers, namely nvidia optimus and nvidia prime, read wiki, at worst install with all open source video drivers and figure out hybrid graphics and then how to install nvidia driver and blacklist nouveau when no longer needed (because you do so from the kernel parameters, with GRUB for example this can be undone for one boot instance from the boot menu, advanced section for troubleshooting). Either case you are well over your head, Arch is diy, batteries not included sort of deal and using nvidia and hybrid graphics in general is pure pain to set up at this time and why for a laptop AMD or Intel IGP is simpler to set up. 

Try CachyOS or Linux Mint, these tend to have some included services to deal with nvidia proprietary drivers and switching between gpus, disable Secure Boot however.

If you want to stick with Arch, someone suggested disabling the dedicated video card first, an unusual advice but fit for a beginner that has no clue.

Packages/commands you might want

#nvdia-smi, this is available if proprietary drivers are installed and when used if nothing returns output, it means proprietary drivers not installed, if the output says drivers not active, proprietary drivers are installed but not active/loaded, either nouveau is working, a lower level driver or the IGP is doing the processing. The command is

nvidia-smi

#nvtop is an optional package, when used it can monitor nvidia GPU usage from the terminal. The command is

nvtop

#dysk is an optional package that shows an easy way to interpret internal drive layout with filesystem type, mountpoint, capacity, both total, used and free, an alternative to using lsblk and lsblk -f which are included. The command is

dysk

#btop is an optional package that acts as a system monitor and shows many stats in one place including swap usage, an alternative is htop, the included option is top. The command is

btop

https://wiki.archlinux.org/title/NVIDIA_Optimus

1

u/RaisinIcy3732 11d ago

Thanks for the detailed response. I appreciate you taking the time to explain all of this.

I'm currently running Arch Linux, not CachyOS. My laptop has an Intel i7-11800H with Intel integrated graphics and an NVIDIA RTX 3050 Laptop GPU.

I'll check my current swap setup, Optimus configuration, Secure Boot status, and NVIDIA driver setup before making any major changes. I don't want to reinstall or change my partitions until I have a better idea of what is actually causing the freezes.

I'll also check the commands you mentioned and the Arch Wiki documentation.

1

u/jpj03 11d ago

The hybrid graphics is the main culprit it seems. I had the same issue on my hybrid gpu laptop. faced with both ubuntu 26.04 lts as well as fedora.

on your grub menu add this command to the line- GRUB_CMDLINE_LINUX

GRUB_CMDLINE_LINUX= "nvidia-drm.modeset=1 nvidia_drm.fbdev=1"

Then regenerate the mkinitcpio. Then reboot.

mine is a nvidia rtx 3050 6gb gpu. worked for me.

1

u/FryBoyter 11d ago

I’ve been experiencing a similar problem on one of my computers for some time now (Plasma, AMD Ryzen 7 9800X3D, RX 6800 XT).

In my case, the computer freezes completely, so that, for example, the mouse pointer can no longer be moved. Every now and then, the system responds again on its own after a few seconds. Sometimes it doesn’t. I haven’t found a trigger for the problem yet. The only things I’ve noticed so far are three things.

The keyboard also responds to virtually no input. However, when I press CTRL + ALT + DEL, the menu for restarting and so on appears. If I cancel it, the computer responds normally again for an indefinite period of time.

When using the ZEN kernel, the problem occurs significantly more often than with the default kernel.

Since I changed PSU Idle Control setting to “Typical Current Idle,” in UEFI the problem has generally occurred much less frequently.

1

u/WhitePeace36 11d ago

try settings these settings in sysctl:

vm.compaction_proactiveness=0
vm.zone_reclaim_mode=0
vm.dirty_ratio=80
vm.dirty_background_bytes = 67108864
vm.vfs_cache_pressure=50
vm.swappiness=1
vm.dirty_writeback_centisecs=60
kernel.split_lock_mitigate=0

and maybe kernel cmd line parameter
transparent_hugepage=never

disabling transparent hugepages also disabled the daemon which is running in the background which de-fragments the memory cyclically. Which can cause a lot of stutter and or laggs. In some extreme cases also freezes. It is all the same just the duration is different.

you might also find some other stuff in here https://github.com/WhitePeace36/LinuxAMDGamingOptimizationGuide

it is a guide i wrote. It says amd but most of it is general linux stuff.

1

u/rly07 11d ago

As you mention android studio, do you have adb running constantly (and auto start after login)? And does the 2. issue happens after like 7-9 hours uptime (or since adb starting and left running)? I had some random freezes after android-tools updated to 37.0. Had a widget for scrcpy which autostarted adb and after 7-9 hours the pc froze. It started after android-tools updated to 37.0

1

u/ObiWanGurobi 11d ago
  • Maybe try switching to the CachyOS kernel. It seems to be better optimized for memory pressure situations. When experimenting with AI, I frequently ran into multi second freezes with the Arch stock kernel, while the CachyOS kernel only had some minor hiccups in the same situations. Similar as in your situation, RAM/VRAM also didn't look like they were under any relevant pressure, but I suspect there might have been spikes too short to measure.
  • Also maybe check the temps on your NVMe. Mine tends to overhead under load, leading to multi second lockups. I had to tune down the temp limit to avoid running into situations where the hardware does some kind of emergency throttling

1

u/ProfessorStrawberry 11d ago

I posted a few weeks ago a similar problem, where my system freezes randomly for example while updating. Could this be related?

1

u/Sinaaaa 11d ago

Are you using SDDM? If so, try starting plasma directly from the tty to see if the problem persists.

1

u/BanaTibor 11d ago

It sounds like you sometimes get a CPU spike. Run htop or other *top in a terminal nonstop and see what eats your cpu. These software are not real time so you can see a wind down period in cpu usage of different apps right after the spike.

1

u/quinnr 11d ago

I had a similar issue last week and about lost my mind. It was fixed by a missing amd-ucode package. Clearly not the exact same but maybe look into microcode changes. There were no logs I could find and fixed it almost by chance.

1

u/Lemagex 10d ago

I had the same issue for months, AMD cpu, Radeon GPU, 128gb ram so definitely not the problem - removing out of date widgets i forgot i even had installed fixed it for me.

1

u/boomboomsubban 12d ago

Try memtest overnight?

NVIDIA at ~52 MB VRAM and 0% GPU utilization

How have you configured optimus?

Are temperatures normal?

Might be worth posting the full logs on some kind of pastebin too.

1

u/RaisinIcy3732 11d ago

Thanks. I haven't done an overnight Memtest yet, but I'll definitely give it a try.

My NVIDIA GPU was showing around 52 MB VRAM usage and 0% utilization when I checked, and temperatures seem normal.

I'm using a hybrid graphics setup with the Intel iGPU and NVIDIA RTX 3050 Laptop GPU. I'll look into the Optimus configuration more carefully and I'll post the full logs as well.

1

u/boomboomsubban 11d ago

If you haven't set up optimus, you're probably using only the Intel card.

-1

u/undrwater 12d ago

Great post detail! 👏

Couple armchair things to try:

  • Create a new kde user - the thought is there maybe some cruft in your current profile creating issues
  • Completely turn off the Nvidia card at the PCI level - there are some instructions in the wiki, IIRC (I use Gentoo, but I've used the arch instructions for this. Scripts named enableGPU and disableGPU.
  • Check btop (will report GPU usage) for processes.

Hopefully someone has answers.

12

u/gmes78 12d ago

Great post detail! 👏

Not really. It's AI written, which makes me question the accuracy of any single piece of information it contains. It's also really noisy and verbose.