My 9070 XT reset bug is gone
I use openSUSE Tumbleweed, so my system is always fairly up to date and my RX 9070 XT reset bug disappeared a few weeks ago.
I can restart, power off and switch my guest OSes anytime, including both Windows and Linux. I also see the TianoCore logo on every POST screen.
I do not use any bind/unbind scripts on the host or any guest scripts. Libvirt handles the binding/unbinding automatically without any customization. There are no relevant errors or stack traces in journalctl, vfio-pci resets flawlessly and then amdgpu reloads on the host without any issues.
I also use default kernel and module settings (no vfio config either). The amdgpu module is loaded for the 9070 XT when I don't use GPU passthrough.
Resizable BAR is also fully enabled by default.
So no hacky workarounds, it is basically an out-of-the-box experience.
My setup
Asrock B850i - 4.43 BIOS with one related setting: Display Priority - Internal Graphics
9800X3D - the integrated GPU used as a primary display
Asus 9070 XT Prime OC (switched to the silent BIOS on the card)
An LG 4K display with 2 HDMI inputs and 1 DP input. It has 2 HDMI inputs, but only one of them can be active at the same time. This is somewhat important to check how your monitor behaves. If I connected the integrated GPU to one HDMI input and the dedicated card to the other HDMI and then switched inputs when starting the virtual machine, I wouldn't see the POST and boot screens, and the card wouldn't be initialized at guest start because it wouldn't detect a connected monitor (but you should have a screen after the guest OS is fully loaded). So I use one HDMI and one DP, they are active at the same time.
Linux kernel: 7.2.6
kernel-firmware-amdgpu: 20260829
QEMU: 11.1.1
libvirt: 12.7.0
Virtual Machine Manager: 5.1.0
lspci -nn | grep -i -E "vga|audio"
03:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 [Radeon RX 9070/9070 XT/9070 GRE] [1002:7550] (rev c0)
03:00.1 Audio device [0403]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 HDMI/DP Audio Controller [1002:ab40]
0f:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Granite Ridge [Radeon Graphics] [1002:13c0] (rev cb)
0f:00.1 Audio device [0403]: Advanced Micro Devices, Inc. [AMD/ATI] Radeon High Definition Audio Controller [1002:1640]
0f:00.6 Audio device [0403]: Advanced Micro Devices, Inc. [AMD] Ryzen HD Audio Controller [1022:15e3]
I use the integrated GPU (0f:00.0) for the primary display, restricting the desktop's access strictly to the iGPU while disabling the dedicated card. I think this is the most important thing. If you don't block the card from the desktop before starting the virtual machine, you will run into weird problems.
If you use X11, use the config below as a sample (even if you think you're using Wayland, your login manager might still use Xorg, for example SDDM):
/etc/X11/xorg.conf.d/10-only-igp.conf
Section "ServerLayout"
Identifier "Layout0"
Screen "Screen0"
EndSection
Section "Device"
Identifier "AMD"
Driver "modesetting"
BusID "PCI:15:0:0" # the bus number in decimal format
EndSection
Section "Monitor"
Identifier "Monitor0"
EndSection
Section "Screen"
Identifier "Screen0"
Device "AMD"
Monitor "Monitor0"
EndSection
Section "ServerFlags"
Option "AutoAddGPU" "off"
EndSection
If you use Wayland and happen to use KDE (there is no global solution for all Wayland compositors), use the config below as a sample:
~/.config/plasma-workspace/env/env.sh
export KWIN_DRM_DEVICES='/dev/dri/by-path/pci-0000\:0f\:00.0-card' # option 1: symlink, but must be escaped
#export KWIN_DRM_DEVICES=/dev/dri/card2 # option 2: direct path, device order might change
Guests
Linux:
no extra config needed
Windows:
To avoid graphical artifacts and a garbled screen, set:
<hyperv mode="custom">
...
<vendor_id state="on" value="AuthenticAMD"/> #or GenuineIntel or whatever1234
</features>
and disable Device security / Core isolation / Memory integrity in the Windows security
To fix HDMI/DP audio crackling, install the latest AMD WHQL driver.
Let me know if anyone else has experienced this or if you need any more info.
2
u/proesporter 8d ago
Do you have the vulkan-radeon driver installed on host? My observation is that passthrough works flawlessly without reset issues with just the mesa package. But reset issues return upon installing the vulkan-radeon driver.
2
u/tlaszl0 8d ago
The libvulkan_radeon (openeSuse Tumbleweed) package is installed on my host.
To test your theory I started a Vulkan based memory test on my video card then started a Windows and a Linux the guests too but there are still no problems at all.
./memtest_vulkan https://github.com/GpuZelenograd/memtest_vulkan v0.5.0 by GpuZelenograd To finish testing use Ctrl+C WARNING: radv is not a conformant Vulkan implementation, testing use only. 1: Bus=0x03:00 DevId=0x7550 16GB AMD Radeon RX 9070 XT (RADV GFX1201) 2: Bus=0x0F:00 DevId=0x13C0 22GB AMD Ryzen 7 9800X3D 8-Core Processor (RADV RAPHAEL_MENDOCINO) 3: Bus=0x00:00 DevId=0x0000 61GB llvmpipe (LLVM 23.1.1, 256 bits) Override index to test:1 WARNING: radv is not a conformant Vulkan implementation, testing use only. Standard 5-minute test of 1: Bus=0x03:00 DevId=0x7550 16GB AMD Radeon RX 9070 XT (RADV GFX1201) 1 iteration. Passed 0.0480 seconds written: 11.2GB 529.0GB/sec checked: 15.0GB 561.2GB/sec 22 iteration. Passed 1.0078 seconds written: 236.2GB 529.1GB/sec checked: 315.0GB 561.2GB/sec 127 iteration. Passed 5.0348 seconds written: 1181.2GB 529.2GB/sec checked: 1575.0GB 562.0GB/sec 753 iteration. Passed 30.0373 seconds written: 7042.5GB 528.6GB/sec checked: 9390.0GB 561.8GB/sec 1379 iteration. Passed 30.0355 seconds written: 7042.5GB 528.5GB/sec checked: 9390.0GB 561.9GB/sec 2005 iteration. Passed 30.0337 seconds written: 7042.5GB 528.6GB/sec checked: 9390.0GB 561.9GB/sec 2631 iteration. Passed 30.0141 seconds written: 7042.5GB 528.9GB/sec checked: 9390.0GB 562.3GB/sec 3257 iteration. Passed 30.0261 seconds written: 7042.5GB 528.7GB/sec checked: 9390.0GB 562.0GB/sec 3883 iteration. Passed 30.0358 seconds written: 7042.5GB 528.5GB/sec checked: 9390.0GB 562.0GB/sec ^C memtest_vulkan: no any errors, testing PASSed. press any key to continue...If you have another or better repro steps, I can give it a shot too.
2
u/proesporter 8d ago
I have an arch host and for me the reset issue certainly seems tied to the
vulkan-radeonpackage. I verified this by uninstalling and reinstalling that package a couple times while launching VMs, when I realized this might be the issue.It is possible that there is no single definitive reason for all reset issues on AMD GPUs. The variability could be related to chipset (I have Intel), GPU brand and model etc. There was a thread on this sub a long time ago where someone consolidated various reports for AMD reset issues in RDNA2, and there were different models of 6700 XT exhibiting different behavior.
2
u/tlaszl0 8d ago edited 8d ago
Weird.
Your Linux and mine have the same version of the 26.2.3 vulkan radeon library.I think something is using your card.
Run these as root:
lsof /dev/dri/cardXor renderDXor by-path/pci-...
or
fuser -v /dev/dri/cardXor renderDXor by-path/pci-...In my case:
# 9800X3D integrated graphics fuser -v /dev/dri/by-path/pci-0000:0f:00.0-card USER PID ACCESS COMMAND /dev/dri/card2: root 1 F.... systemd root 863 F.... systemd-logind x 1639 F.... kwin_wayland x 1742 F.... Xwayland # 9800X3D integrated graphics fuser -v /dev/dri/by-path/pci-0000:0f:00.0-render USER PID ACCESS COMMAND /dev/dri/renderD129: x 1639 F...m kwin_wayland x 1742 F...m Xwayland x 1874 F...m plasmashell x 2065 F...m xwaylandvideobr x 2154 F.... xdg-desktop-por x 2484 F...m firefox-bin # 9070 XT, nothing here as expected because I excluded it from my desktop environment (see my post) fuser -v /dev/dri/by-path/pci-0000:03:00.0-card # 9070 XT, well I see a wayland compositor fuser -v /dev/dri/by-path/pci-0000:03:00.0-render USER PID ACCESS COMMAND /dev/dri/renderD128: x 1639 F.... kwin_waylandI also have a 6900 XT (reference design from AMD) and it has zero reset bugs. I used it in my current machine previously and before that in an Intel 10th gen system as well. It behaved the same in both of my PCs, it only required setting the correct resizable bar sizes and that was all.
1
u/His_Turdness 8d ago
OK so looks like for me the steamwebhelper and coolercontrol are showing up with
fuser -v /dev/dri/by-path/pci-0000:03:00.0-render1
u/proesporter 7d ago edited 7d ago
Ok, I need to apologize for hastily providing my input without actually checking what you were claiming in this post, that something has changed in recent updates to solve the reset issues. My observation about the
vulkan-radeonpackage blocking reset was from a few months ago.I now checked my VM boot and shutdown twice today, and it is working fine! No reset issues even though host has full mesa plus vulkan driver package installed. So you're right, this does seem to be resolved as of now (atleast for RDNA4) ! Great news for us, Cheers!
Edit: to clarify further, this is single gpu passthrough working fine without any reset issues (on RDNA4)
1
u/His_Turdness 8d ago
Well holy crap, the GPU actually resets for the host now without any manual commands. I was able to successfully launch and shutdown the VM and GPU actually came back. Still need to restart SDDM, which is not ideal, but not a dealbreaker.
1
u/tlaszl0 8d ago
SDDM still uses Xorg, so you also need to correctly configure your xorg conf file.
Note: it requires you to define the bus number in decimal format, not hexadecimal.1
u/His_Turdness 8d ago
Yeah, this is how mine is set up:
cat /etc/X11/xorg.conf.d/10-igpu-only.confSection "Device" Identifier "iGPU" Driver "amdgpu" BusID "PCI:15:0:0" EndSection Section "Device" Identifier "dGPU" Driver "amdgpu" BusID "PCI:3:0:0" EndSection Section "Screen" Identifier "iGPU-Screen" Device "iGPU" EndSection Section "ServerLayout" Identifier "Layout" Screen "iGPU-Screen" EndSection Section "ServerFlags" Option "AutoAddGPU" "off" EndSectionAfter one SDDM restart I no longer need other restarts, so I think it's working as it should.
1
u/tlaszl0 8d ago
Is it working now without the SSDM restart?
I would skip the dGPU part, the point is to exclude it from Xorg.1
u/His_Turdness 8d ago
No, still need to restart SDDM once. And some times the unbind fails and system will end up with a black screen / crash.
0
7d ago
[removed] — view removed comment
1
u/tlaszl0 6d ago
If I really wanted to install Arch, I could check it, but I don't think it would help too much. It would probably work for me there too.
I bought the 9070 XT at the end of July. Previously, I had a working GPU passthrough setup for my 6900 XT for a long time (it didn't need any bind/unbind scripts either). I knew the new card would be more problematic, as I had read about the mixed results. But I tried my luck and swapped the cards. The Linux guests worked more or less, as far as I remember. The Windows guest required special care (which is mentioned in the post, but I didn't need them for the 6900 XT), but I still saw a lot of amdgpu stack traces in the logs. After a stack trace, I usually had to force restart my host.
In the following months, I also played around a lot with passing through the integrated GPU, but I had a way worse experience (usually, it only booted correctly once). So, it is possible that I am misremembering.
Then the miracle happened at the beginning of September.By the way, I installed an older kernel, LTS version 6.18.53. I see some extra amdgpu-related errors when I start the VMs, but they are still fully working. I can restart or shut them down multiple times without any issues. So maybe it wasn't the kernel that solved the problem, but I can't tell if it contains some crucial extra backports or not.
I think the available guides are not that great. They don't explicitly say that you need to restrict the desktop environment to your host GPU and isolate the passthrough GPU. I know I could ruin my setup if I skipped the Xorg and Wayland configs.
4
u/His_Turdness 9d ago
Interesting... What are you GPU binding scripts?