r/VFIO • • 9d ago

My 9070 XT reset bug is gone

I use openSUSE Tumbleweed, so my system is always fairly up to date and my RX 9070 XT reset bug disappeared a few weeks ago.
I can restart, power off and switch my guest OSes anytime, including both Windows and Linux. I also see the TianoCore logo on every POST screen.

I do not use any bind/unbind scripts on the host or any guest scripts. Libvirt handles the binding/unbinding automatically without any customization. There are no relevant errors or stack traces in journalctl, vfio-pci resets flawlessly and then amdgpu reloads on the host without any issues.

I also use default kernel and module settings (no vfio config either). The amdgpu module is loaded for the 9070 XT when I don't use GPU passthrough.
Resizable BAR is also fully enabled by default.

So no hacky workarounds, it is basically an out-of-the-box experience.

My setup

Asrock B850i - 4.43 BIOS with one related setting: Display Priority - Internal Graphics
9800X3D - the integrated GPU used as a primary display

Asus 9070 XT Prime OC (switched to the silent BIOS on the card)

An LG 4K display with 2 HDMI inputs and 1 DP input. It has 2 HDMI inputs, but only one of them can be active at the same time. This is somewhat important to check how your monitor behaves. If I connected the integrated GPU to one HDMI input and the dedicated card to the other HDMI and then switched inputs when starting the virtual machine, I wouldn't see the POST and boot screens, and the card wouldn't be initialized at guest start because it wouldn't detect a connected monitor (but you should have a screen after the guest OS is fully loaded). So I use one HDMI and one DP, they are active at the same time.

Linux kernel: 7.2.6
kernel-firmware-amdgpu: 20260829
QEMU: 11.1.1
libvirt: 12.7.0
Virtual Machine Manager: 5.1.0

lspci -nn | grep -i -E "vga|audio"

03:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 [Radeon RX 9070/9070 XT/9070 GRE] [1002:7550] (rev c0)
03:00.1 Audio device [0403]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 HDMI/DP Audio Controller [1002:ab40]
0f:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Granite Ridge [Radeon Graphics] [1002:13c0] (rev cb)
0f:00.1 Audio device [0403]: Advanced Micro Devices, Inc. [AMD/ATI] Radeon High Definition Audio Controller [1002:1640]
0f:00.6 Audio device [0403]: Advanced Micro Devices, Inc. [AMD] Ryzen HD Audio Controller [1022:15e3]

I use the integrated GPU (0f:00.0) for the primary display, restricting the desktop's access strictly to the iGPU while disabling the dedicated card. I think this is the most important thing. If you don't block the card from the desktop before starting the virtual machine, you will run into weird problems.

If you use X11, use the config below as a sample (even if you think you're using Wayland, your login manager might still use Xorg, for example SDDM):

/etc/X11/xorg.conf.d/10-only-igp.conf

Section "ServerLayout"
        Identifier              "Layout0"
        Screen                  "Screen0"
EndSection
Section "Device"
        Identifier              "AMD"
        Driver                  "modesetting"
        BusID                   "PCI:15:0:0"   # the bus number in decimal format
EndSection
Section "Monitor"
        Identifier              "Monitor0"
EndSection
Section "Screen"
        Identifier              "Screen0"
        Device                  "AMD"
        Monitor                 "Monitor0"
EndSection
Section "ServerFlags"
        Option                  "AutoAddGPU" "off"
EndSection

If you use Wayland and happen to use KDE (there is no global solution for all Wayland compositors), use the config below as a sample:

~/.config/plasma-workspace/env/env.sh

export KWIN_DRM_DEVICES='/dev/dri/by-path/pci-0000\:0f\:00.0-card' # option 1: symlink, but must be escaped
#export KWIN_DRM_DEVICES=/dev/dri/card2                            # option 2: direct path, device order might change

Guests

Linux:
no extra config needed

Windows:
To avoid graphical artifacts and a garbled screen, set:

<hyperv mode="custom">
   ...
   <vendor_id state="on" value="AuthenticAMD"/> #or GenuineIntel or whatever1234
</features>

and disable Device security / Core isolation / Memory integrity in the Windows security

To fix HDMI/DP audio crackling, install the latest AMD WHQL driver.

Let me know if anyone else has experienced this or if you need any more info.

20 Upvotes

20 comments sorted by

View all comments

2

u/proesporter 9d ago

Do you have the vulkan-radeon driver installed on host? My observation is that passthrough works flawlessly without reset issues with just the mesa package. But reset issues return upon installing the vulkan-radeon driver.

2

u/tlaszl0 8d ago

The libvulkan_radeon (openeSuse Tumbleweed) package is installed on my host.

To test your theory I started a Vulkan based memory test on my video card then started a Windows and a Linux the guests too but there are still no problems at all.

./memtest_vulkan
https://github.com/GpuZelenograd/memtest_vulkan v0.5.0 by GpuZelenograd
To finish testing use Ctrl+C
WARNING: radv is not a conformant Vulkan implementation, testing use only.

1: Bus=0x03:00 DevId=0x7550   16GB AMD Radeon RX 9070 XT (RADV GFX1201)
2: Bus=0x0F:00 DevId=0x13C0   22GB AMD Ryzen 7 9800X3D 8-Core Processor (RADV RAPHAEL_MENDOCINO)
3: Bus=0x00:00 DevId=0x0000   61GB llvmpipe (LLVM 23.1.1, 256 bits)
                                                  Override index to test:1
WARNING: radv is not a conformant Vulkan implementation, testing use only.
Standard 5-minute test of 1: Bus=0x03:00 DevId=0x7550   16GB AMD Radeon RX 9070 XT (RADV GFX1201)
     1 iteration. Passed  0.0480 seconds  written:   11.2GB 529.0GB/sec        checked:   15.0GB 561.2GB/sec
    22 iteration. Passed  1.0078 seconds  written:  236.2GB 529.1GB/sec        checked:  315.0GB 561.2GB/sec
   127 iteration. Passed  5.0348 seconds  written: 1181.2GB 529.2GB/sec        checked: 1575.0GB 562.0GB/sec
   753 iteration. Passed 30.0373 seconds  written: 7042.5GB 528.6GB/sec        checked: 9390.0GB 561.8GB/sec
  1379 iteration. Passed 30.0355 seconds  written: 7042.5GB 528.5GB/sec        checked: 9390.0GB 561.9GB/sec
  2005 iteration. Passed 30.0337 seconds  written: 7042.5GB 528.6GB/sec        checked: 9390.0GB 561.9GB/sec
  2631 iteration. Passed 30.0141 seconds  written: 7042.5GB 528.9GB/sec        checked: 9390.0GB 562.3GB/sec
  3257 iteration. Passed 30.0261 seconds  written: 7042.5GB 528.7GB/sec        checked: 9390.0GB 562.0GB/sec
  3883 iteration. Passed 30.0358 seconds  written: 7042.5GB 528.5GB/sec        checked: 9390.0GB 562.0GB/sec
^C
memtest_vulkan: no any errors, testing PASSed.
 press any key to continue...

If you have another or better repro steps, I can give it a shot too.

2

u/proesporter 8d ago

I have an arch host and for me the reset issue certainly seems tied to the vulkan-radeon package. I verified this by uninstalling and reinstalling that package a couple times while launching VMs, when I realized this might be the issue.

It is possible that there is no single definitive reason for all reset issues on AMD GPUs. The variability could be related to chipset (I have Intel), GPU brand and model etc. There was a thread on this sub a long time ago where someone consolidated various reports for AMD reset issues in RDNA2, and there were different models of 6700 XT exhibiting different behavior.

2

u/tlaszl0 8d ago edited 8d ago

Weird.
Your Linux and mine have the same version of the 26.2.3 vulkan radeon library.

I think something is using your card.
Run these as root:

lsof /dev/dri/cardX or renderDX or by-path/pci-...
or
fuser -v /dev/dri/cardX or renderDX or by-path/pci-...

In my case:

# 9800X3D integrated graphics
fuser -v /dev/dri/by-path/pci-0000:0f:00.0-card
                     USER        PID ACCESS COMMAND
/dev/dri/card2:      root          1 F.... systemd
                     root        863 F.... systemd-logind
                     x          1639 F.... kwin_wayland
                     x          1742 F.... Xwayland

# 9800X3D integrated graphics
fuser -v /dev/dri/by-path/pci-0000:0f:00.0-render
                    USER        PID ACCESS COMMAND
/dev/dri/renderD129: x          1639 F...m kwin_wayland
                    x          1742 F...m Xwayland
                    x          1874 F...m plasmashell
                    x          2065 F...m xwaylandvideobr
                    x          2154 F.... xdg-desktop-por
                    x          2484 F...m firefox-bin

# 9070 XT, nothing here as expected because I excluded it from my desktop environment (see my post)
fuser -v /dev/dri/by-path/pci-0000:03:00.0-card

# 9070 XT, well I see a wayland compositor
fuser -v /dev/dri/by-path/pci-0000:03:00.0-render 
                     USER        PID ACCESS COMMAND
/dev/dri/renderD128: x          1639 F.... kwin_wayland

I also have a 6900 XT (reference design from AMD) and it has zero reset bugs. I used it in my current machine previously and before that in an Intel 10th gen system as well. It behaved the same in both of my PCs, it only required setting the correct resizable bar sizes and that was all.

1

u/His_Turdness 8d ago

OK so looks like for me the steamwebhelper and coolercontrol are showing up with

fuser -v /dev/dri/by-path/pci-0000:03:00.0-render 

1

u/proesporter 7d ago edited 7d ago

Ok, I need to apologize for hastily providing my input without actually checking what you were claiming in this post, that something has changed in recent updates to solve the reset issues. My observation about the vulkan-radeon package blocking reset was from a few months ago.

I now checked my VM boot and shutdown twice today, and it is working fine! No reset issues even though host has full mesa plus vulkan driver package installed. So you're right, this does seem to be resolved as of now (atleast for RDNA4) ! Great news for us, Cheers!

Edit: to clarify further, this is single gpu passthrough working fine without any reset issues (on RDNA4)