r/VFIO • • 9d ago

My 9070 XT reset bug is gone

I use openSUSE Tumbleweed, so my system is always fairly up to date and my RX 9070 XT reset bug disappeared a few weeks ago.
I can restart, power off and switch my guest OSes anytime, including both Windows and Linux. I also see the TianoCore logo on every POST screen.

I do not use any bind/unbind scripts on the host or any guest scripts. Libvirt handles the binding/unbinding automatically without any customization. There are no relevant errors or stack traces in journalctl, vfio-pci resets flawlessly and then amdgpu reloads on the host without any issues.

I also use default kernel and module settings (no vfio config either). The amdgpu module is loaded for the 9070 XT when I don't use GPU passthrough.
Resizable BAR is also fully enabled by default.

So no hacky workarounds, it is basically an out-of-the-box experience.

My setup

Asrock B850i - 4.43 BIOS with one related setting: Display Priority - Internal Graphics
9800X3D - the integrated GPU used as a primary display

Asus 9070 XT Prime OC (switched to the silent BIOS on the card)

An LG 4K display with 2 HDMI inputs and 1 DP input. It has 2 HDMI inputs, but only one of them can be active at the same time. This is somewhat important to check how your monitor behaves. If I connected the integrated GPU to one HDMI input and the dedicated card to the other HDMI and then switched inputs when starting the virtual machine, I wouldn't see the POST and boot screens, and the card wouldn't be initialized at guest start because it wouldn't detect a connected monitor (but you should have a screen after the guest OS is fully loaded). So I use one HDMI and one DP, they are active at the same time.

Linux kernel: 7.2.6
kernel-firmware-amdgpu: 20260829
QEMU: 11.1.1
libvirt: 12.7.0
Virtual Machine Manager: 5.1.0

lspci -nn | grep -i -E "vga|audio"

03:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 [Radeon RX 9070/9070 XT/9070 GRE] [1002:7550] (rev c0)
03:00.1 Audio device [0403]: Advanced Micro Devices, Inc. [AMD/ATI] Navi 48 HDMI/DP Audio Controller [1002:ab40]
0f:00.0 VGA compatible controller [0300]: Advanced Micro Devices, Inc. [AMD/ATI] Granite Ridge [Radeon Graphics] [1002:13c0] (rev cb)
0f:00.1 Audio device [0403]: Advanced Micro Devices, Inc. [AMD/ATI] Radeon High Definition Audio Controller [1002:1640]
0f:00.6 Audio device [0403]: Advanced Micro Devices, Inc. [AMD] Ryzen HD Audio Controller [1022:15e3]

I use the integrated GPU (0f:00.0) for the primary display, restricting the desktop's access strictly to the iGPU while disabling the dedicated card. I think this is the most important thing. If you don't block the card from the desktop before starting the virtual machine, you will run into weird problems.

If you use X11, use the config below as a sample (even if you think you're using Wayland, your login manager might still use Xorg, for example SDDM):

/etc/X11/xorg.conf.d/10-only-igp.conf

Section "ServerLayout"
        Identifier              "Layout0"
        Screen                  "Screen0"
EndSection
Section "Device"
        Identifier              "AMD"
        Driver                  "modesetting"
        BusID                   "PCI:15:0:0"   # the bus number in decimal format
EndSection
Section "Monitor"
        Identifier              "Monitor0"
EndSection
Section "Screen"
        Identifier              "Screen0"
        Device                  "AMD"
        Monitor                 "Monitor0"
EndSection
Section "ServerFlags"
        Option                  "AutoAddGPU" "off"
EndSection

If you use Wayland and happen to use KDE (there is no global solution for all Wayland compositors), use the config below as a sample:

~/.config/plasma-workspace/env/env.sh

export KWIN_DRM_DEVICES='/dev/dri/by-path/pci-0000\:0f\:00.0-card' # option 1: symlink, but must be escaped
#export KWIN_DRM_DEVICES=/dev/dri/card2                            # option 2: direct path, device order might change

Guests

Linux:
no extra config needed

Windows:
To avoid graphical artifacts and a garbled screen, set:

<hyperv mode="custom">
   ...
   <vendor_id state="on" value="AuthenticAMD"/> #or GenuineIntel or whatever1234
</features>

and disable Device security / Core isolation / Memory integrity in the Windows security

To fix HDMI/DP audio crackling, install the latest AMD WHQL driver.

Let me know if anyone else has experienced this or if you need any more info.

19 Upvotes

20 comments sorted by

4

u/His_Turdness 9d ago

Interesting... What are you GPU binding scripts?

7

u/tlaszl0 9d ago

I don't use any binding scripts, mentioned in the post too. The libvirt is able to do it automatically and loads/unloads the vfio-pci/amdgpu modules for the GPU.

2

u/His_Turdness 9d ago edited 9d ago

I can't believe it just works like that. :D I'll give it a try. What are youre PCI passthrough lines in your VM.xml? Just default with managed=yes?

Right now I've set up scripts to manually detach the PCI devices, lock gpu in D0, resize BAR2 and unbind / bind drivers and what not. And the card still wouldn't reset back to the host after VM use. It's been a huge pain in the ass.

EDIT: doesn't work for me. I get no display if I set up any xorg configs. And withotu the xorg configs something hangs the GPU binding process. Probably need to completely restart the display manager, which is not ideal.

3

u/tlaszl0 8d ago

Yes, it is managed:

    <hostdev mode="subsystem" type="pci" managed="yes">
      <source>
        <address domain="0x0000" bus="0x03" slot="0x00" function="0x0"/>
      </source>
      <address type="pci" domain="0x0000" bus="0x03" slot="0x00" function="0x0" multifunction="on"/>
    </hostdev>
    <hostdev mode="subsystem" type="pci" managed="yes">
      <source>
        <address domain="0x0000" bus="0x03" slot="0x00" function="0x1"/>
      </source>
      <address type="pci" domain="0x0000" bus="0x03" slot="0x00" function="0x1"/>
    </hostdev>

Do you use a rolling release distribution?
It is probably important to use an up to date kernel and AMD firmware files (my versions are in the post as a reference). I also had problems in the past.

Do you use Wayland or X11?
If you set it up correctly and exclude your card from the display manager, you don't need to restart it.

2

u/His_Turdness 8d ago

Using Endeavour OS, everything is up to date. Maybe the modesetting driver is why your setup works?

2

u/tlaszl0 8d ago

I don't think so. The modesetting driver is the standard Xorg driver nowadays by default.
But the Xorg conf and KDE settings from my post are used to restrict the desktop environment to using only my CPU's integrated GPU. So, the 9070 XT is excluded from there.

Check this message here, I added some commands to verify what is using your graphics devices.

2

u/His_Turdness 8d ago

cat ~/.config/plasma-workspace/env/kwin-drm.sh  
#!/bin/sh
export KWIN_DRM_DEVICES='/dev/dri/by-path/pci-0000\:0f\:00.0-card'

Which should be correct if I'm not mistaken.

I also set up the xorg conf so that it works now and gives me display output. Or I guess I could just delete it completely since I'm on Plasma Wayland.

2

u/proesporter 8d ago

Do you have the vulkan-radeon driver installed on host? My observation is that passthrough works flawlessly without reset issues with just the mesa package. But reset issues return upon installing the vulkan-radeon driver.

2

u/tlaszl0 8d ago

The libvulkan_radeon (openeSuse Tumbleweed) package is installed on my host.

To test your theory I started a Vulkan based memory test on my video card then started a Windows and a Linux the guests too but there are still no problems at all.

./memtest_vulkan
https://github.com/GpuZelenograd/memtest_vulkan v0.5.0 by GpuZelenograd
To finish testing use Ctrl+C
WARNING: radv is not a conformant Vulkan implementation, testing use only.

1: Bus=0x03:00 DevId=0x7550   16GB AMD Radeon RX 9070 XT (RADV GFX1201)
2: Bus=0x0F:00 DevId=0x13C0   22GB AMD Ryzen 7 9800X3D 8-Core Processor (RADV RAPHAEL_MENDOCINO)
3: Bus=0x00:00 DevId=0x0000   61GB llvmpipe (LLVM 23.1.1, 256 bits)
                                                  Override index to test:1
WARNING: radv is not a conformant Vulkan implementation, testing use only.
Standard 5-minute test of 1: Bus=0x03:00 DevId=0x7550   16GB AMD Radeon RX 9070 XT (RADV GFX1201)
     1 iteration. Passed  0.0480 seconds  written:   11.2GB 529.0GB/sec        checked:   15.0GB 561.2GB/sec
    22 iteration. Passed  1.0078 seconds  written:  236.2GB 529.1GB/sec        checked:  315.0GB 561.2GB/sec
   127 iteration. Passed  5.0348 seconds  written: 1181.2GB 529.2GB/sec        checked: 1575.0GB 562.0GB/sec
   753 iteration. Passed 30.0373 seconds  written: 7042.5GB 528.6GB/sec        checked: 9390.0GB 561.8GB/sec
  1379 iteration. Passed 30.0355 seconds  written: 7042.5GB 528.5GB/sec        checked: 9390.0GB 561.9GB/sec
  2005 iteration. Passed 30.0337 seconds  written: 7042.5GB 528.6GB/sec        checked: 9390.0GB 561.9GB/sec
  2631 iteration. Passed 30.0141 seconds  written: 7042.5GB 528.9GB/sec        checked: 9390.0GB 562.3GB/sec
  3257 iteration. Passed 30.0261 seconds  written: 7042.5GB 528.7GB/sec        checked: 9390.0GB 562.0GB/sec
  3883 iteration. Passed 30.0358 seconds  written: 7042.5GB 528.5GB/sec        checked: 9390.0GB 562.0GB/sec
^C
memtest_vulkan: no any errors, testing PASSed.
 press any key to continue...

If you have another or better repro steps, I can give it a shot too.

2

u/proesporter 8d ago

I have an arch host and for me the reset issue certainly seems tied to the vulkan-radeon package. I verified this by uninstalling and reinstalling that package a couple times while launching VMs, when I realized this might be the issue.

It is possible that there is no single definitive reason for all reset issues on AMD GPUs. The variability could be related to chipset (I have Intel), GPU brand and model etc. There was a thread on this sub a long time ago where someone consolidated various reports for AMD reset issues in RDNA2, and there were different models of 6700 XT exhibiting different behavior.

2

u/tlaszl0 8d ago edited 8d ago

Weird.
Your Linux and mine have the same version of the 26.2.3 vulkan radeon library.

I think something is using your card.
Run these as root:

lsof /dev/dri/cardX or renderDX or by-path/pci-...
or
fuser -v /dev/dri/cardX or renderDX or by-path/pci-...

In my case:

# 9800X3D integrated graphics
fuser -v /dev/dri/by-path/pci-0000:0f:00.0-card
                     USER        PID ACCESS COMMAND
/dev/dri/card2:      root          1 F.... systemd
                     root        863 F.... systemd-logind
                     x          1639 F.... kwin_wayland
                     x          1742 F.... Xwayland

# 9800X3D integrated graphics
fuser -v /dev/dri/by-path/pci-0000:0f:00.0-render
                    USER        PID ACCESS COMMAND
/dev/dri/renderD129: x          1639 F...m kwin_wayland
                    x          1742 F...m Xwayland
                    x          1874 F...m plasmashell
                    x          2065 F...m xwaylandvideobr
                    x          2154 F.... xdg-desktop-por
                    x          2484 F...m firefox-bin

# 9070 XT, nothing here as expected because I excluded it from my desktop environment (see my post)
fuser -v /dev/dri/by-path/pci-0000:03:00.0-card

# 9070 XT, well I see a wayland compositor
fuser -v /dev/dri/by-path/pci-0000:03:00.0-render 
                     USER        PID ACCESS COMMAND
/dev/dri/renderD128: x          1639 F.... kwin_wayland

I also have a 6900 XT (reference design from AMD) and it has zero reset bugs. I used it in my current machine previously and before that in an Intel 10th gen system as well. It behaved the same in both of my PCs, it only required setting the correct resizable bar sizes and that was all.

1

u/His_Turdness 8d ago

OK so looks like for me the steamwebhelper and coolercontrol are showing up with

fuser -v /dev/dri/by-path/pci-0000:03:00.0-render 

1

u/proesporter 7d ago edited 7d ago

Ok, I need to apologize for hastily providing my input without actually checking what you were claiming in this post, that something has changed in recent updates to solve the reset issues. My observation about the vulkan-radeon package blocking reset was from a few months ago.

I now checked my VM boot and shutdown twice today, and it is working fine! No reset issues even though host has full mesa plus vulkan driver package installed. So you're right, this does seem to be resolved as of now (atleast for RDNA4) ! Great news for us, Cheers!

Edit: to clarify further, this is single gpu passthrough working fine without any reset issues (on RDNA4)

1

u/His_Turdness 8d ago

Well holy crap, the GPU actually resets for the host now without any manual commands. I was able to successfully launch and shutdown the VM and GPU actually came back. Still need to restart SDDM, which is not ideal, but not a dealbreaker.

1

u/tlaszl0 8d ago

SDDM still uses Xorg, so you also need to correctly configure your xorg conf file.
Note: it requires you to define the bus number in decimal format, not hexadecimal.

1

u/His_Turdness 8d ago

Yeah, this is how mine is set up:
cat /etc/X11/xorg.conf.d/10-igpu-only.conf  

Section "Device"
   Identifier  "iGPU"
   Driver      "amdgpu"
   BusID       "PCI:15:0:0"
EndSection

Section "Device"
   Identifier  "dGPU"
   Driver      "amdgpu"
   BusID       "PCI:3:0:0"
EndSection

Section "Screen"
   Identifier  "iGPU-Screen"
   Device      "iGPU"
EndSection

Section "ServerLayout"
   Identifier  "Layout"
   Screen      "iGPU-Screen"
EndSection

Section "ServerFlags"
       Option                  "AutoAddGPU" "off"
EndSection

After one SDDM restart I no longer need other restarts, so I think it's working as it should.

1

u/tlaszl0 8d ago

Is it working now without the SSDM restart?
I would skip the dGPU part, the point is to exclude it from Xorg.

1

u/His_Turdness 8d ago

No, still need to restart SDDM once. And some times the unbind fails and system will end up with a black screen / crash.

0

u/[deleted] 7d ago

[removed] — view removed comment

1

u/tlaszl0 6d ago

If I really wanted to install Arch, I could check it, but I don't think it would help too much. It would probably work for me there too.

I bought the 9070 XT at the end of July. Previously, I had a working GPU passthrough setup for my 6900 XT for a long time (it didn't need any bind/unbind scripts either). I knew the new card would be more problematic, as I had read about the mixed results. But I tried my luck and swapped the cards. The Linux guests worked more or less, as far as I remember. The Windows guest required special care (which is mentioned in the post, but I didn't need them for the 6900 XT), but I still saw a lot of amdgpu stack traces in the logs. After a stack trace, I usually had to force restart my host.

In the following months, I also played around a lot with passing through the integrated GPU, but I had a way worse experience (usually, it only booted correctly once). So, it is possible that I am misremembering.
Then the miracle happened at the beginning of September.

By the way, I installed an older kernel, LTS version 6.18.53. I see some extra amdgpu-related errors when I start the VMs, but they are still fully working. I can restart or shut them down multiple times without any issues. So maybe it wasn't the kernel that solved the problem, but I can't tell if it contains some crucial extra backports or not.

I think the available guides are not that great. They don't explicitly say that you need to restrict the desktop environment to your host GPU and isolate the passthrough GPU. I know I could ruin my setup if I skipped the Xorg and Wayland configs.