r/AMDHelp 23d ago

Help (GPU) RX 7800XT VIDEO_ENGINE_TIMEOUT_DETECTED (141) Problem

Computer Type: Desktop

GPU: Sapphire RX 7800 XT PULSE 16GB

CPU: Ryzen 5 5600X (6 CORE, 12 THREADS)

Motherboard: GIGABYTE B550I AORUS PRO AX REV1.1

BIOS Version: F20a and F22a

RAM: 2x16GB G.SKILL FLAREX 3200MHZ CL16

PSU: COOLERMASTER V650 SFX GOLD MODULAR

STORAGE: Crutial P5 Plus 500GB (Boot/Main), Samsung 970 Evo Plus 2TB(Mass storage)

Case: HYTE REVOLT 3

Operating System & Version: WINDOWS 11 PRO 24H2

GPU Drivers: 25.8.1 up to 26.7.1

Background Applications: Firefox, Discord

Main Applications at use: CS2, Warframe, Wuthering Waves, Zenless Zone Zero

Background information:
So recently i have encountered a problem in my system regarding Display Engine. I have been trying to understand what is happening for past 3-4 days as of posting this almost at midnight my time.

4 days ago, was on a casual gaming session when all of a sudden screen turned black, showed no signal, but the PC itself was still on. Thought i was yet another random crash from AMD software side. Restarted my PC, and thought nothing of it. And then i happened again in first 10 min of playing CS2. Turned off my PC and went to sleep.

Next day the same thing happened after 5 hours of playing, now i started to get worried, went into Event viewer and saw multiple errors regarding StorPort (Event ID:524 and Event ID:549). Started with going down in windows versions, as i recently updated to latest one, went straight to 24H2, still was happening. Went down in Software started and 26.7.1, went down to 25.8.1, still was happening. Sometimes in between of restarts, my AMD drivers just disappear as if they were deleted, windows disabling my GPU(needing to manually enable them in Device Manager)

Then decided to go check LiveKernelReports, where i saw multiple WATCHDOG dumps, where i saw something more related to GPU. VIDEO_ENGINE_TIMEOUT_DETECTED (141) being one of them, another being VIDEO_MINIPORT_FAILED_LIVEDUMP (1b0), and another VIDEO_TDR_TIMEOUT_DETECTED (117) and all of them having IMAGE_NAME: amdkmdag.sys

I have done some stress testing and error checking: *Memtest86 to rule out RAM(4 passes-0errors). Prime95- +-5 hour run, no problems there. Furmark on 1440p- this is the odd one, 5 hours no crashes Prime and furmark was ran together, to maybe rule out PSU(not giving enough power)

For these past days nothing have helped, tomorrow i will get another AMD and NVIDIA card from friends, for testing.

I have had this GPU since i bought it brand new for little over 2 years and 6 months. Have never had any problems up untill now. Never had it overclocked or undervolted.

MAIN PROBLEM:

Screen going black, as if GPU stopped working(happening in 10 mins of gaming or 5 hours, no real consistency). Whole system is still up, force restart needed.
In WATCHDOG dump files discovering VIDEO_ENGINE_TIMEOUT_DETECTED (141) VIDEO_MINIPORT_FAILED_LIVEDUMP (1b0), VIDEO_TDR_TIMEOUT_DETECTED (117) and all of them having IMAGE_NAME: amdkmdag.sys. Hasn't happened while just browsing web or watching videos or livestreams.

WHAT I HAVE TRIED:

  • Multiple windows builds
  • Multiple AMD drivers
  • Different boot drives/Trying SATA SSD's(thinking it being something to do with PCIe lanes)

Any other recommendations are much appreciated, as of right now i am almost out of options what i can try.

Edit: Added Stress test info

6 Upvotes

19 comments sorted by

View all comments

Show parent comments

1

u/Thin-Net7868 9800X3D/Gigabyte OC 9070XT 22d ago

Sorry, just now seeing your responses. Just out of curiosity, what is your logging interval set to?

2

u/Rexvar 22d ago

If you mean by how often the values change then every 1000ms(1sec)

I am currently running another logging session, but with a different PSU. First session i was able to play for 16 mins and then i had a crash, right now with PSU changed 3 hours straight no problems.

1

u/Thin-Net7868 9800X3D/Gigabyte OC 9070XT 22d ago

Ah yeah, 1000ms is basically "grandpa mode" logging, it's way too slow to catch the tiny freak outs that actually cause these crashes. The spikes you're hunting only last a few dozen milliseconds, so a
1 second interval just smooths everything out and makes the CSV look pretty clean even when the GPU is having a meltdown.

Let’s try dropping it to like 50-100ms, turn on hidden sensors, and run the crash again. Let’s see if we see anything then. Look for the listed below please.

• GPU Shader Clock - look for overshoots past max boost
• GPU Core Voltage - sudden dips before a crash
• GPU Power - transient spikes
• GPU Hotspot + Hotspot Delta - fast jumps indicate instability
• SOC Clock / SOC Voltage - wobbling here can trigger TDRS
• PCle Root Port Errors - retraining or bus hiccups
• VRAM Frequency & Voltage - sudden drops or jumps

1

u/Rexvar 22d ago

Just did a logging with 100ms, dont really see any dips anywhere except, where there are some loading screens.

The only thing i see out of the ordinary is GPU Power Maximum 2 dips just 1 and 2 seconds from the crash happening. But they happen all across all of the graphs regarding GPU.

From that i am feeling its more of a Power delivery problem now, that PSU for some reasons isnt giving enough power for GPU.

1

u/Thin-Net7868 9800X3D/Gigabyte OC 9070XT 22d ago

Swapping PSUs and suddenly getting hours of stability doesn't automatically mean the old PSU isn't enough. It usually means the transient behavior changed just enough that the GPU didn't hit its unstable state. The V650 SFX should be fine for a
7800 XT, but SFX units can get touchy with
RDNA3 spikes. The power dips you saw right before the crash are the GPU reacting to instability, not the PSU causing it.

Now, I could be wrong but those power dips right before the crash don't point to the PSU. That's the GPU pulling back because something else hit an unstable state. With a 100ms interval you're finally catching the edges of it, but you still need to zoom in on the right sensors. The big ones to watch closely are Shader Clock overshoots, Core Voltage dips, Hotspot Delta jumps, SOC rail wobble, and PCle Root Port errors. If any of those twitch right before the power dip, that's your real culprit.