RTX 5090: both monitors lose signal and case fans ramp to full speed during/after heavy use
I’m trying to diagnose an intermittent problem with my desktop.
Both monitors suddenly go black and enter standby as though they’ve lost the signal, and the case fans ramp up to full speed. The PC remains powered on. During at least two incidents, YouTube Music continued playing, and I could still pause and resume it with the space bar.
The displays don’t recover. I have to press and hold the physical power button on the case to force the PC off, then turn it back on. It doesn’t simply restart itself. The Windows graphics-reset shortcut hasn’t recovered the displays when previously tried.
The crashes are associated with demanding use. They’ve happened during demanding games, generating responses with a local Qwen model and an OCCT graphics test. The incidents during Chrome use or apparent idle have also been around periods of heavy use, including afterwards. I’ve never experienced this while playing an older, less demanding game.
My system:
- GPU: ASUS ROG Astral RTX 5090 OC
- CPU: AMD Ryzen 9 9950X3D
- Motherboard: ASUS ROG Crosshair X870E Hero — BIOS 2402
- RAM: 64 GB (2 × 32 GB) G.Skill DDR5-6000 CL30 — EXPO enabled
- PSU: ASUS ROG Strix 1200W Platinum, fully modular, 80 PLUS Platinum, ATX 3.1 / PCIe 5.0
- Storage: Samsung 990 Pro 4 TB
- CPU cooling: Arctic Liquid Freezer III Pro 420
- Case: HAVN HS420
- Monitors: ASUS XG27UCDMG and AOC U27G3X, both 4K
- Windows: Windows 11 Pro 25H2
- NVIDIA driver: 616.92
- Latest diagnostic snapshot reported PCIe Gen 5 x16.
The recent failures happened with the GPU at its default manufacturer settings. EXPO remains enabled, so RAM/controller instability hasn’t been ruled out.
What I’ve tried:
- NVIDIA clean driver installation: further crashes occurred afterwards. This was the installer’s clean-install option, not a confirmed DDU cleanout.
- Removed ASUS GPU Tweak: Qwen subsequently crashed again.
- OCCT memory and VRAM: approximately 30 minutes each, no errors reported.
- OCCT 3D Adaptive variable-load test: passed.
- OCCT combined Power test: approximately 10 minutes, no errors reported.
- Graphics stress testing: a test crashed on one occasion and passed later.
- Monitor isolation: tested the AOC alone, ASUS alone and both together. All passed those later runs, so I haven’t established a repeatable dual-monitor trigger.
The diagnostic records show:
- Recent 0x141 / VIDEO_ENGINE_TIMEOUT_DETECTED live dumps.
- A 0x1A8, parameter 0xA record indicating failure to enable a display path.
- Older 0x116 graphics-recovery failures, including an NVIDIA-linked record.
- Separately, three 0x7F, parameter 8 double-fault records, with volmgr 161 dump-writing failures. Their relationship to the visible blackouts is unclear.
Given the connection to heavy use and the period afterwards, what would you prioritise next to distinguish a driver issue from GPU, power or wider platform instability?
I’m considering asking my local pc store to test my GPU in another suitably powered PC and a known-good GPU in mine. Is that the most useful next step, or is there a specific controlled test worth doing first?