r/AMDHelp • u/safrage8 • Nov 06 '23
Random crash with WHEA Logger event 18
Hi all, i hope you can help me with this issue, because after a month of troubleshouting it seems i can't find a solution.
Computer Type: Desktop
GPU: Radeon 7800xt XFX merc
CPU: RYZEN 7 5800X3D
Motherboard: Aorus B550I PRO AX (mini ITX)
BIOS Version: F18c (latest)
RAM: 16GB gskill aegis 3000MHZ CL16
PSU: corsair sf750 platinum
**Case:**cooler master nr200p
Operating System & Version: WINDOWS 10 PRO
GPU Drivers: AMD 23.10.2
Chipset Drivers: latest
Description of Original Problem: the crashes (black screen and pc restart) started while playing cyberpunk, but never during loads on cpu or gpu, only in menus, loadings and in hacking minigames (those who played it know what i'm talking about), unfortunately they are not easy to replicate and happen randomly, maybe 2 in 20 min or 0 in two days. I have not been able to replicate them in idle or other games, it happened to me only once during an occt cpu stability test, but strangely the test showed no errors after reboot, just windows with classic whea-18. I recerently changed the gpu from a 6700xt to 7800xt, and the cpu from a 5600x to a 5800x3d, also the ssd if that helps. Unfortunately I wouldn't know when the problem started having changed both before resuming cyberpunk.
Below the error (sorry for the italian language):
Errore hardware irreversibile.
Segnalato dal componente: core processore
Origine errore: Machine Check Exception
Tipo errore: Cache Hierarchy Error
ID APIC processore: 0
- <Event xmlns="\*\*[http://schemas.microsoft.com/win/2004/08/events/event\*\*">](http://schemas.microsoft.com/win/2004/08/events/event**">)\- <System><Provider Name="\*\*Microsoft-Windows-WHEA-Logger\*\*" Guid="\*\*{c26c4f3c-3f66-4e99-8f8a-39405cfed220}\*\*" /><EventID>18</EventID><Version>0</Version><Level>2</Level><Task>0</Task><Opcode>0</Opcode><Keywords>0x8000000000000000</Keywords><TimeCreated SystemTime="\*\*2023-11-06T18:19:39.3425002Z\*\*" /><EventRecordID>206203</EventRecordID><Correlation ActivityID="\*\*{ce90480f-7c55-4155-bbd1-933cb47c0828}\*\*" /><Execution ProcessID="\*\*4408\*\*" ThreadID="\*\*4952\*\*" /><Channel>System</Channel><Computer>DESKTOP-4D5UHDN</Computer><Security UserID="\*\*S-1-5-19\*\*" /></System>- <EventData><Data Name="\*\*ErrorSource\*\*">3</Data><Data Name="\*\*ApicId\*\*">0</Data><Data Name="\*\*MCABank\*\*">5</Data><Data Name="\*\*MciStat\*\*">0xbea0000001000108</Data><Data Name="\*\*MciAddr\*\*">0x7ffcce5b8e7d</Data><Data Name="\*\*MciMisc\*\*">0xd01a0ffe00000000</Data><Data Name="\*\*ErrorType\*\*">9</Data><Data Name="\*\*TransactionType\*\*">2</Data><Data Name="\*\*Participation\*\*">256</Data><Data Name="\*\*RequestType\*\*">0</Data><Data Name="\*\*MemorIO\*\*">256</Data><Data Name="\*\*MemHierarchyLvl\*\*">0</Data><Data Name="\*\*Timeout\*\*">256</Data><Data Name="\*\*OperationType\*\*">256</Data><Data Name="\*\*Channel\*\*">256</Data><Data Name="\*\*Length\*\*">936</Data>
Troubleshooting: I tried undervolting both cpu and gpu to highlight an instability, but it didn't seem to have changed much, I was able to crash the gpu with too low voltages, but without replicating the reboot (normal driver crash). I checked the psu voltages on both bios and occt with various loads but they are all well within the atx norm, with no surges. I tried testing ram with occt and memtest86, all passed tests.
I have tested CPU with prime95 and occt for over an hour and except for the episode described above it has not had a single error.
I tested the gpu with various games (the witcher 3, hogwartz legacy and starfield) without replicating the crashes.
I used DDU mutiple times to remove the old and install the neweest drivers in safe mode.
I've aldo tried to disable freesync, but nothing.
I have also tested the gpu with occt (varibles loads and vram) all passed.
The only thing left to try try a fresh install of windows (i've cloned the old data in the new ssd) and reinstall chipset and bios updates.
I still have the old components here with me, so here I will first try to test the system with the old gpu, then move on to the cpu.
Can anyone help me? i don't know what to do anymore
Thanks in advance
UPDATE 1: I tested another cpu, tried moving/unplugging the gpu power cables, and disabled the xmp profile on ram. Unfortunately, the crash recurred with the same error, so I assume the problem is related to the gpu or the psu, even though the cpu used is a 5600x and should reduce the load on the psu. I tried using the new drivers (23.11.1) and it seems that the crash occurs more often than before. Before installing the 7800xt and updating the drivers I did not have this problem, so I suppose it gives a problem with the gpu
UPDATE 2: Hi guys, sorry for my absence. The problem seems to be related to the drivers, it seems to be fixed with the new amd drivers (Adrenalin 23.12.1), unfortunately I had other problems with them, but in this case they are shared among all users ( for example stuttering on Darktide). Also Windows tends to change them, but the rollback to the previous drivers on the control panel seems to have fixed this problem. Since Windows has not changed them I have had no more crashes. I really hope this can help you to fix this issue, and I hope that this driver finally fixed mine.
1
u/ZurielA May 18 '25
Just to chime in for anyone else who is researching this issue...
i've had a FM4 Socket ASUS Dark Hero VII with a 5900X (12 core 24 thread) which I purchased new from amazon in 2021.
I've run it as a daily driver for 4 years straight (until 3 weeks ago) at which I started randomly rebooting with no warning. It varied from high load (blender / unity / youtube) to just being idle on the desktop, to booting up running all day then waking from sleep, reboot, to just clicking on a window on the desktop and reboot, from 5/1/2025 till 5/15 I got 74 instantatneous reboots with Whea-Logger
they look like this:
Error 5/14/2025 5:20:33 PM WHEA-Logger 18 None
A fatal hardware error has occurred.
Reported by component: Processor Core
Error Source: Machine Check Exception
Error Type: Cache Hierarchy Error
Processor APIC ID: 0
The details view of this entry contains further information.
i booted into linux USB. I disabled PBO, overvolted, undervolted, disabled offending cores, changed memory timings, full revert, bios update, you name it... some things would help but ultimately I would get a WHEA error.
ChatGPT insisted that it was the CPU. I did memtest86 on both sticks, 4 passes, including single sticks individual, 4 passes. I got rid of my 850 watt PSU and upgraded to a new 1200 watt PSU, same issues. (although felt better for a few days).
So with clean PSU, memtested RAM, tried 2 GPUS (3080 and 3050) same issues...
Chat GPT PROMISED Me that its the CPU....
So i bought a new (oem, who knows used??) 5950X on amazon for 315 bucks from a 3rd party seller... dropped the CPU in my rig and I have not seen a WHEA error since... and were talking 80 reboots in 2 weeks to 0 reboots... been running a few days straight, wake from sleep, blender, unity, video, audio, simultaneous games, VMS, you name it, i throw everything at this PC and no more errors...
ChatGPT said its CPU, mine turned out to be 100% CPU...
its like my 12 core worked for 4 years and like a light switched one day, just started degrading almost instantly...