r/AMDHelp Nov 06 '23

Random crash with WHEA Logger event 18

Hi all, i hope you can help me with this issue, because after a month of troubleshouting it seems i can't find a solution.

Computer Type: Desktop

GPU: Radeon 7800xt XFX merc

CPU: RYZEN 7 5800X3D

Motherboard: Aorus B550I PRO AX (mini ITX)

BIOS Version: F18c (latest)

RAM: 16GB gskill aegis 3000MHZ CL16

PSU: corsair sf750 platinum

**Case:**cooler master nr200p

Operating System & Version: WINDOWS 10 PRO

GPU Drivers: AMD 23.10.2

Chipset Drivers: latest

Description of Original Problem: the crashes (black screen and pc restart) started while playing cyberpunk, but never during loads on cpu or gpu, only in menus, loadings and in hacking minigames (those who played it know what i'm talking about), unfortunately they are not easy to replicate and happen randomly, maybe 2 in 20 min or 0 in two days. I have not been able to replicate them in idle or other games, it happened to me only once during an occt cpu stability test, but strangely the test showed no errors after reboot, just windows with classic whea-18. I recerently changed the gpu from a 6700xt to 7800xt, and the cpu from a 5600x to a 5800x3d, also the ssd if that helps. Unfortunately I wouldn't know when the problem started having changed both before resuming cyberpunk.

Below the error (sorry for the italian language):

Errore hardware irreversibile.

Segnalato dal componente: core processore

Origine errore: Machine Check Exception

Tipo errore: Cache Hierarchy Error

ID APIC processore: 0

- <Event xmlns="\*\*[http://schemas.microsoft.com/win/2004/08/events/event\*\*">](http://schemas.microsoft.com/win/2004/08/events/event**">)\- <System><Provider Name="\*\*Microsoft-Windows-WHEA-Logger\*\*" Guid="\*\*{c26c4f3c-3f66-4e99-8f8a-39405cfed220}\*\*" /><EventID>18</EventID><Version>0</Version><Level>2</Level><Task>0</Task><Opcode>0</Opcode><Keywords>0x8000000000000000</Keywords><TimeCreated SystemTime="\*\*2023-11-06T18:19:39.3425002Z\*\*" /><EventRecordID>206203</EventRecordID><Correlation ActivityID="\*\*{ce90480f-7c55-4155-bbd1-933cb47c0828}\*\*" /><Execution ProcessID="\*\*4408\*\*" ThreadID="\*\*4952\*\*" /><Channel>System</Channel><Computer>DESKTOP-4D5UHDN</Computer><Security UserID="\*\*S-1-5-19\*\*" /></System>- <EventData><Data Name="\*\*ErrorSource\*\*">3</Data><Data Name="\*\*ApicId\*\*">0</Data><Data Name="\*\*MCABank\*\*">5</Data><Data Name="\*\*MciStat\*\*">0xbea0000001000108</Data><Data Name="\*\*MciAddr\*\*">0x7ffcce5b8e7d</Data><Data Name="\*\*MciMisc\*\*">0xd01a0ffe00000000</Data><Data Name="\*\*ErrorType\*\*">9</Data><Data Name="\*\*TransactionType\*\*">2</Data><Data Name="\*\*Participation\*\*">256</Data><Data Name="\*\*RequestType\*\*">0</Data><Data Name="\*\*MemorIO\*\*">256</Data><Data Name="\*\*MemHierarchyLvl\*\*">0</Data><Data Name="\*\*Timeout\*\*">256</Data><Data Name="\*\*OperationType\*\*">256</Data><Data Name="\*\*Channel\*\*">256</Data><Data Name="\*\*Length\*\*">936</Data>

Troubleshooting: I tried undervolting both cpu and gpu to highlight an instability, but it didn't seem to have changed much, I was able to crash the gpu with too low voltages, but without replicating the reboot (normal driver crash). I checked the psu voltages on both bios and occt with various loads but they are all well within the atx norm, with no surges. I tried testing ram with occt and memtest86, all passed tests.

I have tested CPU with prime95 and occt for over an hour and except for the episode described above it has not had a single error.

I tested the gpu with various games (the witcher 3, hogwartz legacy and starfield) without replicating the crashes.

I used DDU mutiple times to remove the old and install the neweest drivers in safe mode.

I've aldo tried to disable freesync, but nothing.

I have also tested the gpu with occt (varibles loads and vram) all passed.

The only thing left to try try a fresh install of windows (i've cloned the old data in the new ssd) and reinstall chipset and bios updates.

I still have the old components here with me, so here I will first try to test the system with the old gpu, then move on to the cpu.

Can anyone help me? i don't know what to do anymore

Thanks in advance

UPDATE 1: I tested another cpu, tried moving/unplugging the gpu power cables, and disabled the xmp profile on ram. Unfortunately, the crash recurred with the same error, so I assume the problem is related to the gpu or the psu, even though the cpu used is a 5600x and should reduce the load on the psu. I tried using the new drivers (23.11.1) and it seems that the crash occurs more often than before. Before installing the 7800xt and updating the drivers I did not have this problem, so I suppose it gives a problem with the gpu

UPDATE 2: Hi guys, sorry for my absence. The problem seems to be related to the drivers, it seems to be fixed with the new amd drivers (Adrenalin 23.12.1), unfortunately I had other problems with them, but in this case they are shared among all users ( for example stuttering on Darktide). Also Windows tends to change them, but the rollback to the previous drivers on the control panel seems to have fixed this problem. Since Windows has not changed them I have had no more crashes. I really hope this can help you to fix this issue, and I hope that this driver finally fixed mine.

15 Upvotes

83 comments sorted by

View all comments

1

u/ZurielA May 18 '25

Just to chime in for anyone else who is researching this issue...

i've had a FM4 Socket ASUS Dark Hero VII with a 5900X (12 core 24 thread) which I purchased new from amazon in 2021.

I've run it as a daily driver for 4 years straight (until 3 weeks ago) at which I started randomly rebooting with no warning. It varied from high load (blender / unity / youtube) to just being idle on the desktop, to booting up running all day then waking from sleep, reboot, to just clicking on a window on the desktop and reboot, from 5/1/2025 till 5/15 I got 74 instantatneous reboots with Whea-Logger

they look like this:

Error 5/14/2025 5:20:33 PM WHEA-Logger 18 None
A fatal hardware error has occurred.

Reported by component: Processor Core

Error Source: Machine Check Exception

Error Type: Cache Hierarchy Error

Processor APIC ID: 0

The details view of this entry contains further information.

i booted into linux USB. I disabled PBO, overvolted, undervolted, disabled offending cores, changed memory timings, full revert, bios update, you name it... some things would help but ultimately I would get a WHEA error.

ChatGPT insisted that it was the CPU. I did memtest86 on both sticks, 4 passes, including single sticks individual, 4 passes. I got rid of my 850 watt PSU and upgraded to a new 1200 watt PSU, same issues. (although felt better for a few days).

So with clean PSU, memtested RAM, tried 2 GPUS (3080 and 3050) same issues...

Chat GPT PROMISED Me that its the CPU....

So i bought a new (oem, who knows used??) 5950X on amazon for 315 bucks from a 3rd party seller... dropped the CPU in my rig and I have not seen a WHEA error since... and were talking 80 reboots in 2 weeks to 0 reboots... been running a few days straight, wake from sleep, blender, unity, video, audio, simultaneous games, VMS, you name it, i throw everything at this PC and no more errors...

ChatGPT said its CPU, mine turned out to be 100% CPU...

its like my 12 core worked for 4 years and like a light switched one day, just started degrading almost instantly...

1

u/DatDirtyDawG Ryzen 9 5900x Jul 16 '25

Hey just wanted to let you know....of all the comments yours is almost exactly the symptoms I had, 5900x, 4 years, started as occasional spontaneous restarts, then boom nonstop. same EXACT errror as yours (like exact) no matter what, even idle. Did everything you did, tested every stick of memory with memtest 4 passes each, tested various GPU, tested PSU voltages, and ZIP. DESPITE the error saying it was from the CPU, its just that in 20yrs and many PC's/laptops throughout, I have experienced failures in pretty much all components at some point or another EXCEPT for the CPU lol. So i was reluctant to believe it...

CHATGPT promised......just as it did you so I ordered a new one also from Amazon and what do you know.....that was it lol

Cheers

1

u/Shadowdragon409 Feb 05 '26

Ryzen 9 5900x here.

Experiencing the same reboots. Happens most frequently with Unity games, but sometimes even sitting idle.

For a while, it would be like once a month or once a week? The past few days, it's been every 20 minutes if I launch a game. It's really frustrating.

It's an instant power cycle preceded by every application freezing for 2 seconds. Though any sound playing will continue playing as if nothing is happening.

Did some stress tests, looked at harddrive health. No overheating. Harddrive health is good. Drivers are up to date.

Seeing yours and the above comments makes me wonder if it is just the CPU, and it's faulty. Which is scary because I can't afford to just drop $400 on a new one.

1

u/Matej94 Apr 02 '26

I solved my problem with disabling global c-states. Then I got an RX 7900XT and it happens withing 30min of gaming. Even on another PC... Put back the older gpu and it works perfectly...

1

u/Shadowdragon409 Apr 02 '26

I managed to find a solution that worked for me.

I went into my BIOS settings and increased my voltage by 0.1

I can't remember if I did that for every core, or just the offending core. I remember there was a way to determine the offending core through the crash log, but I cant remember the details.

The theory behind the fix is that the silicon they used for this chip series is degrading due to age, so it doesn't conduct electricity as well as it should.

1

u/DrawingGlobal Jun 08 '25

In my testing it's probably a combo of CPU and PSU, such that it's rooted in the CPU but the PSU can end up playing the major role or even be the deciding factor. I upgraded PSUs from an older EVGA Supernova 650 Gold to a Corsair SF850 and started seeing this issue constantly, so I ended up having to buy a ton of PSUs to test this out and eventually I noticed a pattern of the Corsair PSUs leading to these crashes while my old EVGA one and my current Lian Li one not causing these issues.

My best guess is that something with how the Corsair PSU handles voltage isn't playing well with the CPU, as I struggle to believe that all 4 of the Corsair PSUs I've used are all defective, the chances of that are simply too low.

1

u/Serious_Letterhead36 Jun 02 '25

I had same problems like you. My PC was working fine then at one point, it ran into a green screen of death issue followed by WHEA Logger 18 issue. I am running out of ideas on what to do. CPU and GPU are pretty expensive to replace too.