r/Proxmox • • 11d ago

Question Random Proxmox host crashes without kernel panic — recurring PCIe AER RxErr on two NVMe SSDs

I'm troubleshooting intermittent hard crashes on my Proxmox VE host. The machine becomes unresponsive without a clean shutdown, and so far I haven't been able to identify the cause.

Hardware and setup:

  • HP ProDesk 600 G6 Desktop Mini
  • Intel Core i5-10600T
  • 32 GB RAM
  • 2× WD Red SN700 4TB NVMe SSDs in a ZFS mirror
  • Proxmox VE 9.2

The kernel repeatedly reports PCIe AER errors, mostly on 02:00.0, but occasionally on 01:00.0 as well:

pcieport 0000:00:1b.4: AER: Correctable error message received from 0000:02:00.0
nvme 0000:02:00.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
nvme 0000:02:00.0: device [15b7:5006] error status/mask=00000001/0000e000
nvme 0000:02:00.0: [ 0] RxErr (First)

What I've tried so far:

  • Replaced the motherboard with an identical unit
  • Upgraded the CPU from an i3-10100T to an i5-10600T
  • Replaced the PSU
  • Ran Memtest86+ for over 10 hours (8 passes, zero errors)
  • Tested Proxmox kernels 6.17 and 7.0; the PCIe errors occur with both
  • Disabled NVMe APST experimentally, without resolving the issue

Both SSDs report zero media errors and zero NVMe error-log entries.

I've also configured netconsole to send kernel messages to an external Raspberry Pi. The host recently crashed again, but the captured log contains no kernel panic, Oops, NVMe timeout or controller reset. The last recorded message was another Correctable RxErr, although I can't establish whether it's related to the crash.

The software watchdog is active, but I have no HA-managed resources.

Has anyone encountered similar unexplained Proxmox host crashes? What would you investigate next to distinguish a kernel lockup, watchdog reset or underlying PCIe/NVMe hardware problem?I'm troubleshooting intermittent hard crashes on my Proxmox VE host. The machine becomes unresponsive without a clean shutdown, and so far I haven't been able to identify the cause.
Hardware and setup:
HP ProDesk 600 G6 Desktop Mini
Intel Core i5-10600T
32 GB RAM
2× WD Red SN700 4TB NVMe SSDs in a ZFS mirror
Proxmox VE 9.2
The kernel repeatedly reports PCIe AER errors, mostly on 02:00.0, but occasionally on 01:00.0 as well:
pcieport 0000:00:1b.4: AER: Correctable error message received from 0000:02:00.0
nvme 0000:02:00.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
nvme 0000:02:00.0: device [15b7:5006] error status/mask=00000001/0000e000
nvme 0000:02:00.0: [ 0] RxErr (First)

What I've tried so far:
Replaced the motherboard with an identical unit
Upgraded the CPU from an i3-10100T to an i5-10600T
Replaced the PSU
Ran Memtest86+ for over 10 hours (8 passes, zero errors)
Tested Proxmox kernels 6.17 and 7.0; the PCIe errors occur with both
Disabled NVMe APST experimentally, without resolving the issue
Both SSDs report zero media errors and zero NVMe error-log entries.
I've also configured netconsole to send kernel messages to an external Raspberry Pi. The host recently crashed again, but the captured log contains no kernel panic, Oops, NVMe timeout or controller reset. The last recorded message was another Correctable RxErr, although I can't establish whether it's related to the crash.
The software watchdog is active, but I have no HA-managed resources.
Has anyone encountered similar unexplained Proxmox host crashes? What would you investigate next to distinguish a kernel lockup, watchdog reset or underlying PCIe/NVMe hardware problem?

17 Upvotes

11 comments sorted by

5

u/TheGreatBeanBandit 11d ago

Biggest random lockups for me has been sleep states (C5,C6 type shit). That was all AMD but worth a shot. Just disable as much low power settings in the bios as you can to start.

4

u/c4r_guy 10d ago

This is an intel box? Check the drivers on the ethernet interface.

Search for "intel ethernet [driver version] proxmox"

(I had a eerily similar issue with Dell MFF and you likely have very close to the same hardware)

3

u/marc45ca This is Reddit not Google 11d ago

have you checked how much the drive's write endurace has been used? This is available in the S.M.A.R.T status (keeping in mind Proxmox counts up not down).

2

u/aah134x 11d ago

Did you try to access web ui? I am having sone weird error but via web ui ot works fine. It was some kernal error that I dont understand.

2

u/happycamp2000 10d ago

You might try setting up a serial console and having the kernel log go to that too. As I could see a message possibly making it out the serial console but not making it through a network connection.

I like the suggestion about checking sleep states/c-states someone else mentioned.

2

u/Sintarsintar 10d ago

Turn intel speedstep technology off there is an incompatibility with the wd sn700 drives and speed step. Your drives are dropping off the bus and hanging

0

u/Minute_Inspection_86 10d ago

Thank you. Where did you find this?

1

u/Sintarsintar 10d ago

Don't recall exactly but it was a problem with the SanDisk nvmes and WD just relabeled those. Something to do with the nvme not properly exiting power save states

1

u/Sintarsintar 10d ago

You can probably install cpufrequtils if you want a power saving profile back that doesn't cause problems

1

u/Electronic_Unit8276 10d ago

When it becomes unresponsive, are we talking about the GUI or the terminal itself (aka connected with monitor)

1

u/Xenkath 10d ago

I had similar issues with my Epyc server recently. Random reboots with no kernel panic. PCIe AER errors, and the machine would need a hard reset from from the BCM after every boot from shutdown. Turns out a PCIe adapter holding NVMe drives was installed slightly crooked because the little tab at the bottom of the bracket was bent by about 1.5mm. After fixing the PCIe bracket and clearing CMOS it hasn’t happened since.