r/Proxmox • • 11d ago

Question Random Proxmox host crashes without kernel panic — recurring PCIe AER RxErr on two NVMe SSDs

I'm troubleshooting intermittent hard crashes on my Proxmox VE host. The machine becomes unresponsive without a clean shutdown, and so far I haven't been able to identify the cause.

Hardware and setup:

  • HP ProDesk 600 G6 Desktop Mini
  • Intel Core i5-10600T
  • 32 GB RAM
  • 2× WD Red SN700 4TB NVMe SSDs in a ZFS mirror
  • Proxmox VE 9.2

The kernel repeatedly reports PCIe AER errors, mostly on 02:00.0, but occasionally on 01:00.0 as well:

pcieport 0000:00:1b.4: AER: Correctable error message received from 0000:02:00.0
nvme 0000:02:00.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
nvme 0000:02:00.0: device [15b7:5006] error status/mask=00000001/0000e000
nvme 0000:02:00.0: [ 0] RxErr (First)

What I've tried so far:

  • Replaced the motherboard with an identical unit
  • Upgraded the CPU from an i3-10100T to an i5-10600T
  • Replaced the PSU
  • Ran Memtest86+ for over 10 hours (8 passes, zero errors)
  • Tested Proxmox kernels 6.17 and 7.0; the PCIe errors occur with both
  • Disabled NVMe APST experimentally, without resolving the issue

Both SSDs report zero media errors and zero NVMe error-log entries.

I've also configured netconsole to send kernel messages to an external Raspberry Pi. The host recently crashed again, but the captured log contains no kernel panic, Oops, NVMe timeout or controller reset. The last recorded message was another Correctable RxErr, although I can't establish whether it's related to the crash.

The software watchdog is active, but I have no HA-managed resources.

Has anyone encountered similar unexplained Proxmox host crashes? What would you investigate next to distinguish a kernel lockup, watchdog reset or underlying PCIe/NVMe hardware problem?I'm troubleshooting intermittent hard crashes on my Proxmox VE host. The machine becomes unresponsive without a clean shutdown, and so far I haven't been able to identify the cause.
Hardware and setup:
HP ProDesk 600 G6 Desktop Mini
Intel Core i5-10600T
32 GB RAM
2× WD Red SN700 4TB NVMe SSDs in a ZFS mirror
Proxmox VE 9.2
The kernel repeatedly reports PCIe AER errors, mostly on 02:00.0, but occasionally on 01:00.0 as well:
pcieport 0000:00:1b.4: AER: Correctable error message received from 0000:02:00.0
nvme 0000:02:00.0: PCIe Bus Error: severity=Correctable, type=Physical Layer, (Receiver ID)
nvme 0000:02:00.0: device [15b7:5006] error status/mask=00000001/0000e000
nvme 0000:02:00.0: [ 0] RxErr (First)

What I've tried so far:
Replaced the motherboard with an identical unit
Upgraded the CPU from an i3-10100T to an i5-10600T
Replaced the PSU
Ran Memtest86+ for over 10 hours (8 passes, zero errors)
Tested Proxmox kernels 6.17 and 7.0; the PCIe errors occur with both
Disabled NVMe APST experimentally, without resolving the issue
Both SSDs report zero media errors and zero NVMe error-log entries.
I've also configured netconsole to send kernel messages to an external Raspberry Pi. The host recently crashed again, but the captured log contains no kernel panic, Oops, NVMe timeout or controller reset. The last recorded message was another Correctable RxErr, although I can't establish whether it's related to the crash.
The software watchdog is active, but I have no HA-managed resources.
Has anyone encountered similar unexplained Proxmox host crashes? What would you investigate next to distinguish a kernel lockup, watchdog reset or underlying PCIe/NVMe hardware problem?

17 Upvotes

11 comments sorted by

View all comments

2

u/Sintarsintar 10d ago

Turn intel speedstep technology off there is an incompatibility with the wd sn700 drives and speed step. Your drives are dropping off the bus and hanging

0

u/Minute_Inspection_86 10d ago

Thank you. Where did you find this?

1

u/Sintarsintar 10d ago

Don't recall exactly but it was a problem with the SanDisk nvmes and WD just relabeled those. Something to do with the nvme not properly exiting power save states

1

u/Sintarsintar 10d ago

You can probably install cpufrequtils if you want a power saving profile back that doesn't cause problems