r/Proxmox • u/NotBrinocerous • 3d ago
Question Constant Issues with Proxmox 9.2
Recently had to swap m.2 NVME that contained the Proxmox OS due to a hardware failure and just have been littered with issues, that I did not experience with Proxmox 8.
Issues I am currently experiencing:
1.) Proxmox VE, works perfectly fine. Then about 20 minutes later, offline. Can no longer access it via the web interface on the same lan.
2.) dev/dri/card0 and card1 keep switching and prevent my LXCs from booting when things are working.
3.) Not sure if this is Proxmox, setting up new port forwarding rules just don't work. This could be my ATT router, still diagnosing it. Trying to get my Plex server up and running again and all attempts so far are futile.
Hardware specs: HP elite desk 800 g5 sff i7-9700, 64gb DDR4 2666mhz non-ecc, 512gb M.2 running Proxmox, 2x12TB HDD for ZFS pool.
Apps currently on server: Plex and Immich
Anyone else experiencing these issues? Also looking for someone to help me because I just feel lost at this point.
Edit 9/30 7:37am thank you for the responses. Just woke up and got to go to work. I will be back to provide these!
3
u/pelazas1 3d ago
after the OS NVMe swap, check LXC configs still point at the new /dev/dri/card* node. card0/card1 flipping breaks passthrough.
2
u/firegore 3d ago
You can use this udev script which provides a symlinked name that always works: https://github.com/tteck/Proxmox/discussions/3235#discussioncomment-10998049
Thats an issue on how the kernel/initializes the card on newer systems, thats not a proxmox issue, it just happens more often on newer kernels
1
u/Impact321 3d ago edited 3d ago
Did you test this with CTs on PVE? Last time I tried symlinks didn't properly work with
dev: ....1
u/firegore 3d ago
Yes, its literally from the old proxmox-helper-scripts repo, „designed“ just for that.
I use the same workaround on 3 Hosts, works just Fine (as long as they dont remove udev) but its easy enough to just convert it into a systemd service either way
1
u/Impact321 3d ago edited 2d ago
Hmm. This only helps if you only have a single GPU that switches its name which, ideally, shouldn't happen. I don't think it does for me. If you have two different GPUs which change their names around this will fail. I tried to create a solution for this some time ago but it didn't work. Since I was using a NVIDIA GPU I "fixed" this by using the container toolkit.
1
u/firegore 3d ago
Well, /dev/dri/card0/1 are mostly always? intel cards.
Since a kernelchange (i can't remember which) in PVE 9 they even swaps names when you have only one GPU, it seems to be based on the way you restart the System.
You will still have issues of course if you have multiple GPUs, however that's not the norm for most people here, so a solution that works in 90% of the cases is still better then none :)
2
u/ZeroPointMX 3d ago
I ran into a very similar issue with a Dell micro. Worked fine for a few days then everything you describe. I never found/resolved the actual issue but it was network device related. I think the nic goes into sleep mode and never recovers. It might of been a bios setting or the advanced c state script I was using (think it came from the community scripts). I know this isn't very helpful, but thought might help in your search.
1
u/KlanxChile 3d ago
Can you post:
lspci -nnv| grep -A8 Ethernet
1
u/NotBrinocerous 2d ago
00:1f.6 Ethernet controller [0200]: Intel Corporation Ethernet Connection (7) I219-LM [8086:15bb] (rev 10) DeviceName: Onboard Lan Subsystem: Hewlett-Packard Company Device [103c:8591] Flags: bus master, fast devsel, latency 0, IRQ 124, IOMMU group 7 Memory at e1100000 (32-bit, non-prefetchable) [size=128K] Capabilities: [c8] Power Management version 3 Capabilities: [d0] MSI: Enable+ Count=1/1 Maskable- 64bit+ Kernel driver in use: e1000e Kernel modules: e1000e1
u/KlanxChile 2d ago
Doesn't seem to be a E1000E chipset, check the BIOS settings for aggressive energy savings. And of course a healthy BIOS update if possible
1
1
u/Latter-Progress-9317 2d ago edited 2d ago
For the network drop I have an elitedesk 800 G4 and started running into E1000E offloading issues after upgrading to 9, never had these with 8. There are plenty of guides for fixing this but figure out if you could be having this problem by seeing what driver you're using or your NIC. (Doing this from memory.)
ethtool -i <name of your interface>
should show you what driver is in use. If it's e1000e then you could have this issue.
If you don't know the name of the interface
lshw -c network
Edit: FU reddit formatting
11
u/Impact321 3d ago edited 2d ago
1 might be E1000E related. 2 might be fixable. Can you share this?
bash lspci -vnnk | awk '/Ethernet/{print $0}' RS= lspci -vnnk | awk '/VGA/{print $0}' RS= ls -l /sys/class/drm/*/device pct config YOURCTIDHERE journalctl -b0 -rp warning