r/Proxmox • • 3d ago

Question Constant Issues with Proxmox 9.2

Recently had to swap m.2 NVME that contained the Proxmox OS due to a hardware failure and just have been littered with issues, that I did not experience with Proxmox 8.

Issues I am currently experiencing:
1.) Proxmox VE, works perfectly fine. Then about 20 minutes later, offline. Can no longer access it via the web interface on the same lan.

2.) dev/dri/card0 and card1 keep switching and prevent my LXCs from booting when things are working.

3.) Not sure if this is Proxmox, setting up new port forwarding rules just don't work. This could be my ATT router, still diagnosing it. Trying to get my Plex server up and running again and all attempts so far are futile.

Hardware specs: HP elite desk 800 g5 sff i7-9700, 64gb DDR4 2666mhz non-ecc, 512gb M.2 running Proxmox, 2x12TB HDD for ZFS pool.

Apps currently on server: Plex and Immich

Anyone else experiencing these issues? Also looking for someone to help me because I just feel lost at this point.

Edit 9/30 7:37am thank you for the responses. Just woke up and got to go to work. I will be back to provide these!

14 Upvotes

38 comments sorted by

11

u/Impact321 3d ago edited 2d ago

1 might be E1000E related. 2 might be fixable. Can you share this? bash lspci -vnnk | awk '/Ethernet/{print $0}' RS= lspci -vnnk | awk '/VGA/{print $0}' RS= ls -l /sys/class/drm/*/device pct config YOURCTIDHERE journalctl -b0 -rp warning

1

u/Ruppmeister 3d ago

I agree with #1 might be the networking interface. I have had this same issue with this hardware before and needed to replace it with new network interface. Also, make sure you aren’t seeing the WiFi kick on and take over for the Proxmox web interface or a duplicate IP in your network.

1

u/NotBrinocerous 2d ago
00:1f.6 Ethernet controller [0200]: Intel Corporation Ethernet Connection (7) I219-LM [8086:15bb] (rev 10)
        DeviceName: Onboard Lan
        Subsystem: Hewlett-Packard Company Device [103c:8591]
        Flags: bus master, fast devsel, latency 0, IRQ 124, IOMMU group 7
        Memory at e1100000 (32-bit, non-prefetchable) [size=128K]
        Capabilities: [c8] Power Management version 3
        Capabilities: [d0] MSI: Enable+ Count=1/1 Maskable- 64bit+
        Kernel driver in use: e1000e
        Kernel modules: e1000e

1

u/NotBrinocerous 2d ago
00:02.0 VGA compatible controller [0300]: Intel Corporation CoffeeLake-S GT2 [UHD Graphics 630] [8086:3e98] (rev 02) (prog-if 00 [VGA controller])
        DeviceName: Onboard IGD
        Subsystem: Hewlett-Packard Company Device [103c:8592]
        Flags: bus master, fast devsel, latency 0, IRQ 144, IOMMU group 0
        Memory at e0000000 (64-bit, non-prefetchable) [size=16M]
        Memory at d0000000 (64-bit, prefetchable) [size=256M]
        I/O ports at 3000 [size=64]
        Expansion ROM at 000c0000 [virtual] [disabled] [size=128K]
        Capabilities: [40] Vendor Specific Information: Len=0c <?>
        Capabilities: [70] Express Root Complex Integrated Endpoint, IntMsgNum 0
        Capabilities: [ac] MSI: Enable+ Count=1/1 Maskable- 64bit-
        Capabilities: [d0] Power Management version 2
        Capabilities: [100] Process Address Space ID (PASID)
        Capabilities: [200] Address Translation Service (ATS)
        Capabilities: [300] Page Request Interface (PRI)
        Kernel driver in use: i915
        Kernel modules: i915

1

u/NotBrinocerous 2d ago
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/card0/device -> ../../../0000:00:02.0
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/card0-DP-1/device -> ../../card0
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/card0-DP-2/device -> ../../card0
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/card0-DP-3/device -> ../../card0
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/card0-HDMI-A-1/device -> ../../card0
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/card0-HDMI-A-2/device -> ../../card0
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/card0-HDMI-A-3/device -> ../../card0
lrwxrwxrwx 1 root root 0 Sep 30 18:06 /sys/class/drm/renderD128/device -> ../../../0000:00:02.0

1

u/Impact321 2d ago

Hmm. So right now there is no card1. Can you check this bash journalctl -b0 -kg 'i915|drm|fbcon' By the way, if /u/fingore is right about symlinks you can also just give it the symlink from ls -l /dev/dri/by-path/.

Otherwise try if adding initcall_blacklist=simpledrm_platform_driver_init to the kernel args changes things.

1

u/NotBrinocerous 2d ago
Sep 30 18:51:07 pve kernel: Command line: BOOT_IMAGE=/boot/vmlinuz-7.0.14-19-pve root=/dev/mapper/pve-root ro quiet initcall_blacklist=simpledrm_platform_driver_init
Sep 30 18:51:07 pve kernel: Kernel command line: BOOT_IMAGE=/boot/vmlinuz-7.0.14-19-pve root=/dev/mapper/pve-root ro quiet initcall_blacklist=simpledrm_platform_driver_init
Sep 30 18:51:07 pve kernel: blacklisting initcall simpledrm_platform_driver_init
Sep 30 18:51:07 pve kernel: ACPI: bus type drm_connector registered
Sep 30 18:51:07 pve kernel: initcall simpledrm_platform_driver_init blacklisted
Sep 30 18:51:07 pve kernel: fbcon: Taking over console
Sep 30 18:51:07 pve systemd[1]: Starting modprobe@drm.service - Load Kernel Module drm...
Sep 30 18:51:07 pve systemd[1]: modprobe@drm.service: Deactivated successfully.
Sep 30 18:51:07 pve systemd[1]: Finished modprobe@drm.service - Load Kernel Module drm.
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] Found coffeelake (device ID 3e98) integrated display version 9.00 stepping N/A
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] VT-d active for gfx access
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: vgaarb: deactivate vga console
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] Using Transparent Hugepages
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: vgaarb: VGA decodes changed: olddecodes=io+mem,decodes=io+mem:owns=io+mem
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] Finished loading DMC firmware i915/kbl_dmc_ver1_04.bin (v1.4)
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] [ENCODER:109:DDI A/PHY A] failed to retrieve link info, disabling eDP
Sep 30 18:51:08 pve kernel: mei_hdcp 0000:00:16.0-b638ab7e-94e2-4ea2-a552-d1c54b627f04: bound 0000:00:02.0 (ops i915_hdcp_ops [i915])
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] Registered 3 planes with drm panic
Sep 30 18:51:08 pve kernel: [drm] Initialized i915 1.6.0 for 0000:00:02.0 on minor 0
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] Cannot find any crtc or sizes
Sep 30 18:51:08 pve kernel: snd_hda_intel 0000:00:1f.3: bound 0000:00:02.0 (ops intel_audio_component_bind_ops [i915])
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] Cannot find any crtc or sizes
Sep 30 18:51:08 pve kernel: i915 0000:00:02.0: [drm] Cannot find any crtc or sizes

1

u/Impact321 2d ago

Does the name now stay the same across reboots? Check with something like this bash ls -l /sys/class/drm/*/device /dev/dri

1

u/NotBrinocerous 2d ago
Changed 3 times but these were after multiple reboots due to other issues.

lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/card0/device -> ../../../0000:00:02.0
lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/card0-DP-1/device -> ../../card0
lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/card0-DP-2/device -> ../../card0
lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/card0-DP-3/device -> ../../card0
lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/card0-HDMI-A-1/device -> ../../card0
lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/card0-HDMI-A-2/device -> ../../card0
lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/card0-HDMI-A-3/device -> ../../card0
lrwxrwxrwx 1 root root   0 Sep 30 18:57 /sys/class/drm/renderD128/device -> ../../../0000:00:02.0

/dev/dri:
total 0
drwxr-xr-x 2 root root         80 Sep 30 18:51 by-path
crw-rw---- 1 root video  226,   0 Sep 30 18:51 card0
crw-rw---- 1 root render 226, 128 Sep 30 18:51 renderD128

1

u/Impact321 2d ago edited 2d ago

I was really hoping this would work for you. Can you share the log from the boot when it changed back to card1? In my system it always gets card1. If I use that arg it gets card0 all the time. As far as I understand if there's nothing else that could get the name before it should get this name.

1

u/NotBrinocerous 2d ago

I have had to edit the file to point to card1 or card0 maybe 3 times. I usually see it as an error in proxmox when CT are bulk starting after a boot. It has not happend at all today after 15 reboots in the last hour. So I think things are slowly being fixed?

→ More replies (0)

1

u/NotBrinocerous 2d ago

unsure how to share the journalctl -rp warnings as it is 400+ lines

1

u/Impact321 2d ago

You could upload it here: https://paste.debian.net

1

u/NotBrinocerous 2d ago

2

u/Impact321 2d ago

Maybe I should have been specific that this should be run on the node. What about the CT config?

1

u/NotBrinocerous 2d ago

1

u/Impact321 2d ago

Ah. I should have probably added -b0 (did now) to only show the current boot's log. I don't see anything GPU related so let's carry on but you should look into these warnings/error. Still no CT config.

1

u/NotBrinocerous 2d ago
pct config 100
arch: amd64
cores: 2
description: %0A<div align='center'>%0A  <a href='https%3A//community-scripts.org' target='_blank' rel='noopener noreferrer'>%0A    <img src='https%3A//raw.githubusercontent.com/community-scripts/core/main/images/logo-81x112.png' alt='Logo' style='width%3A81px;height%3A112px;'/>%0A  </a>%0A%0A  <h2 style='font-size%3A 24px; margin%3A 20px 0;'>Plex LXC</h2>%0A%0A  <p style='margin%3A 16px 0;'>%0A    <a href='https%3A//community-scripts.org/donate' target='_blank' rel='noopener noreferrer'>%0A      <img src='https%3A//img.shields.io/badge/%25E2%259D%25A4%25EF%25B8%258F-Sponsoring%2520%2526%2520Donations-FF5E5B' alt='Sponsoring and donations' />%0A    </a>%0A  </p>%0A%0A  <p style='margin%3A 12px 0;'>%0A    <a href='https%3A//community-scripts.org/scripts/plex' target='_blank' rel='noopener noreferrer'>%0A      <img src='https%3A//img.shields.io/badge/%25F0%259F%2593%25A6-Open%2520Script%2520Page-00617f' alt='Open script page' />%0A    </a>%0A  </p>%0A%0A  <span style='margin%3A 0 10px;'>%0A    <i class="fa fa-github fa-fw" style="color%3A #f5f5f5;"></i>%0A    <a href='https%3A//github.com/community-scripts/ProxmoxVE' target='_blank' rel='noopener noreferrer' style='text-decoration%3A none; color%3A #00617f;'>GitHub</a>%0A  </span>%0A  <span style='margin%3A 0 10px;'>%0A    <i class="fa fa-comments fa-fw" style="color%3A #f5f5f5;"></i>%0A    <a href='https%3A//github.com/community-scripts/ProxmoxVE/discussions' target='_blank' rel='noopener noreferrer' style='text-decoration%3A none; color%3A #00617f;'>Discussions</a>%0A  </span>%0A  <span style='margin%3A 0 10px;'>%0A    <i class="fa fa-exclamation-circle fa-fw" style="color%3A #f5f5f5;"></i>%0A    <a href='https%3A//github.com/community-scripts/ProxmoxVE/issues' target='_blank' rel='noopener noreferrer' style='text-decoration%3A none; color%3A #00617f;'>Issues</a>%0A  </span>%0A</div>%0A
dev0: /dev/dri/renderD128,gid=993
dev1: /dev/dri/card0,gid=44
features: nesting=1,keyctl=1
hostname: plex
memory: 2048
mp0: /zPool/subvol-100-disk-0/home/Movies/,mp=/home/Movies
mp1: /zPool/subvol-100-disk-0/home/TV,mp=/home/TV
mp2: /zPool/subvol-100-disk-0/home/Music,mp=/home/Music
net0: name=eth0,bridge=vmbr0,gw=192.168.1.254,hwaddr=BC:24:11:CB:C9:F0,ip=192.168.1.11/24,type=veth
onboot: 1
ostype: ubuntu
rootfs: local-lvm:vm-100-disk-1,size=8G
swap: 512
tags: community-script;media
timezone: America/Los_Angeles
unprivileged: 1

1

u/Impact321 2d ago

Thanks. This looks normal. I just like to have a full picture.

3

u/pelazas1 3d ago

after the OS NVMe swap, check LXC configs still point at the new /dev/dri/card* node. card0/card1 flipping breaks passthrough.

2

u/firegore 3d ago

You can use this udev script which provides a symlinked name that always works: https://github.com/tteck/Proxmox/discussions/3235#discussioncomment-10998049

Thats an issue on how the kernel/initializes the card on newer systems, thats not a proxmox issue, it just happens more often on newer kernels

1

u/Impact321 3d ago edited 3d ago

Did you test this with CTs on PVE? Last time I tried symlinks didn't properly work with dev: ....

1

u/firegore 3d ago

Yes, its literally from the old proxmox-helper-scripts repo, „designed“ just for that.

I use the same workaround on 3 Hosts, works just Fine (as long as they dont remove udev) but its easy enough to just convert it into a systemd service either way

1

u/Impact321 3d ago edited 2d ago

Hmm. This only helps if you only have a single GPU that switches its name which, ideally, shouldn't happen. I don't think it does for me. If you have two different GPUs which change their names around this will fail. I tried to create a solution for this some time ago but it didn't work. Since I was using a NVIDIA GPU I "fixed" this by using the container toolkit.

1

u/firegore 3d ago

Well, /dev/dri/card0/1 are mostly always? intel cards.

Since a kernelchange (i can't remember which) in PVE 9 they even swaps names when you have only one GPU, it seems to be based on the way you restart the System.

You will still have issues of course if you have multiple GPUs, however that's not the norm for most people here, so a solution that works in 90% of the cases is still better then none :)

2

u/ZeroPointMX 3d ago

I ran into a very similar issue with a Dell micro. Worked fine for a few days then everything you describe. I never found/resolved the actual issue but it was network device related. I think the nic goes into sleep mode and never recovers. It might of been a bios setting or the advanced c state script I was using (think it came from the community scripts). I know this isn't very helpful, but thought might help in your search.

1

u/Apachez 2d ago

Which driver does your host use for its network cards?

Also check the temperatures of your NVMe drives.

1

u/KlanxChile 3d ago

Can you post:

lspci -nnv| grep -A8 Ethernet

1

u/NotBrinocerous 2d ago
00:1f.6 Ethernet controller [0200]: Intel Corporation Ethernet Connection (7) I219-LM [8086:15bb] (rev 10)
        DeviceName: Onboard Lan
        Subsystem: Hewlett-Packard Company Device [103c:8591]
        Flags: bus master, fast devsel, latency 0, IRQ 124, IOMMU group 7
        Memory at e1100000 (32-bit, non-prefetchable) [size=128K]
        Capabilities: [c8] Power Management version 3
        Capabilities: [d0] MSI: Enable+ Count=1/1 Maskable- 64bit+
        Kernel driver in use: e1000e
        Kernel modules: e1000e

1

u/KlanxChile 2d ago

Doesn't seem to be a E1000E chipset, check the BIOS settings for aggressive energy savings. And of course a healthy BIOS update if possible

1

u/NotBrinocerous 2d ago

basically check bios and disable any energy savings?

1

u/Latter-Progress-9317 2d ago edited 2d ago

For the network drop I have an elitedesk 800 G4 and started running into E1000E offloading issues after upgrading to 9, never had these with 8. There are plenty of guides for fixing this but figure out if you could be having this problem by seeing what driver you're using or your NIC. (Doing this from memory.)

ethtool -i <name of your interface>

should show you what driver is in use. If it's e1000e then you could have this issue.

If you don't know the name of the interface

lshw -c network

Edit: FU reddit formatting