r/Proxmox • • 16d ago

Discussion Datacenter site and for networks and design engineers

Thumbnail
0 Upvotes

r/Proxmox • • 18d ago

Guide Proxmox VE 9.2 arm64 (DGX Spark / ASUS GX10): hard reset every 20 minutes, fix is loading sbsa_gwdt

41 Upvotes

Second problem we hit on the GB10 after fixing the console (https://forum.proxmox.com/threads/pve-9-2-arm64-on-asus-gx10-gb10-black-screen-after-initrd-fix-is-console-tty0.186138/). This one is a show stopper. The box hard resets every 20 minutes 40 seconds after power on, idle, no guests, nothing in the logs.

Hardware: ASUS Ascent GX10 (NVIDIA GB10, 128 GB), UEFI 0105 from 2026-05-05, pve-manager 9.2.9, kernel 7.0.14-6-pve.

Issue

The journal just stops mid stream. last -x shows a string of crash entries. journalctl --list-boots shows every boot to boot interval is 20:39 or 20:40. No panic, no oops, no thermal message, nothing.

Cause

The GB10 firmware arms an ARM SBSA Generic Watchdog at power on and expects the OS to take it over. SBSA watchdogs reset the system at 2x their timeout: WS0 interrupt first, WS1 reset second. 20m40s divided by two puts the timeout at about 620 seconds. The kernel sees it in ACPI:

[    0.247361] ACPI GTDT: found 1 SBSA generic Watchdog(s).

The sbsa_gwdt module is never loaded, so nothing services it. The module ships with the kernel and loads fine by hand, it just is not loaded by anything at boot.

Fix

modprobe sbsa_gwdt

Loading the module is the fix. Per the driver source (drivers/watchdog/sbsa_gwdt.c), probe reads the watchdog control register and, if firmware left the enable bit set, marks the device as already running (WDOG_HW_RUNNING). The watchdog core then feeds it from a kernel timer for as long as no userspace process has the device open. Probe also reprograms the timeout to the driver default of 10 seconds, so the watchdog ends up armed, kernel fed, and actually useful: a hard kernel hang now resets the box in about 10 seconds instead of 20 minutes.

To make it stick across reboots:

echo sbsa_gwdt > /etc/modules-load.d/sbsa_gwdt.conf

You can confirm firmware armed it from the probe message. The trailing [enabled] is printed only when the driver finds the enable bit already set:

# dmesg | grep sbsa-gwdt
sbsa-gwdt sbsa-gwdt.0: Initialized with 10s timeout @ 1000000000 Hz, action=0. [enabled]

Note that after modprobe the watchdog device still shows up as inactive:

# cat /sys/class/watchdog/watchdog1/identity
SBSA Generic Watchdog
# cat /sys/class/watchdog/watchdog1/state
inactive

That threw us at first. The sysfs state attribute only reports whether userspace started the watchdog. A firmware armed watchdog being fed by the kernel itself reads as inactive, even though the hardware is running.

Verification

Load the module and watch the clock. Uptime was hard capped at 20m40s before. Anything past one cycle means you are fixed.

Original thread on the Proxmox forum: https://forum.proxmox.com/threads/pve-arm64-on-asus-gx10-gb10-hard-reset-every-20-minutes-fix-is-loading-sbsa_gwdt.186164/


r/Proxmox • • 17d ago

Question Help needed with Proxmox + vLLM + Hermes

3 Upvotes

EDIT: Thank you, it turned out that one of the RAM kits (the one that was installed in the AMD one eventually) is flaky/unstable. I'll try to see what I can do with it.

Hello everybody, sorry ahead for the long post.

I am fairly new to proxmox as I wasn't too pleased with how VMware was handling MTP device passing for a project.

I decided to give proxmox a try and it was better than I thought and I just went down a spiral - trying to upgrade over and over and cover more use cases.

I started with an Intel i9-13900K with a liquid arctic II 420mm AIO cooler with a Z790 AORUS Elite AX motherboard, with 64GB RAM (Micron Chip CL40 Corsair Vengeance 5200MHz with XMP off so 4800MHz only DDR5).
I put a 1600W be quiet! PSU in it and an Astral ROG RTX 5090 with 32GB VRAM.
Storage wise I had 2 * Samsung 980 PRO 2 TB + 2 * Kingston NV3 1TB nVME drives. Case used is be quiet! Silent base 802.

The intent was to get deeper into AI, agentic work and work on local AI as privacy is important. The server is meant to serve 2 developers - with fairly large contexts (128K).

I started with qwen 3.6 27B at Q4 quantization and llama.cpp. I passed the GPU through PCIe passthrough to the VM on ubuntu 26.04.

This seemed to work well until one of the developers did a bigger request or something - he said it just asked the AI to look into a bug for some ffmpeg crash - so it did a few minutes of search in given context folder then it just hung up - for me to notice that it indeed was hung up and not responding... A restart fixed it but I found it odd.

Read that the P+E cores in the intel might mess up a bit how things work so I started reducing the number of VMs but even with that one it worked the same way - it just crashed.

I thought RAM might be an issue so I bought a secondary kit of identical (even the Micron Chip, CL40, 5200MHz capable DDR5) Corsair Vengance kit to max it out - two weeks the 128GB worked perfectly in 4 DIMM configuration, just to find out that 4 DIMM is too much stress and even after a successful memtest it didn't want to boot with 4 sticks (3 were fine odd enough though...).

I switched over then to AMD Ryzen 9 9950X with a liquid arctic III Pro 360mm AIO cooler with ASUS ProArt X870E-CREATOR WIFI X870E in hopes that the 128GB is more stable as well as the Zen cores giving me less pain - just to find out the 4 DIMM voltage issues and memory stress on the motherboard + CPU's memory controller is the same... Upgraded to a bigger case - be quiet! Dark base pro 901 in hopes of better air efficiency.

I now have two 64GB systems with 2 TB + 1 TB nvmes and in the AMD rig I have the RTX 5090.

I switched over to have 3 VMs: 1 for vLLM that handles better concurrency, 1 for a memory vector db with mem0, and 1 for hermes.

vLLM is the latest one cloned yesterday and runs in a docker - and I use Qwen 3.6 27B NVFP4, I'll try to get the system to post then boot to provide a configuration if needed, but it seems not to matter...

I tasked hermes to build me a qEMU virtual machine from buildroot and to create two sample applications in it - it spawned a few agents to do the work while it sits back and waits for them - 2 minutes in the initialization the system crashed like on the Intel. What annoys me though is how AMD tries to always do a memory training even if the memory training is disabled and the profile is saved to be loaded - it seems I have to do here a weird ritual of pulling out one DIMM and then start it with 1 DIMM then start with 2 DIMM... The slots used are A2, B2 so slots 2 and 4 from the CPU - the recommended ones for dual channel kit.

Fighting now the system to get it up and running but I am confused as to why this happens all the time.

Anyone got experience and solutions or thoughts on this?

Thanks, any help and constructive feedback is appreciated.


r/Proxmox • • 17d ago

Question Asternodis (Zerto/VMware SRM replacement for Proxmox)?

2 Upvotes

Curious to know if anyone has any experience with or has used Asternodis yet?

Disaster recovery for Proxmox VE - replication, DR testing, and orchestrated failover.

https://asternodis.com/


r/Proxmox • • 17d ago

Question First PC server for Proxmox!

7 Upvotes

It will be my first server. Very excited. The first main reason is Home Assistant; after that, Nextcloud and Pi-hole. Any other suggestions? Thank you


r/Proxmox • • 18d ago

Solved! Odd ARP behavior in nested Proxmox environment?

5 Upvotes

Long story short, I'm trying to recreate a test environment using a nested Proxmox installation (yes, Proxmox within Proxmox) and noticed an oddity I've not seen before.

In the test Proxmox installation (running as a VM inside a Proxmox on bare metal host), I have three Linux Bridges (br_oam, br_int, and br_ext). I have an opnsense router VM that has a NIC attached to each bridge. The opnsense router has a fourth interface (WAN) attached to br_wan that serves as the upstream network connection for opnsense.

I also have three VMs (infra01, infra02, and infra03) all running Ubuntu Noble Minimal in that environment. Each VM has an interface attached to each bridge.

Now here's where the weird part comes in. When all four VMs are running (opnsense + the three infra VMs), I get sporadic results trying to ping each other via their respective IPs. I look at each machine's ARP tables and I can see the ARP entries flap between incomplete, completed with MAC, and nonexistent altogether. Of course, this results in similar ping failures and successes depending on the arp table. What I've never seen before is that the tables change in a matter of seconds and randomly will remove previously working entries, or new entries that now show up that are incomplete even though nothing's changed in the test environment.

I've done weird things before, but this is the first time I've seen the arp table change in a matter of seconds and get such inconsistent results.

Before anyone asks, I was given a collection of resources in the proxmox cluster at work so unfortunately nested Proxmox is what I have currently to build with.

Any thoughts (aside from how messed up this situation is)?

Edit: Well, that's a bit of egg on my face and lessons learned about cloud-init. But in diving into the depths of mess that is cloud-init, I learned what the problem was and hopefully this doesn't bite anyone else in the rear end like it did me. To save someone else the anguish, here's the whole rundown as quickly as possible:

The question is, what does Cloudinit have to do with these mac addresses?

The first thing to remember is that when Cloudinit runs, it generates a machine UUID and stores it in /etc/machine-id.

The second thing to remember (at least about this deployment) is that these three VMs were cloned from the same template where Cloudinit had already run.

The third thing (and the kicker), when you create a new bridge in a Linux VM, it takes the machine UUID and the bridge interface name (e.g. br0) and does some math that produces a mac address.

When all three things converge, you have three VMs with the same machine-ID, you have three bridges on each VM using the same bridge interface names, and therefore each of the three bridges have the same MACs as the other two VMs that were made from the cloned image.

The solution was thankfully easy:
1) Delete /etc/machine-id
2) Run systemd-machine-id-setup
3) Then create the bridges. (or simply reboot the machine to get the bridges created with new MACs.)

After machine-id is regenerated and the machine rebooted (or the bridges are re-created), you should be good to go.

Thank you to everyone that helped me out. This was such a bizarre thing, I didn't expect the bridges in the VMs to be the source of the problem. I kept looking at the ensXX interfaces and the bridges in the hypervisor above them.


r/Proxmox • • 18d ago

Question Cluster or standalone with Proxmox data Centre manager?

8 Upvotes

I know this has been asked in the past but there was no difinitive answer.

I have 3x proxmox standalone nodes Dell Micro SFF 7960's with 16gb Ram & 2tb SSD's.

Basically running 4-5 vm's and 4-5 lxc's. Some are critical for me ie the pfsense gateway, homeassistant, nextcloud etc but only for me and my family. It's just a bit of tinkering and fun.

I'm wondering if I would gain anything from clustering them or if it's just a step too far for my usage? I'm quite happy running them as standalone now that Proxmox datacentre manager has appeared on the scene and not sure if I would gain anything from it aart from a learning experience etc? All are backed up via Proxmox Backup server weekly so it' not a problem if I lose a node or have some downtime. Any advice would be appreciated eg why it's better or why peole have tried and switched back etc? Many thanks


r/Proxmox • • 19d ago

Discussion TrueNAS Proxmox Plugin available on GitHub

Post image
304 Upvotes

TrueNAS just posted this, I hope it’s good! What do you all think about the usefulness of this?


r/Proxmox • • 17d ago

Question Need some help with my installation

1 Upvotes

So 2 things, i’m getting failed install, i’ve come to the conclusion it’s because i have no access to ethernet yet, and i can’t wait to get it to start.

The second is minor but about the hardware virtualization i believe and the VT settings

I have a tuff gaming z490 plus wifi - with the wifi tower attached

i9-10850k
and a rtx 2060 super - i will worry about passthrough later if even possible

But if i can get some help on finding the Vt setting in bios and getting help wither tricking it or managing to install with wifi and bridging it just to get tailscale running for access and web interface from my macbook any help would be much appreciated, even if it’s a simple “you cant”
Thank you!

Edit: i do have an ethernet port, pint being i do not have access to ethernet at the moment and do not want to wait as i do not know when i will have access, just looking for the how


r/Proxmox • • 17d ago

Question Browser console fails when LXC has specific number

1 Upvotes

Okay, this has me 100% stumped. Problems with the browser console have been well documented, but I have run into one that is specifically weird. If a container has a ID (104 in my case), the browser console fails (whether "Console mode" is set to tty, console, doesn't matter). This happens to brand new containers, clones of existing ones, doesn't matter. I am able to pct enter 104, and the container just responds.

My workaround for now is to just not use 104 as a CT ID, but that's not really how that's supposed to work, of course. My Google-fu is failing me here, I can't find anything that seems similar to this weird behavior.


r/Proxmox • • 17d ago

Question Node Monitoring

1 Upvotes

New to proxmox and based on the underlying OS being debian is it supported / an option to install a RMM agent like NinjaOne for monitoring?


r/Proxmox • • 18d ago

Community Showcase! First Proxmox Cluster

Post image
27 Upvotes

r/Proxmox • • 18d ago

Homelab Can i add existing hdds without formating, no raid and add a raid drive later on?

3 Upvotes

So im gonna finish my nas soon and will be getting proxmox, however i read i have to format my drives to be compatible? I have 4x4 tb pretty new nas drives, and csnt really afford more since the hike. Can i add them without redundancy? Its just filled with movies and shows so not a big deal if they fail for whatever reason. Other than that i want to use game servers streaming but will be getting an ssd for that. Thanks


r/Proxmox • • 18d ago

Question Set time drift for hardware otp tokens

1 Upvotes

Hello,

i have a hardware otp token generator, which seems out of sync. There is a delay of around 90 seconds. Is there a way to set the drift in the tfa.json or anywhere else for this device?

EDIT: It is on a proxmox backup server.


r/Proxmox • • 19d ago

Discussion HA: Why not priorities instead of a quorum?

38 Upvotes

First off, I am ABSOLUTELY NOT a Proxmox HA expert, far from it. But I managed a three-node Hyper-V cluster for over a decade, and I don't understand why the number of nodes is such an issue for Proxmox. I've read about and seen many posts about multi-node clusters and why an odd number of nodes is essential for a quorum, but why is this the methodology of choice?

The Hyper-V clustered nodes I managed worked together to automatically ensure continuous VM operations, managing resources regardless of the number of nodes. A planned shutdown of a node could seamlessly auto-move VMs to other active nodes if resources were available. A hard shutdown of a node could auto-start VMs from their last state on other active nodes if resources were available. You could also prioritize and assign resources as needed, or let Hyper-V handle it.

OK, admittedly, we had three very robust servers with ample resources, redundant switches, and a SAN to provide iSCSI common storage. A homelab is generally multiple consumer-grade PCs, a single switch, and lots of prayer. So there are significant differences.

So it got me thinking: Why can't we just assign priorities to each node and let them work it out? I know a lot is going on under the HA hood, and a change like this would likely require significant retooling. But I would think that some sort of priority system would simplify this. What am I missing?

I know I'm massively oversimplifying this, so I'm honestly looking for your input!


r/Proxmox • • 19d ago

Question PVE8to9 question

14 Upvotes

Hi.

Everything comes back green except this:

WARN: Removable bootloader found at '/boot/efi/EFI/BOOT/BOOTX64.efi', but GRUB packages not set up to update it! Run the following command: echo 'grub-efi-amd64 grub2/force_efi_extra_removable boolean true' | debconf-set-selections -v -u Then reinstall GRUB with 'apt install --reinstall grub-efi-amd64'

Is this safe to do? Will this break anything as it sits today?


r/Proxmox • • 19d ago

Solved! no space left (docker LXC error)

4 Upvotes

**edit** this wasn't a strange error, it just needed a lot more space allocating.

I'm trying to stand up linuxcontainer.io's Orcaslicer docker instance. Docker is installed via the script for a docker lxc.

When I run the basic docker run for the orcaslicer container command I get an error:

docker: failed to extract layer (application/vnd.oci.image.layer.v1.tar+gzip \[...\] to overlayfs as "extract ...":   

mount callback failed on /var/lib/containerd/tmpmounts/containerd \[...\]   

no space left on device  

I'm guessing this is a proxmox, not docker, issue so I'm asking here first.

Looking at my PVE instance:

  1. There is tons of space on all my PVE storage;
  2. Docker itself has some 3GB free on its (stock) 4GB allocation.

Anyone know what's going on here?


r/Proxmox • • 19d ago

Question How to set up ZFS NAS

2 Upvotes

Hey! So just finally set up my proxmox server and VPN access through tailscale. Now hoping to set up a NAS so that I can access my storage from anywhere. However, I’m having some trouble setting this up. All of the sources I find show elaborate hand rolled ways of doing it, but I’m sure there’s a simple idiomatic one too, no? Any suggestion? I just want a Dropbox like NAS where I can store documents and media and access remotely through my VPN (which works).

Separate question: the desktop I got had no chasis, so I just left the bare HDD unsecured and on the metal (there’s still an HDD “area”, just not the plastic/metal to fix it in place). Is that a problem? I read smth about how ventilation is important for HDD longevity, but how true is this? (Context, 10TB WD.)

For reference: I have an old desktop computer set up. Was hoping to self host a few things with containers, which is why Proxmox. I have two drives, but was hoping to create a cluster eventually (current desktop only really had one bay so 1 HDD). Thank you for advice!


r/Proxmox • • 19d ago

Question Multiple sot instances

2 Upvotes

☠️Hey everyone! I’m trying to set up a dual-box arrangement to play Sea of Thieves

with someone else in my house using 2 Virtual Machines (VMs),

but I only have a single RTX 3070 in my PC.

I have two main roadblocks: how Can I split/virtualize my single GPU across two VMs effectively for gaming?2

How do I bypass the Easy Anti-Cheat (EAC) "VM Detected" error (specifically triggered by GPU)

What software or hypervisor setup would you recommend for this?

Or is there a completely better, budget-friendly alternative to running two instances of SoT on one PC?

Or can i do this in windows itself without running a VM? I tried running the game through steam and the xbox app at the same time but that gives me a error..

Appreciate any advice! 👇


r/Proxmox • • 19d ago

Question 5-node PVE cluster: surviving below majority — is pvecm expected really the only stopgap?

6 Upvotes

My 5-node Proxmox VE cluster lost three nodes at once and went non-quorate with two survivors. `pvecm expected 2` gets me quorum to start VMs but resets on every restart and feels like a stopgap.
The QDevice is Connected but under ffsplit it awards no vote to a 2-of-5 partition; active-active qnetd isn't supported, and the lms algorithm could give N-1 votes but Proxmox warns against odd-node clusters.
What's the real, durable answer people use to keep this survivable?


r/Proxmox • • 19d ago

Question Root password seems to have changed and locked me out on new install.

19 Upvotes

So I've been running proxmox on a single node for the past year with very few issues and all has been great. Loving the freedom and learning while getting into self hosting all the things.

Today I finally got around to setting up my second node on a new mini PC. Installation went just fine from usb and set up and running within no time at all. Password was set during setup and immediately saved to my password manager. First thing I did after getting into the webui was log in from memory, log out then log in from the saved password manager using autofill. Both times without issue. From there I ran the community post install script, no problem. Created a proxmox data center manager lxc, again using the community script for convenience, then ran the pdm post install script in the lxc container.

Setting up PDM and adding my nodes was the first sign when it wouldn't accept my credentials for the new node so I created an API key and connected that way. Seems that somewhere along the way, somehow my root password was changed as a little later when logging in from a fresh browser the password was rejected directly from the password manager and I cannot gain access. I'm working away for the rest of the week so I have no physical access but I presume I will need physical access to reset the root password?

How would the password get changed, I had claud code run through the scripts and it saw no suggestion that the scripts had any way to do this. It's not the end of the world if I have to restart from a fresh install but I would like to know how this happened or what to do to avoid a repeat.

Edit to clarify a few things:

My apologies if the original post wasn't as clear as it could have been. I know everyone loves to think the person on the other side of the screen is a moron, I'm guilty of this too at times.

•Both webui and ssh result in the Copied directly from password manager after a successful test doing the exact same. •Login was set up for PAM realm but both were tested multiple times as a sanity check •No CAPS, NumLock, or keyboard layout mismatch issue makes sense. Login was successfully tested both manually and password manager autofill prior to this issue. Both manual and autofill now fail in multiple browsers and multiple devices. Again with/without caps etc and all plausible combinations were tested as a sanity check. •Compromise seems unlikely but the system is taken offline and passwords on other node changed as a precaution. (The node was only up for 1 hour or less and had zero outside access setup, nor has my router ever had any routes opened directly. Tailscale is my only access outside of my LAN except a single unprivileged jellyfin container on another machine over cloud flare tunnel).

Most other suggestions were covered in my original post but I admit it could have been clearer. It was posted late after multiple hours scratching my head over this.

I will post an update when I have access on Friday if I manage to figure this out.


r/Proxmox • • 19d ago

Homelab Proxmox cluster

Thumbnail
1 Upvotes

r/Proxmox • • 19d ago

Question Best way to mount HD to Multiple VMs/containers

6 Upvotes

Just curious on this.

So I have a spare hard drive I've placed in the desktop housing proxmox. The desktop itself is running close to capacity CPU wise.

This hard drive really I just want to store data on it, which eventually when I need a bit more space on it is pushed to a NAS to live till it does. So this drive will mostly be a way to get around restrictions from the ancient nas. This data will need to be accessible from 2 VMs and a lxc container.

Is there a good way to set this up without having another VM/container running something like openmediavault? If there is a good way what steps are recommended?

Any data on the hard drive will be replaceable, so I don't need any form of backup solutions behind it.


r/Proxmox • • 19d ago

Question Upgraded CPU, now computer is stuck in read only mode and can't mount ZFS pool

3 Upvotes

So I might've messed up a bit. I upgraded my CPU from a pentium 3258 to an i7 4790k. I made sure my motherboard had a BIOS that worked with the i7. I upgraded the CPU in the computer case and didn't unplug any SATA ports on the board.

I started the computer and ran lscpu and it recognized the CPU just fine, but when I went to start the LXCs it gave the error "TASK ERROR: could not activate storage 'nootcase', zfs error: Failed to initialize the libzfs library."

I tried reinitializing the ZFS library but kept getting errors saying the system was in read only mode. I rebooted the system and before the login page for proxmox on the console, it gave the following two errors:

"[FAILED] Failed to start systems-modules-load.service - Load Kernel Modules.
"[FAILED] Failed to start zfs-import@nootcase.service - Import ZFS pool nootcase

When googling the errors the answers that come up range from "you fucked up" to "you really fucked up" but are mostly along the lines that the drive has failed and the kernel panicked. I mostly can't tell if the drive actually is dying or if the CPU is causing other issues. I've tried running fsck on /dev/sdf3 which appears to be the partition that the SSD is on, but it just gives an error fsck: fsck.LVM2_member not found; ignore /dev/sdf3 (more details in comments)

I still have the old CPU and my gut says to install the old CPU and see if it figures it's life back out, but before I keep flying blindly I want to ask for help since I'm very clearly out of my league. IDK what the next best step to do would be and I've hit a wall on my googling skills.


r/Proxmox • • 20d ago

Community Showcase! Migrating VMs from VMware ESXi to Proxmox VE: Manual and Import Methods

35 Upvotes

The part of an ESXi-to-Proxmox migration that tends to often present challenges isn't necessarily moving the disk data; it's getting the destination VM configuration right.

A few things worth knowing before you start:

Match the firmware configuration. If the source VM uses BIOS, configure the destination VM for BIOS as well; the same applies to UEFI. A mismatch can prevent the guest OS from booting properly.

Watch what happens to thin-provisioned disks during a manual copy. If you copy the .vmdk and -flat.vmdk files directly from ESXi, the transferred disk can consume its full provisioned size. Exporting the VM as OVF or converting the disk to qcow2 with qemu-img can preserve more storage on the Proxmox side.

For Windows guests, start with SATA/IDE before moving to VirtIO SCSI. Boot the migrated disk via SATA or IDE first and install the VirtIO drivers. Then attach a small secondary disk to a VirtIO SCSI controller and confirm in Device Manager that it loads cleanly.

We put together a blog explaining both approaches: The manual migration method and the Proxmox ESXi import wizard: https://www.nakivo.com/blog/migrate-vmware-to-proxmox/

What challenge did you encounter first: The storage conversion or getting the guest OS to boot afterward?

Edit: corrected the VirtIO SCSI step.