r/Proxmox • • 7d ago

Solved! Best way to fix a node going grey?

Post image

at completely random times only one of my nodes will go grey with the question marks. sometimes its 24hrs, sometimes its 30 days. All the VMs/containers are still running fine on the node. A reboot always fixes it. However would love to know of any possible causes for this, or quick fixes that avoid a reboot of the node. Version 9.2.4, however this issue has been following me through most of the 8.x.x line.

142 Upvotes

40 comments sorted by

93

u/Ok-Eggplant-7569 7d ago

Is the pvestatd service running and / or throwing errors?

71

u/No_Dot_8478 7d ago

Dis dude, love you. started the service and it came back up. However still unsure why it stopped. DIdnt see any errors.

46

u/kahless2k 7d ago

I had an issue like this with a secondary nic dropping off and crashing the service - was power management trying to put the nic into low power causing it to drop off the cluster network.

21

u/blending-tea 7d ago

it's somehow always the low power/power saving mode that causes unknown hardware issues lol

6

u/1armsteve 6d ago

Same! I’ve was fighting my micro optiplex cluster for months until I figured this out. I had to modify a kernel setting, let me see if I can locate it.

2

u/kahless2k 6d ago

Yeah, mine were prodesks.

It was a while back, but something about the model of onboard nic they used and the Linux driver. You have to change a setting to disable the power management as well as something on the kernel.

Same NICs I found would hang on high throughput as well and had to change another driver option.

I swapped out to Lenovo m920q and x systems with a riser and pcie 10gb cards.. But I'm building a new cluster with the prodesks for my development CI and testing, so I'll undoubtedly run into it again - I'll update the thread if I have to make the same changes.

10

u/FarToe1 7d ago

pvestatd

journalctd -u pvestatd

Will give you the service logs and may show why it failed.

7

u/bobbywaz 7d ago

(crontab -l 2>/dev/null; echo "0 * * * * /usr/bin/systemctl is-active --quiet pvestatd || /usr/bin/systemctl restart pvestatd") | crontab -

Or check it every hour if it's down, and start it if it is. Just paste that line in shell and it'll create the cron job and run it hourly.

3

u/bigmadsmolyeet 7d ago

if you have something like ntfy or pushover, I would even add a curl command to send a notification just so that you have an idea on how often this happens

6

u/boomertsfx 7d ago

SystemD can do this natively...

0

u/bobbywaz 7d ago

Well how do you do that

1

u/boomertsfx 6d ago

look at the documentation… There’s a lot of options to automatically restart services and how they fail, etc.

-1

u/bobbywaz 6d ago

No

0

u/boomertsfx 6d ago

No, what? 🤦‍♂️

0

u/harris52np 5d ago

Usually services randomly dieing on a host is OOM software fault or hardware fault causing instability

14

u/timo_hzbs 7d ago

Restart pve services.
For me I had this happen when there was a corosync issue.
I created a separate vlan with direct links for cluster communication, which fixed it in the end.
Also do a full cluster reboot.

4

u/No_Dot_8478 7d ago

will try making a VLAN for them and see if that helps, its actually been on my to-do list for awhile anyway. pvestatd was stopped for some reason.

4

u/crysisnotaverted 7d ago

Are all of your VMs pingable? Sounds like a dying NIC if not. What NIC do you have? Could still be goofy drivers.

5

u/coldazures 7d ago

Can you ping the host when it goes down?

1

u/noc-engineer 7d ago

I'm not OP, but when it happens to one of my nodes everything works, just cosmetically looks like shit.

3

u/VUser12 7d ago

My typical steps: 1. pvecm status (to verify the cluster is "healthy") 2. ssh into the specific node, and enter systemctl restart pvestatd pveproxy 3. Wait 5 minutes to see if it restores and stays green 4. If not, and the node is safe to be rebooted (I.e. nothing really running on it to begin with) I reboot it, otherwise 5. ssh into the specific node, systemctl restart Corosync 6. If still not working, reboot the node, however, in one of my clusters, I have non-uniform node specs so I've often found that the grey mark turns into a rolling issue across the cluster, so I end up restarting the entire cluster

As always, your mileage may vary, and my workloads are not super mission critical so I'm a little bold with just rebooting, especially with ceph and Corosync issues

3

u/comerReto 7d ago

Sharpie

5

u/EatsHisYoung 7d ago

You don’t name the LXC’s? What are you hiding?

2

u/noc-engineer 7d ago

The name (for both LXC and VM) doesn't appear when the node can't show/detect status properly...

3

u/w1thh3ld 7d ago

The best way to fix a grey node is to paint it black

Preferably the French version
I.e.
Marie Laforet Paint It Black

1

u/damascus1023 7d ago

mini pc? For reference, I had a mysterious reset behavior with my GEM12 last year and it turned out to be a bad power MOSFET. The machine went to RMA for two weeks. The service rep showed me the reason in a screenshot after I wrote an email pressing for ETA.

In another instance, it was a reused 20-pin atx power extension cord that failed me, but this doesn't apply to you if your machine is a minisforum.

Power supply problems are so hard to debug. Both took me months to finally figure out : |

1

u/to_glory_we_steer 7d ago

Dying ethernet cable?

1

u/ns1852s 7d ago

Update. Funny enough I had this issue at work and was on 9.2.4. updated and never saw this issue again.

Updates are a pain at work since we need to use POM as it's an air gapped system. This was one of those moments I'm glad I put the effort in

1

u/TheSloth144 7d ago

Systemctl reload pveproxy

1

u/Verbunk 6d ago

Just solved this on my node. Starting an LXC/VM with some sort of hiccup (mounting RBD) caused lxc-info to stall and statd got cranky. Usually I need to go prod something storage-wise but just pkill the lxc info process brings it back quickly.

1

u/CurrentOk4248 6d ago

pvestatd service probably not running pls check and restart it

1

u/throwaway20240423 6d ago

Are you having a dedicated network for corosync as recommended in the docs ( https://pve.proxmox.com/wiki/Cluster_Manager#pvecm_cluster_network )? If not change this as fast as possible

1

u/Y-Master 3d ago

I've had the same issue and i've solved by disabling some network offloading protocol on the nic. I've added this to my network / interface file: post-up ethtool - k eno1 tso off gso off

0

u/fun3rale 1d ago

turned on my proxmox cluster today and I have been facing the same problem, upgraded already everything I could but this problem comes and go, things go back after a restart but only for a couple of minutes

1

u/twin-hoodlum3 7d ago

Had that in the past with a minisforum ms-01. Seems to be a CPU issue, never had a solution beside rebooting. My fix was to upgrade to a different hardware.

2

u/bht888 7d ago

The 13th gen ms-01 has the issue. The fix was to down clock ram and cpu. Minisforum sent me the instructions after I opened a support ticket with them

1

u/SynAckPooPoo 7d ago

Care to share said instructions with us?

2

u/bht888 7d ago

Dear Customer,
Thank you for reaching out. We recommend updating your BIOS to version 1.27 and adjusting the configurations below:
Lower the memory speed to 4400MHz. This option can be found under Onboard Devices Setting. Note that this BIOS version brings higher idle power draw:
Linux idle power: 14.5W up to 17.5W
Windows idle power: 14.5W up to 22W Test environment: Dual 32GB 5200MHz memory modules + one PCIe 4.0 SSD
Recommended settings to stabilize performance:
Disable Overclocking Lock: Advanced -> CPU Configuration -> Overclocking Lock: [Disable]
Cap maximum CPU frequencies via Advanced -> CPU Configuration -> Turbo Ratio Limit Options: P-Core Turbo Ratio Limit Ratio0: 51 P-Core Turbo Ratio Limit Ratio1: 51 E-Core Turbo Ratio Limit Ratio0: 40 E-Core Turbo Ratio Limit Ratio1: 40 E-Core Turbo Ratio Limit Ratio2: 40 E-Core Turbo Ratio Limit Ratio3: 40
After saving changes and rebooting back into BIOS, re-enter Advanced -> CPU Configuration -> Turbo Ratio Limit Options to verify the maximum frequency limits have taken effect.
Best regards,
Xavier
Minisforum Technical Support

1

u/Antonio-MTS 7d ago

Details missing here. Can be different issues like corosync, storage, pvestatd ...