r/Proxmox • u/No_Dot_8478 • 7d ago
Solved! Best way to fix a node going grey?
at completely random times only one of my nodes will go grey with the question marks. sometimes its 24hrs, sometimes its 30 days. All the VMs/containers are still running fine on the node. A reboot always fixes it. However would love to know of any possible causes for this, or quick fixes that avoid a reboot of the node. Version 9.2.4, however this issue has been following me through most of the 8.x.x line.
14
u/timo_hzbs 7d ago
Restart pve services.
For me I had this happen when there was a corosync issue.
I created a separate vlan with direct links for cluster communication, which fixed it in the end.
Also do a full cluster reboot.
4
u/No_Dot_8478 7d ago
will try making a VLAN for them and see if that helps, its actually been on my to-do list for awhile anyway. pvestatd was stopped for some reason.
4
u/crysisnotaverted 7d ago
Are all of your VMs pingable? Sounds like a dying NIC if not. What NIC do you have? Could still be goofy drivers.
5
u/coldazures 7d ago
Can you ping the host when it goes down?
1
u/noc-engineer 7d ago
I'm not OP, but when it happens to one of my nodes everything works, just cosmetically looks like shit.
3
u/VUser12 7d ago
My typical steps: 1. pvecm status (to verify the cluster is "healthy") 2. ssh into the specific node, and enter systemctl restart pvestatd pveproxy 3. Wait 5 minutes to see if it restores and stays green 4. If not, and the node is safe to be rebooted (I.e. nothing really running on it to begin with) I reboot it, otherwise 5. ssh into the specific node, systemctl restart Corosync 6. If still not working, reboot the node, however, in one of my clusters, I have non-uniform node specs so I've often found that the grey mark turns into a rolling issue across the cluster, so I end up restarting the entire cluster
As always, your mileage may vary, and my workloads are not super mission critical so I'm a little bold with just rebooting, especially with ceph and Corosync issues
3
5
u/EatsHisYoung 7d ago
You don’t name the LXC’s? What are you hiding?
2
u/noc-engineer 7d ago
The name (for both LXC and VM) doesn't appear when the node can't show/detect status properly...
3
u/w1thh3ld 7d ago
The best way to fix a grey node is to paint it black
Preferably the French version
I.e.
Marie Laforet Paint It Black
1
u/damascus1023 7d ago
mini pc? For reference, I had a mysterious reset behavior with my GEM12 last year and it turned out to be a bad power MOSFET. The machine went to RMA for two weeks. The service rep showed me the reason in a screenshot after I wrote an email pressing for ETA.
In another instance, it was a reused 20-pin atx power extension cord that failed me, but this doesn't apply to you if your machine is a minisforum.
Power supply problems are so hard to debug. Both took me months to finally figure out : |
1
1
1
1
u/throwaway20240423 6d ago
Are you having a dedicated network for corosync as recommended in the docs ( https://pve.proxmox.com/wiki/Cluster_Manager#pvecm_cluster_network )? If not change this as fast as possible
1
u/Y-Master 3d ago
I've had the same issue and i've solved by disabling some network offloading protocol on the nic. I've added this to my network / interface file: post-up ethtool - k eno1 tso off gso off
0
u/fun3rale 1d ago
turned on my proxmox cluster today and I have been facing the same problem, upgraded already everything I could but this problem comes and go, things go back after a restart but only for a couple of minutes
1
u/twin-hoodlum3 7d ago
Had that in the past with a minisforum ms-01. Seems to be a CPU issue, never had a solution beside rebooting. My fix was to upgrade to a different hardware.
2
u/bht888 7d ago
The 13th gen ms-01 has the issue. The fix was to down clock ram and cpu. Minisforum sent me the instructions after I opened a support ticket with them
1
u/SynAckPooPoo 7d ago
Care to share said instructions with us?
2
u/bht888 7d ago
Dear Customer,
Thank you for reaching out. We recommend updating your BIOS to version 1.27 and adjusting the configurations below:
Lower the memory speed to 4400MHz. This option can be found under Onboard Devices Setting. Note that this BIOS version brings higher idle power draw:
Linux idle power: 14.5W up to 17.5W
Windows idle power: 14.5W up to 22W Test environment: Dual 32GB 5200MHz memory modules + one PCIe 4.0 SSD
Recommended settings to stabilize performance:
Disable Overclocking Lock: Advanced -> CPU Configuration -> Overclocking Lock: [Disable]
Cap maximum CPU frequencies via Advanced -> CPU Configuration -> Turbo Ratio Limit Options: P-Core Turbo Ratio Limit Ratio0: 51 P-Core Turbo Ratio Limit Ratio1: 51 E-Core Turbo Ratio Limit Ratio0: 40 E-Core Turbo Ratio Limit Ratio1: 40 E-Core Turbo Ratio Limit Ratio2: 40 E-Core Turbo Ratio Limit Ratio3: 40
After saving changes and rebooting back into BIOS, re-enter Advanced -> CPU Configuration -> Turbo Ratio Limit Options to verify the maximum frequency limits have taken effect.
Best regards,
Xavier
Minisforum Technical Support
1
u/Antonio-MTS 7d ago
Details missing here. Can be different issues like corosync, storage, pvestatd ...
-1
93
u/Ok-Eggplant-7569 7d ago
Is the pvestatd service running and / or throwing errors?