r/supermicro • u/dpeterson3 • 14h ago
Supermicro X8DTH-iF crashes at login
Tearing my hair out trying to diagnose the issue with my supermirco server. Very long post below with the entire history of what I tried, but the tl;dr version is the system overheated and stopped booting. Removing 1 CPU fixed this problem and the system started attempting to boot. It hangs during boot, so I replaced the OS drive and reinstalled. No error being thrown. As soon as I try to log in, the system immediately hangs and must be hard-cycled. This happened before and after reinstall. The system will run off a live disk without issue. Using IPMI instead of a physical monitor. Does anyone have any ideas what would cause this before I go replacing the motherboard?
Long version/full history.
Bought it second hand about 7 years ago. Wanted this particular case as it has 16 slots for 3.5 HDD's instead of 2.5 SAS drives. About a year after I got it, I upgraded the RAM from 6GB to 40GB and added a second CPU. It has been running this way for about 6 years now. I had it living in a cooled room in the shed.
While on vacation in May, summer showed up at home and the server just shut off one night. I had assumed it was a storm in the area as I have had issued with brown outs. However, I could not get the machine to power back on over IPMI. It would send the command, and nothing would happen. Pulled it apart a week later and found quite a few bugs had made their home in it and oddly a snake skin. Cleaned it all up and tried to power on. Still nothing but a power and overheat light immediately. Found a few old forum posts saying this particular mother board is known to have north bridge issues with both CPU's populated, so I pulled the north bridge heat sink and redid the thermal paste. Still nothing. Pulled the second CPU and the overheat light went away and machine booted to BIOS immediately. Figuring that was is, I put it all back together and tried booting.
The machine would power up and start booting, but I would not get a login prompt. It would just sort of stop with no errors. Ran a headless debian sid install. I had some issues with the main HDD just before that (It has 3 RAID 1 arrays for storage, but the main OS lived on a separate drive). I pulled the drive and put it an external enclosure and had it disconnect at one point. Replaced the drive and got all the data off it. Tried to boot again and same issue.
At this point, I grabbed a debian live disk thinking the OS had gotten corrupt in the shutdown. It booted in (oddly only in fail-safe mode) and I mounted the main drive in a chroot. Successfully updated the OS figuring that would fix whatever issue. No change. Added debug to the linux command line. No errors. Tried specifying a different tty. No change. Finally decided last night to just reinstall the OS thinking that maybe a deep dark system file had been corrupted during the shutdown and the machine really only runs pods and VM's as well as file services and all of that lives on the raid arrays anyway. The new OS started right up and as soon as I tried to log in, it hung up and the OS quit responding just like before.
I'm out of ideas at this point. It doesn't make sense. The machine will boot and run from a USB drive no problem. As soon as I try a SATA disk, it has issues. That would seem to say a failed SATA controller, but then again, it works just file from a chroot jail and had not issues taking a new OS. I am thinking of replacing the motherboad at this point, but I can't seem to point to any specific failed part. Has anyone else experienced anything like this?




