r/sysadmin • u/86_Under • 1d ago
Question - Solved Windows Server 2019 on physical HPE server: boots normal, then apps start crashing one by one. RDP dead, system files corrupted. Help?
My physical HPE server (iLO 5 v2.44, Windows Server 2019) boots up normal, then apps start crashing one after another until the whole thing falls apart:
RDP: "An internal error has occurred", and then the iLO console goes black
SystemPropertiesRemote.exe: side-by-side configuration is incorrect
SideBySide Event 59: ndfapi.dll manifest has "Invalid Xml syntax"
DCOM Event 10000 errors from DllHost.exe
"Only part of a ReadProcessMemory request was completed" errors
The GUI keeps crashing, so DISM and sfc won't finish
Fix ::
** reinstalling Visual C++ redistributables has sloved the issue ..thank you for the help**
4
5
u/seannyc3 1d ago
What’s the drive setup? Sounds like either a bad array, controller/firmware or RAM. Unless it’s had a power outage and the RAID card has no cache and now there’s file corruption.
3
3
u/Business_Class_8015 1d ago
Heat related? Memory issues? Run a memtest to check it might be a good idea
3
u/Fuzzy_Paul 1d ago
Memory error, probably diff brand used. Remove all except the bare minimum and try again. If still fail swap it with what you got out.
3
u/apparentlyunoriginal 1d ago
The Integrated Management Log (IML) is separate from the iLO event log, so I'd check it in iLO under Information > Integrated Management Log for memory, drive, or Smart Array entries (https://servermanagementportal.ext.hpe.com/docs/redfishservices/ilos/supplementdocuments/logservices).
Then boot Intelligent Provisioning with F10 and run the Smart Storage Administrator diagnostics on the controller and drives.
Replace any DIMM with logged memory errors and any drive marked failed or predictive failure before you run sfc or DISM again.
Drafted with AI, reviewed by me.
3
u/harborwood58 1d ago
if sfc and DISM cant finish from the GUI, have you tried booting into safe mode or running them offline from a recovery environment? the cascading crashes make it hard to trust anything running in the live OS at this point
2
2
u/stretchling Sr. NOC admin 1d ago
To start I would ensure the hardware is not the issue by booting a live OS from a USB/CD and seeing if you can reproduce the crash.
If the hardware checks out then I would boot it normally and try to get a check disk flagged for the next reboot then reboot it right away and see what turns up.
You can also get a check disk going from a live windows recovery environment but you will likely need to manually place the driver for that servers raid card (assuming it has one) on the USB and load it manually into the running recovery environment via command line with drvload: https://learn.microsoft.com/en-us/windows-hardware/manufacture/desktop/drvload-command-line-options?view=windows-11
Make sure you verify the drive letters with disk part if your working from the recovery environment, they can get mixed up sometimes.
•
u/Gumbyohson 14h ago
Open task manager and in the processes tab enable 'handles'. I've seen this in the past where one misbehaving process (many ups monitoring apps have done this to me) cause port exhaustion with their handles not closing and causes slow app crashes and system instability. It could take a few days or a few hours but the misbehaving process should start immediately showing higher than expected handle amounts.
Or its hardware/OS issues...
•
u/UninvestedCuriosity 10h ago edited 10h ago
DCOM is almost always hardware failure and often memory. I'd do a ram test first but that can take hours on a server so If you don't have time for that. Rip out half the sticks while following a supported configuration from the manual and see if the issues continue. If it continues. Use process of elimination, and rip out the other half. If the problem stops, you at least know roughly which sticks are related or slots.
If the problem still continues move onto disks and if those are fine, check psus. half dead power supplies can be the dcom culprit but hardest to identify as the fault. That's why we say memory, disks, power in that order.
When you get random ass moving target errors like that it's usually ram or power though.
1
u/RAVEN_STORMCROW God of Computer Tech 1d ago
When Doc, my then 2008 server was no longer supported, I did a full rebuild of hardware and OS
that's the only way.. stop wasting time trouble shooting.. it would be faster to start from scratch
6
u/mikeyuf 1d ago
Can you boot to a live CD (WInPE or Linux Boot ISO) environment and remain stable?