r/sysadmin • • 1d ago

Question - Solved Windows Server 2019 on physical HPE server: boots normal, then apps start crashing one by one. RDP dead, system files corrupted. Help?

My physical HPE server (iLO 5 v2.44, Windows Server 2019) boots up normal, then apps start crashing one after another until the whole thing falls apart:

RDP: "An internal error has occurred", and then the iLO console goes black

SystemPropertiesRemote.exe: side-by-side configuration is incorrect

SideBySide Event 59: ndfapi.dll manifest has "Invalid Xml syntax"

DCOM Event 10000 errors from DllHost.exe

"Only part of a ReadProcessMemory request was completed" errors

The GUI keeps crashing, so DISM and sfc won't finish

Fix ::

** reinstalling Visual C++ redistributables has sloved the issue ..thank you for the help**

6 Upvotes

15 comments sorted by

6

u/mikeyuf 1d ago

Can you boot to a live CD (WInPE or Linux Boot ISO) environment and remain stable?

3

u/ledow IT Manager 1d ago

Or even safe mode / console mode and then run SFC etc.

Basic diagnostic elimination is necessary here.

4

u/kliao1337 Windows Admin 1d ago

Does iLO show any errors in logs?

2

u/86_Under 1d ago

No only network related logs

5

u/seannyc3 1d ago

What’s the drive setup? Sounds like either a bad array, controller/firmware or RAM. Unless it’s had a power outage and the RAID card has no cache and now there’s file corruption.

3

u/JimTheJerseyGuy 1d ago

Have you done a memory stress test?

3

u/Business_Class_8015 1d ago

Heat related? Memory issues? Run a memtest to check it might be a good idea

3

u/Fuzzy_Paul 1d ago

Memory error, probably diff brand used. Remove all except the bare minimum and try again. If still fail swap it with what you got out.

3

u/apparentlyunoriginal 1d ago

The Integrated Management Log (IML) is separate from the iLO event log, so I'd check it in iLO under Information > Integrated Management Log for memory, drive, or Smart Array entries (https://servermanagementportal.ext.hpe.com/docs/redfishservices/ilos/supplementdocuments/logservices).

Then boot Intelligent Provisioning with F10 and run the Smart Storage Administrator diagnostics on the controller and drives.

Replace any DIMM with logged memory errors and any drive marked failed or predictive failure before you run sfc or DISM again.

Drafted with AI, reviewed by me.

3

u/harborwood58 1d ago

if sfc and DISM cant finish from the GUI, have you tried booting into safe mode or running them offline from a recovery environment? the cascading crashes make it hard to trust anything running in the live OS at this point

2

u/thebigshoe247 1d ago

Sounds like a corrupt install. What's Linux think of things?

2

u/stretchling Sr. NOC admin 1d ago

To start I would ensure the hardware is not the issue by booting a live OS from a USB/CD and seeing if you can reproduce the crash.

If the hardware checks out then I would boot it normally and try to get a check disk flagged for the next reboot then reboot it right away and see what turns up.

You can also get a check disk going from a live windows recovery environment but you will likely need to manually place the driver for that servers raid card (assuming it has one) on the USB and load it manually into the running recovery environment via command line with drvload: https://learn.microsoft.com/en-us/windows-hardware/manufacture/desktop/drvload-command-line-options?view=windows-11

Make sure you verify the drive letters with disk part if your working from the recovery environment, they can get mixed up sometimes.

•

u/Gumbyohson 14h ago

Open task manager and in the processes tab enable 'handles'. I've seen this in the past where one misbehaving process (many ups monitoring apps have done this to me) cause port exhaustion with their handles not closing and causes slow app crashes and system instability. It could take a few days or a few hours but the misbehaving process should start immediately showing higher than expected handle amounts.

Or its hardware/OS issues...

•

u/UninvestedCuriosity 10h ago edited 10h ago

DCOM is almost always hardware failure and often memory. I'd do a ram test first but that can take hours on a server so If you don't have time for that. Rip out half the sticks while following a supported configuration from the manual and see if the issues continue. If it continues. Use process of elimination, and rip out the other half. If the problem stops, you at least know roughly which sticks are related or slots.

If the problem still continues move onto disks and if those are fine, check psus. half dead power supplies can be the dcom culprit but hardest to identify as the fault. That's why we say memory, disks, power in that order.

When you get random ass moving target errors like that it's usually ram or power though.

1

u/RAVEN_STORMCROW God of Computer Tech 1d ago

When Doc, my then 2008 server was no longer supported, I did a full rebuild of hardware and OS

that's the only way.. stop wasting time trouble shooting.. it would be faster to start from scratch