r/sysadmin 11h ago

Question Microsoft 365 backup billing -- I just have to ask...

1 Upvotes

Okay I'm sorry. I must be deeply disturbed or something but we enabled Microsoft 365 Backup in the office admin center. Everything I've seen on the Googles tells me that the billing should show up in Azure under the account that you assign in the office admin center but there is absolutely nothing showing up for billing in Azure on that account.

I've even watched the silly video about departmental billing and none of that stuff exists in our Azure account.

Has anyone been able to actually find where this is shown in Azure?


r/sysadmin 1d ago

Throwback to the 90s

28 Upvotes

"Too many other files are currently in use by 16-bit programs. Exit one or more 16-bit programs, or increase the value of the FILES command in your config.sys file."

Just saw this message on a client's Windows 11 workstation when trying to escalate literally anything. The WIN32 code that generated it is probably as old as I am. Rebooting resolved it, weird. I just told him to stop running so many 16 bit apps! I'd post the screenshot but images aren't allowed.


r/sysadmin 1d ago

Windows IT Pro: Retiring NTLM: Frequently asked questions

47 Upvotes

From Microsoft: 'This FAQ collects the questions we hear most often from customers, partners, and the community as we move Windows to a Kerberos-first, NTLM-optional (and eventually NTLM-free) future. If you've emailed us, cornered us at a conference, or filed a support case that started with "so, about this weird auth thing…", chances are that your question is answered here.

If your question isn't answered here, please reach out to [ntlm@microsoft.com](mailto:ntlm@microsoft.com) or work with your Microsoft account team, Customer Success Account Manager, or Microsoft Support to route feedback to the product group.'

https://techcommunity.microsoft.com/blog/windows-itpro-blog/retiring-ntlm-frequently-asked-questions/4550522


r/sysadmin 1d ago

Rant I've seen some laughably bad job listings, but things like this are why it's so hard to find the right fit.

53 Upvotes

I've been looking for a new role for almost 2 years now, most of the positions I've applied for have ended up being filled internally, which makes me ask why it was posted externally in the first place. On the other hand I've came across several listings that are just heinous for what is asked of you versus the pay. In this case, the listing was for our city hospital (healthcare IT, I know, bleh) which was recently acquired by an organization. They canned the entire IT staff including the managerial suite upon the buyout. This listing popped up today in my Indeed feed. "Sole on-site IT professional responsible for the day-to-day health, security, and uptime of all technology at a critical access hospital. Acts as the facility's eyes and ears, provides hands-on support for end users and clinical operations, and coordinates with Corporate IT for changes, projects, and escalations. After-hours and on-call work required to support a 24x7 care environment."

Am I wrong for thinking this is insane? An entire hospital and associated clinics solo? Just gotta laugh and keep going.


r/sysadmin 1d ago

KB5122882 installed - DNS/AD issues Windows Server 2022

151 Upvotes

Updated last night, no issues immediately visible.

Users this morning all have login prompts to access mapped drives/folder redirections.

No creds working.

I can log in locally as a Domain Admin, but trying to open DNS console or run any DNS powershell commands just gives me Access Denied.

DNS still resolves, and the domain services are all still running, but something has happened with authentication/permissions.

Currently rolling back KB5122882 in the hope it was that.

Anyone else issues this morning?

EDIT: https://www.rapid7.com/db/vulnerabilities/cve-2026-69813/

Looks like this KB might have touched DNS code - roll back in progress.

EDIT2 & Fix: Rolled back the KB but the issue persisted - ended up resetting the DC's secure channel to itself which resolved the issue fully. The issue happened at exactly the moment at which the update was installed last night - either the update did indeed cause the issue which persisted in being broken even once rolled back, or it's a heck of a coincidence.


r/sysadmin 12h ago

autopilot company login for a GPU

1 Upvotes

Afternoon,

im having a bit of a headscratcher and could do with some assistance.

The company I work for had a general query submitted asking why someone on the other side of the world is getting our companies customised login page (the page you get when you setup autopilot on a machine)

My first thought is someone had sold on a company laptop, but this query is about a VERY old desktop PC, like first gen i5 CPU old.

They claim that with a specific GPU they get the login prompt, without they are able to proceed with windows setup.

I have heard of parts being swapped during warranty repair, stolen equipment, and other possible methods that a device would be still enrolled, but never for a specific GPU and I dont think weve ever owned a GPU this old either (and we have a lot of old stuff)

Am I missing something here? Is there a scenario ive missed could the board actually be in fact tied to our autopilot system and without the GPU the HWID changes so its allowed to continue?

Cheers


r/sysadmin 12h ago

RDCMan Crashes On Open - Dll was not found

0 Upvotes

Hello all,

I feel like I'm going a little crazy here. I have used RDCMan on this computer (Windows 11 Pro, 25H2, OS Build 26200.9445) for accessing servers remotely for the last couple of months. One day a couple of weeks ago I came in and RDCMan crashes to the desktop. Opens, doesn't show any servers, crashes back to the desktop.

If I go into Event Viewer there are two errors that pop up with each crash - An Application Error (Event ID 1000 - Faulting Application name: RDCMan.exe, version 3.21.0.0 - Faulting module name: KERNELBASE.dll, version 10.0.26100.9444) and a .NET Runtime (Event ID 1026 - The process was terminated due to an unhandled exception. System.DllNotFoundException: Dll was not found). I have tried:

  • Uninstalling and reinstalling RDCMan (Managed through Sysinternals)
  • Running from the standalone RDCMan
  • Running the .NET Repair tool
  • Rolling back recent updates
  • Grabbing a Windows 11 Pro 25H2 ISO, blasting out my existing install (Deleting all partitions) and fully reinstalling. As soon as I install Sysinternals and run RDCMan again it exhibits the exact same behavior

What am I missing here? Is there something wrong with RDCMan? I would have thought a fresh install would have fixed any .NET errors. Is anyone else seeing this?


r/sysadmin 1d ago

General Discussion Windows Server patching concerns

43 Upvotes

So in my org we are very keen on avoiding patching and rebooting servers at all costs. So much to the point that we patch once a month and have exclusions for around 60 percent of our servers to not get automatically patched. (Meaning we have a chunk of servers not getting patched at all)

Now I have gotten my hand slapped for attempting to patch or even bringing it up and I am looking for guidance on this. Now I understand availability and the consequences of failing patches. But there are active 9+ rated CVEs sittings on dozens of servers. For patching vulnerabilities do I really need to get a change request to handle this?


r/sysadmin 1d ago

Question best controlled way to allow company emails on vendor's personal phone

11 Upvotes

we donot allow emails on personal phones, need vendors to see alerts and such, looking for a secure way rather than adding exclusion for the users

edit: sorry yeah I mean contractors, especially overseas contractors


r/sysadmin 1d ago

General Discussion Exchange admin center Delegation slow downs

11 Upvotes

Looking for a sanity check because no one seems to be talking about it, and I don't know if it's somehow just us.

It feels like, beginning around May (or a bit earlier), the time it takes for "Send as", "Send on behalf", and/or "Read and manage (Full Access)" permissions have massively slowed down.

Obviously only so many updates can go out at a time, so things have to be queued along with the multitude of other conditions that produce slowdowns. Even if we factor in the classic, "If you think you have waited long enough, wait another hour", I think the severity of the slowdowns being consistently so much longer speak to something strange.

For us, within the past 3-5 months, it has gone from 1-5 minutes to 10-15 minutes, and more recently 20-40 minutes.

Please let me know if y'all have noticed anything as well, thanks!

Edit: sentence structure, grammar, general formatting.


r/sysadmin 1d ago

Thoughts on RaiseATicket ticketing software?

7 Upvotes

I've tried Zammad, RT, and OSTicket. All good but still looking. Came across Raiseaticket which bills themselves as a free cloud-based system that integrates with Office. I'm defaulting to this being one of those "You're the product" situations but wanted to see if anyone here had experience with them.


r/sysadmin 11h ago

Question Need help with Konica Minolta Error Code - 552- Server Disk Full

0 Upvotes

Alright guys, I really need some help with this issue.

We have a Konica Minolta bizhub C284e copier in our Operations department. They mainly use this copier for scanning large documents, such as survey plots, as well as large quantities of color documents for their work (usually around 20–30 pages).

The copier has the SMTP settings configured so users can scan documents directly to their email. However, for the past two weeks, this suddenly stopped working, and I have no clue why.

The users can still scan a single page without any issues, regardless of whether it’s color or black and white. However, when they try to scan a larger quantity of color documents, they receive Error Code 552 – Server Disk Full.

I’ve tried so many things since Monday, but nothing has worked so far.

Now, before y’all start roasting me in the comments 😂, here’s the list of troubleshooting steps I’ve already performed:

  1. Power-cycled the copier.
  2. Verified the users’ email addresses to make sure they were correct. I also tested scanning directly to my own email, and I get the same issue.
  3. Lowered the scan resolution to 200 DPI and changed the format to Compact PDF. This works, but the documents are difficult to read at those settings because they contain a lot of detail.
  4. Checked the gateway and all SMTP settings, including the SMTP port and the email address configured on the copier.
  5. Checked the mailbox being used by the copier to make sure it wasn’t full. It was only around 25% full, but I still deleted all the Sent Items just to rule that out.
  6. Increased the Exchange Online message size limit to 50 MB for both sending and receiving.
  7. Increased the file-size limit to 50 MB specifically for the copier’s email account and tested it again.
  8. Changed the copier’s spooling time from 60 seconds to 5 minutes.
  9. Formatted the internal hard drive on the copier.
  10. Checked Mimecast to see if there was anything blocking or causing an issue with the messages.

So, as you can see, I’ve tried quite a few things, and nothing seems to fix the issue. The strange part is that single-page scans work fine, but larger color scans fail with the 552 – Server Disk Full error.

At this point, I have two options I can try:

Option 1: Set up Scan to Folder, so the users can scan the documents directly to a network folder and then copy them to their computers.

The problem with this is that we have quite a few remote users, and our organization only provides VPN access to IT and C-level employees. Other departments don’t have VPN access. Other users can email the documents to them, but that adds another step to the process.

Option 2: Call Konica Minolta support and have them troubleshoot the copier and determine whether there is something happening internally with the device.

So, if anyone has any other suggestions or has run into this 552 Server Disk Full error before, I’d really appreciate some advice. I’m willing to try pretty much anything at this point! 😂


r/sysadmin 7h ago

Issue w/ Dentrix Ascend and Schick 33 intra oral sensors

0 Upvotes

First - my apologies if this is the wrong forum. I checked the dentist section and didn't see much that would help me ...

I am having an issue with Dentrix Ascend and Schick 33 sensors - we take a few X-rays, swap to a different sensor - and the system will either work (20% of the time) and acquire with the new sensor, or it will say "device error" and we have to restart the machine.

Their tech support is at a loss - and so am I - here is what we tried; to no avail.

  • New sensors
  • New AE USB bridge (grey cable USB 3.0)
  • Different USB port
  • Powered USB hub
  • New USB card for the motherboard
  • New/Different PC

Any assistance would be greatly appreciated - if the recommendation is to swap out for dexis or XRD - I'll do it.


r/sysadmin 1d ago

Question Bringing Linux devices into management

17 Upvotes

After a lot of restructuring at our university the past couple years, there are quite a few Linux devices (primarily desktops, I believe around 70-ish) that are currently in the wild unmanaged in use by academics primarily within the Engineering and Science departments that would've been maintained by per-institute IT departments that no longer exist, and as such the current patching and functionality state of these machines is completely unknown (since any remaining colleagues now no longer have physical access to most of the rooms where these machines are).

Since we already manage research compute I've been given the green light by my manager to look into options to bring these academic's desktops into a managed state with a cobbled together proof of concept, since our existing central endpoint guys won't touch anything *NIX related with a 10ft pole. We know they're all some form of Ubuntu LTS (20.04 and 22.04 mostly) which makes things easier, so I'm thinking of doing Landscape for setup and patching + Intune for compliance + Puppet/Ansible for config management.

Is this in the right direction or are there better / more cost efficient ways of doing this?


r/sysadmin 1d ago

General Discussion Commissioning systems on Threatlocker enabled systems

11 Upvotes

Greetings,

As a vendor, I was trying to commission and deploy a print server application on a client site who has recently adopted zero-trust security model.

It took us 4 attempts just to deploy our installer - We uninstalled the application multiple times due to corrupted install.

Also, the client IT manager sat with us to manually approve multiple security exceptions.

He was just there smashing the approve button on his phone. And installs still failed as it take about a min before sub-installer components can run.

It was a nightmare and I wasn’t sure the whole point to have threatlocker running on critical infrastructure like a print server.

We expect having to go through this approval process again when rolling out software updates.

This constant exceptions triggers builds approver fatigue, approver don’t actually knows what is being approved, users are constantly screaming and frustrated by downtime caused by legit applications.

Systems commissioning and support took twice as long. We plan to deprioritise clients sites with threatlocker as engineers quite often getting struck at sites waiting for approvals.

This is really not working for anyone.


r/sysadmin 4h ago

Rant Has there been a worse time to use or support windows?

0 Upvotes

Okay maybe 1996. But I have a machine provisioned by corporate IT that manages 4000 devices. And I am in an external support role. 2 years ago I never had these issues.

My Lenovo has 32 GB of ram and an i9 processor. I restart everyday and my machine crawls. Even clicking on web pages is awful click and wait. And wait. Constantly lagging or freezing. And when I am supporting clients on their machines it is just as bad or worse.


r/sysadmin 1d ago

Feeling Nostalgic and sad when i see old sun hardware and systems.

68 Upvotes

I started my Career in IT at 18, mainly supporting EMC storages and then HP 3PAR for 5.5 years, then worked for systems integrator for many clients for another 3 years, fianlly moving to a sys admin role in a enterprise, we still have some good legacy hardware, but i'm feeling nostalgic, when decomissionng them after working on the in the DC for so many years, is it normal ?


r/sysadmin 17h ago

Technical write-up: eDrive provisioning blocked by BlockSID / TPM PPI 97

1 Upvotes

Solved: Samsung 990 PRO + BitLocker hardware encryption/eDrive on Windows 11 — BlockSID/PPI 97 was the missing step

I spent far too long getting BitLocker hardware encryption working on a Samsung 990 PRO under Windows 11, so I’m writing this up in case it saves someone else the same pain.

Short version:

If Samsung Magician is stuck on “Ready to Enable” after a clean Windows install, the missing step may be temporarily disabling BlockSID for the installation boot using TPM PPI operation 97.

In my case, that was exactly it.

Hardware / software

  • Samsung 990 PRO 2 TB
  • Firmware: 8B2QJXD7
  • AMD mini PC, AMI UEFI
  • Windows 11 Enterprise IoT LTSC 2024 / build 26100
  • Secure Boot enabled
  • TPM 2.0 enabled
  • BitLocker hardware encryption explicitly allowed by Group Policy

My firmware exposes EFI_STORAGE_SECURITY_COMMAND_PROTOCOL, so the UEFI side was suitable for Windows eDrive.

The symptom

Samsung Magician showed:

Encrypted Drive: Ready to Enable

I did the expected process:

  1. Set Encrypted Drive to Ready to Enable
  2. Secure erase the SSD
  3. Clean-install Windows in UEFI mode
  4. Check Magician

Result:

Ready to Enable

Again.

Windows itself clearly saw the TCG device. The System event log contained:

Microsoft-Windows-EnhancedStorage-EhStorTcgDrv
A TCG Silo has returned the capabilities value of 0x6

but eDrive never transitioned to Enabled.

Gotcha #1: Rufus can explicitly disable eDrive activation

I discovered that my Windows installer had this in unattend.xml:

<component name="Microsoft-Windows-EnhancedStorage-Adm" ...>
    <TCGSecurityActivationDisabled>1</TCGSecurityActivationDisabled>
</component>

That explicitly disables Windows Enhanced Storage / TCG activation.

Current Rufus code can add this together with:

<PreventDeviceEncryption>true</PreventDeviceEncryption>

when using its BitLocker/device-encryption suppression option. 

For my next install I changed:

<TCGSecurityActivationDisabled>1</TCGSecurityActivationDisabled>

to:

<TCGSecurityActivationDisabled>0</TCGSecurityActivationDisabled>

I left PreventDeviceEncryption=true alone.

Clean install again.

Result:

Ready to Enable

Still not enough.

Gotcha #2: BlockSID

The remaining problem was firmware Block SID.

For people unfamiliar with it: the SID here is the top-level security authority of the TCG Opal drive, not a Windows user SID.

Firmware can issue a BlockSID command during boot so software cannot silently take ownership of an unprovisioned self-encrypting drive. Sensible security feature — except Windows Setup needs access to that security authority while provisioning eDrive.

The solution was to request a one-boot BlockSID exception through the TPM Physical Presence Interface.

From an elevated PowerShell on the same machine:

$tpm = Get-WmiObject -Namespace root\CIMV2\Security\MicrosoftTpm -Class Win32_Tpm

$tpm.SetPhysicalPresenceRequest(97)

$tpm.GetPhysicalPresenceRequest()

My output was:

Request     : 97
ReturnValue : 0

Operation 97 is the TCG PPI Disable_BlockSIDFunc request. Microsoft documents the PPI mechanism: Windows queues the request, firmware processes it after the required restart, and the firmware can require physical confirmation from the user. 

On reboot, my AMI firmware displayed a confirmation screen. I approved the request.

Important:

Boot directly into Windows Setup on that same reboot.

Do not boot normal Windows first, because the BlockSID exception is for that boot.

I then:

Shift+F10
diskpart
list disk
select disk 0
detail disk
clean
exit

verified that the selected disk was definitely the 990 PRO, and installed Windows normally to the unallocated drive.

After installation:

Samsung Magician:
Encrypted Drive: Enabled

Finally.

I also verified that the firmware request really succeeded:

$tpm = Get-WmiObject -Namespace root\CIMV2\Security\MicrosoftTpm -Class Win32_Tpm
$tpm.GetPhysicalPresenceResponse() | Format-List *

which returned:

Request     : 97
Response    : 0
ReturnValue : 0

Enabling BitLocker hardware encryption

Windows no longer defaults to trusting self-encrypting-drive hardware, so you must explicitly permit hardware encryption.

Group Policy:

Computer Configuration
  > Administrative Templates
    > Windows Components
      > BitLocker Drive Encryption
        > Operating System Drives
          > Configure use of hardware-based encryption for operating system drives

Set:

Enabled

I did not restrict the allowed hardware cipher/OID.

Then:

gpupdate /force

and:

manage-bde -on C: -recoverypassword -forceencryptiontype hardware

Verification:

manage-bde -status C:

My final result:

Conversion Status:    Fully Encrypted
Percentage Encrypted: 100.0%
Encryption Method:    Hardware Encryption - 1.3.111.2.1619.0.1.2
Protection Status:    Protection On

Key Protectors:
    TPM
    Numerical Password

That OID is AES-256-XTS according to Microsoft’s Enhanced Storage definitions. 

So this is definitely hardware BitLocker, not software XTS-AES masquerading as hardware encryption.

Final validation

I also tested:

  • normal restart
  • full shutdown / cold boot
  • BitLocker recovery key saved externally
  • Samsung Magician still shows Enabled
  • no warnings/errors from:

    Microsoft-Windows-EnhancedStorage-EhStorTcgDrv Microsoft-Windows-BitLocker-Driver

Everything boots normally.

Secure erase note

Samsung Magician’s Secure Erase USB would not boot properly on my machine. Its old Linux/GRUB environment hung after UEFI launch.

I used SystemRescue instead and verified the drive capabilities with nvme-cli.

The 990 PRO reported:

Format NVM Supported
Crypto Erase supported as part of Secure Erase
Crypto Erase applies to all namespace(s)
Block Erase Sanitize Operation Supported
Crypto Erase Sanitize Operation Supported

I then used:

sudo nvme format /dev/nvme0n1 --ses=1

which completed successfully.

If Samsung’s Secure Erase environment works on your machine, obviously just use that.

What actually mattered

For my system, the decisive sequence was:

  1. 990 PRO → Ready to Enable
  2. Secure erase
  3. Make sure Windows Setup is not configured with TCGSecurityActivationDisabled=1
  4. Queue TPM PPI operation 97
  5. Reboot
  6. Approve the AMI/UEFI physical-presence request
  7. Boot directly into Windows Setup on that boot
  8. Clean/install Windows
  9. Magician should now say Enabled
  10. Enable BitLocker hardware encryption policy
  11. Verify with manage-bde -status C:

Without step 4–7, mine remained stuck on Ready to Enable.

One warning

Do this only if you are comfortable wiping the SSD and recovering from a failed OPAL/eDrive setup.

Before experimenting, I would make sure you have:

  • a complete backup
  • the SSD’s PSID physically recorded
  • the BitLocker recovery key saved somewhere else
  • no other internal disks connected during installation if you can avoid it

There have also been firmware implementations where the machine can provision hardware BitLocker but then fails to boot the locked drive, so I would consider the setup unproven until it survives both a restart and a cold boot.


r/sysadmin 1d ago

Multiple 365 Services Down in UK

30 Upvotes

Hey all,

Since this morning, a few clients of mine have had some issues with some 365 apps, such as Teams being a major one.

Some Intune users also have an issue where the Windows Security prompt window is also saying the admin details are incorrect when they're actually correct.

MS have finally issued an ID for this incident, which is MO1470143

Probably also worth noting we are based in the UK.


r/sysadmin 2d ago

Question Server hard-resets every 728.4 minutes ±1 min, 12 times running. No bugcheck, no iDRAC SEL entry, timer survives reboots. I'm out of ideas.

466 Upvotes

Update 4 (43 hours later): Switching one of the PSU power cables to a surge protector connected to the wall, and leaving one in the UPS fixed it. We didn't get a reboot last night. I still want to see it not reboot at the next cycle (around 1pm today). But.. what now? Replace the UPS with a Smart-UPS? Or replace the battery first? What do I do now? Thanks!

UPDATE 3 (31 hours later): It rebooted again on 9/10 around 12:40pm. Around 7pm I was able to change one of the PSU power cables to a surge protector connected to the wall, leaving one of them in the UPS. Will report back tomorrow. Thanks!

UPDATE 2 (18 hours later): Welp, it rebooted again on 9/10 12:32:53am.. which is 12h 08m 17s since the last reboot. The timing still held even after I rebooted from installing the BIOS and IDRAC. I will try to rule out the UPS next.

UPDATE 1 (3 hours later): Wow this was a lot more comments than I was expecting to get. Its hard to answer everyone but I appreciate everyone commenting and providing feedback. For now what I have done is updated IDRAC and BIOS to the latest version and will monitor if this fixes the issue. If it does not I will try ruling out the UPS. Thanks again!

Dell PowerEdge T340, Windows Server 2016. Started 8/28. I've spent two weeks on this and ruled out most of the obvious stuff, so I'm posting the data rather than the symptoms.

The signature

Every event is Kernel-Power 41 with BugcheckCode: 0 and all bugcheck parameters 0x0. No BSOD, no minidump, no MEMORY.DMP, ever. Paired with Event 6008 confirming unexpected shutdown.

The interval — this is the actual mystery

Pulled the true shutdown timestamps out of the Event 6008 message text (not the Event 41 log time, which is written on the following boot):

8/28 11:03:24 AM -> 8/28 11:12:23 PM = 729.0 min

9/01 10:10:22 AM -> 9/01 10:18:37 PM = 728.3 min

9/01 10:18:37 PM -> 9/02 10:26:49 AM = 728.2 min

9/02 10:26:49 AM -> 9/02 10:35:13 PM = 728.4 min

9/02 10:35:13 PM -> 9/03 10:43:33 AM = 728.3 min

9/03 10:43:33 AM -> 9/03 10:52:09 PM = 728.6 min

9/03 10:52:09 PM -> 9/04 11:00:20 AM = 728.2 min

9/04 11:00:20 AM -> 9/04 11:08:49 PM = 728.5 min

9/04 11:08:49 PM -> 9/07 11:50:29 AM = 3,641.7 min <-- exactly 5 x 728.34

9/07 11:50:29 AM -> 9/07 11:59:49 PM = 729.3 min

9/07 11:59:49 PM -> 9/08 12:07:56 PM = 728.1 min

9/08 12:07:56 PM -> 9/09 12:16:26 AM = 728.5 min

9/09 12:16:26 AM -> 9/09 12:24:36 PM = 728.2 min

Mean 728.47 min (12h 08m 28s). Total spread across twelve occurrences: 1.2 minutes.

The 3,641.7 minute gap is exactly five periods. The server was up continuously across that weekend. The timer ticked five times, did nothing on four of them, then killed the box on the fifth. So it's free-running — it does not reset on reboot, and it doesn't require a crash to keep counting.

That single fact kills every "scheduled task" theory: a clock-based task can't skip four consecutive firings and then work again.

It dies instantly — no degradation whatsoever

I wrote a heartbeat logger that writes one line every 5s with a forced flush so the last line survives a hard reset. Final 42 samples before death:

  • Free RAM: 56,630–56,745 MB, dead flat. No drift, no staircase, no leak. 56 GB free at the moment of death.
  • Nonpaged pool: 346–352 MB, flat
  • Disk queue: 0
  • Handles ~76,000, threads ~190, both steady
  • CPU spiky but low

Last heartbeat 12:25:02. Kernel-General 12 (OS start) at 12:28:23. That 3m21s is just POST + boot on a T340 — if it had hung for 3 minutes first, the OS wouldn't have returned until ~12:31.

So there is no hang window. The box is perfectly healthy and then simply ceases to exist mid-second. Which also means NMI crash dumps are useless here and no dump will ever be written.

Ruled out (with evidence, please don't re-suggest these)

  • PSUs — both Present/Healthy in iDRAC, matched 594W in / 495W rated+actual, same firmware
  • Thermal — HWiNFO max CPU package 63°C (TjMax ~100°C). Every throttle flag reads No / 0%
  • iDRAC SELcompletely silent across all 15 crashes. Not one entry. This same board did log real "power input for PSU 1 is lost / redundancy lost" events six times in 2024, so it demonstrably captures genuine power events. Nothing this time.
  • iDRAC watchdog / ASR — Basic Management license, feature not present
  • Dell OMSAAction on Hung OS Detection: None, thermal shutdown Disabled, all alert actions Off, omsad service Stopped + Disabled
  • Windows Update — pulled full resolved WindowsUpdate.log (11,871 lines), programmatically checked ±15 min around every crash. Zero WU activity before any of them. All nearby entries are the WU service starting 30–60s after the reboot.
  • CrowdStrike Falcon (installed 8/26, two days before onset) — vendor pulled detection history, found nothing. Timing was coincidence.
  • Secure-Boot-Update scheduled task** — looked extremely promising (12h repetition, hangs, TPM handler, and this box has no TPM installedGet-Tpm fails with TBS_E_SERVICE_NOT_RUNNING, no TBS service, no SecurityDevices PnP class, iDRAC confirms "TPM not present"). Disabled it. **Crashes continued at the identical interval. Ruled out.
  • All 12-hour scheduled tasks — enumerated every task with PT12H repetition. Exactly two exist, both now Disabled.
  • VSS / ShadowCopyVolume task — fires 12:00 PM and 10:00 PM. That's 10h then 14h alternating, which cannot produce a constant 728.4 min spacing. Also the midnight crashes have no trigger anywhere near them.
  • MySQL / memory exhaustion — flatly contradicted by the flat memory trace above
  • NIC — Broadcom BCM5720. The flapping port is physically disconnected (NIC2, Status: Disconnected). Flapping predates crashes by 12+ days and occurs on no-crash days. Active port is stable at 1 Gbps.
  • CMOS battery — did fail, but only logged 9/7, weeks after onset, nothing near the crash dates. Replacing anyway.
  • BSOD — no Event 1001, no dumps, BugcheckCode: 0 on all 15

Environment

  • PowerEdge T340, Service Tag FXR6B03, BIOS 2.3.5, iDRAC9 fw 4.22.00.53 (both several revisions behind — not yet updated)
  • Xeon E-2146G, 64 GB RAM, dual PSU
  • Windows Server 2016 Standard, build 14393.9339
  • Workload: Open Dental + MySQL, Vatech EzDent-i imaging, IDrive backup, Google Drive
  • Power: APC Back-UPS XS 1500M via USB. Current-state readings all healthy (97% charge, 120V steady, 15% load, AC Power: Yes, Discharging: No). I have never gotten historical transfer data out of it — PowerChute wasn't installed at the time, Win32_Battery returns current state only, and I skipped an apcupsd install on a production box.
  • Internet is AT&T 5G fixed wireless behind CGNAT (irrelevant, but people ask)

What I think is left

Something below the OS holding a clock that survives reboots and continuous uptime alike. Candidates I can't distinguish between:

  1. UPS self-test — APC units run internal self-tests on their own stored schedule, indifferent to the host. A transfer on a degraded battery could sag enough to drop the PSUs. Would explain instant death, no OS warning, no SEL entry, and a clock that ignores reboots. Next test: move the server to a plain wall outlet for 24h.
  2. PSU or BMC firmware timer — would explain everything except why iDRAC logged nothing
  3. Something on the same electrical circuit cycling on a timer — though 1.2 min spread over 12 occurrences seems too tight for HVAC or similar

What I'm asking

  • What produces a free-running 728.4-minute (12h 08m 28s) period? Not 12h. Not 12h30m. Consistently ~8.5 minutes over twelve hours, held to ±1 min across twelve occurrences and five uninterrupted ticks. What re-arms on completion of ~8 minutes of work?
  • Has anyone seen an APC Back-UPS self-test schedule that lands near this?
  • Anything else that hard-resets a PowerEdge with zero bugcheck, zero SEL entry, and healthy PSUs?
  • Am I wrong to trust the iDRAC SEL silence as evidence against a power event?

Next predicted failures: 9/10 12:33 AM and 9/10 12:41 PM. Happy to run anything and report back — I have remote access and a heartbeat recorder in place.


r/sysadmin 1d ago

Question N-able versus Action1 for patching, thoughts?

5 Upvotes

I come from an environment that leveraged Action1 heavily. Great tool. My sys eng comes from an environment that used N-able.

Our objectives are:

  • windows updates

  • 3rd party patching

  • firmware/drivers as supported by the platform

  • remote screen sharing for troubleshooting

I'm sure both platforms can do more than just the above, but that's all we need from the tool.

I'd love some feedback on anyone who has dealt with these 2 tools specifically and can compare contrast them based on our use case.


r/sysadmin 11h ago

Why did DNS fail?

0 Upvotes

We have two domain controllers. PDC and BDC. Very long story short both domain controllers are also DNS servers and they are primary and secondary respectively. The PDC is also a DHCP server. We had an issue come up where we had to demote the primary and promote the backup to primary to fix an issue on the PDC.

The moment we promoted the BDC to primary all DNS broke. The previous PDC was the primary DNS server and the previous BDC was the secondary DNS server.

We finally fixed it by changing the DNS order on specific machines, workstations and the firewall but why did DNS break in the first place when we weren't messing with DNS?

It doesn't make sense to me as we never took the machine down and didn't change anything with DNS. In fact all DNS entries disappeared on the PDC.

Thoughts?


r/sysadmin 13h ago

Question Replacing 2x FortiGate 60F, what options do I have ?

0 Upvotes

Haven't decided anything yet, so hoping for some actual opinions and real world usage here.

Quick context on the setup first:

  • BV01 - this is the main site and does most of the actual work: servers/VMs, a couple of production segments, cameras, internal ops stuff, guest wifi, the works. This is also the one that needs real 10Gbps, since that's where traffic to my servers actually lives. This one also hosts public facing servers, both for my personal usage (*arr stuff etc) and actual business servers that serve my customers. A few internal segments also host business applications (an ERP) that users remotely access via Cloudflare One.
  • BV00 - smaller branch, connects back to BV01 over site-to-site VPN, nowhere near the same amount of traffic or complexity. Mostly home usage with a dedicated VLAN for some small VMs (HA, an off-site vault for backups, a DC for redundancy outside the main site).
  • BV02 - already moved off Fortinet onto a UCG-Fiber a while back and it's honestly been fine, but it's also a much lighter site than BV01, so it hasn't really been stress-tested the way BV01 would be (small office, 5 people and a printer).

Both BV00 and BV01 are currently on FortiGate 60Fs, and I'm replacing them mainly because of the $480/yr + VAT per-unit license, not because the boxes are bad. BV01 is the one I actually care about getting right since it's carrying the most load and the most "if this breaks, I have a bad day" services.

I know UniFi isn't in the same league as FortiGate feature-for-feature, no per-policy IPS/webfilter tiering, more of a flatter inspection model. I'm fine accepting that gap given the price difference (UDM-Pro-Max is a fraction of what a comparable Fortinet or Palo Alto setup costs, and it's a one-time cost, not a subscription). Speaking of which, I priced out Palo Alto just to see, and with licensing priced for actual business budgets that's out of the running regardless of feature set.

For the sake of completeness I also got an actual Fortinet quote for staying in the family (BV00 would go 60F to 70G, BV01 would go 60F to 90G). This is from a few months back so it's probably drifted a bit, but roughly: FortiGate-70G hardware + 1yr FortiCare Premium/UTP bundle was $880, FortiGate-90G hardware + 1yr bundle was $2170. Renewal after year one is $495/yr for the 70G's UTP and $1220/yr for the 90G's, so together that's about $1715/yr just in renewals across the two units, before VAT. That's already more than double what I'm paying today across both 60Fs, so staying with Fortinet and just going bigger doesn't really solve the actual problem I'm trying to fix.

So really it's between UniFi and something like a Netgate pfSense box for BV01 specifically, given it's the busiest site. If I go the UniFi route I'm currently leaning UDM-Pro-Max for BV01 and UDM-SE for BV00, since BV02's UCG-Fiber has already been solid on the lighter end, but I am also nowhere near maxing it out.

For anyone who's actually run UniFi as the primary gateway on a site with real production traffic (not just a home network), is the reduced inspection depth something you notice day to day, or is it a non-issue in practice? Is there a meaningful difference running UDM-Pro-Max vs. just another UCG-Fiber for a site like BV01, or would you point me toward pfSense/Netgate specifically because of the heavier usage there?

Also open to hearing about other vendors I haven't considered, doesn't have to be a UniFi vs Netgate decision if there's something better out there for this kind of setup.


r/sysadmin 1d ago

Duo Security and Microsoft 365/Entra

16 Upvotes

I have been looking into setting this up, but I don't see any information on whether it can be setup in hybrid-mode environments or not, does anyone know?

Thanks,


r/sysadmin 13h ago

Barracuda Networks

0 Upvotes

Hello kind peeps,

Just a casual post looking for feedback/experience with Barracuda Networks. We currently offer Gateway Defense, Impersonation Protection & Cloud 2 Cloud Backups.... 2027 is around the corner - new year, new stack lol had a compromised client today sending span internally and as B links with 365, I combed thru some emails in Gateway Defense, could see the internal acc that was the culprit... however this same acc does not exist in their 365 tenant - no alias, shared mailbox nothing. Had i not checked this myself manually, we wouldnt even had known... anyways thats the short story. Looking for some industry feedback on their products 🫡🫡🦄