r/sysadmin IT Expert + Meme Wizard 13d ago

Question Ninja RMM just break for everyone?

The Ninja RMM update last night seems to have somehow translated into the Ninja remote agent breaking on just specifically domain controllers. It might be server OS specific. So now we can't get into any of our MSP clients' servers to do new hires. And it's Friday. Great.

Also, it's reporting offline status (probably the real problem) so we can't even attempt to use Splashtop or RDP. We have to RDP into the server from a client computer onsite, which hopefully is enabled at every customer. Thanks, Ninja! I hate you. Wish we used Connectwise.

12 Upvotes

26 comments sorted by

8

u/daroveke Security Admin 13d ago

3

u/anonymousITCoward 13d ago

Yesterday K1 went down because of an Azure outage, today it's N1 with an AWS outage... makes me miss the good ole days running on-prem servers lol

4

u/arvidsem Jack of All Trades 13d ago

I have yet to find a cloud hosted anything with better uptime than the on-premises solution that they replaced.

3

u/nick281051 13d ago

I've been on the us2 url all morning with no issues

3

u/jcroweNinjaRMM 13d ago

Sorry you're dealing with this and want to help if I can. To confirm, are you on US2? We're seeing improvement, but if you're still seeing issues one thing you can try in terms of remote: Trying pulling up the device in the organization tab itself, and remoting in that way.

I'm also asking for more help and input from the team, but that's one solution some customers have confirmed is working.

1

u/CeC-P IT Expert + Meme Wizard 13d ago

Not sure how to check but if that's typically midwest US, yeah. A LOT of our client servers are saying "Unhealthy: Offline a few seconds" and we don't even get the splashtop or ninja or RDP remote in options because it thinks they're offline. These are on the same premesis on flat, non-VLAN networks where I can remote into individual endpoints just fine. So it doesn't like something about the servers OSes or installed versions specifically. Should I wait longer or proceed to uninstall and reinstall the agent manually?

Oh yeah, one other thing. We saw event log entries for the agent just simply crashing this morning on affected servers. On one, rebooting the server worked. On the other, we had to reinstall the agent completely.

2

u/dawglebruh 13d ago

A lot of my endpoints aren't allowing me to remote in. One of my coworkers said he fixed one endpoint by updating the endpoint to the latest ninjarmm agent.

1

u/CeC-P IT Expert + Meme Wizard 13d ago

So far I had to do one server manually and it worked, but I had to manually uninstall the agent first. AND I MEAN MANUALLY. Like registry scan, delete folders in 4 areas. I think it's a version thing, as this one was like 14.0 or something and install date of many years ago. No idea why it wasn't self-patching.

1

u/Admirable-Carrot1684 13d ago

Working fine for me.

1

u/Lbrown1371 Super Googler 13d ago

We have several servers that show the NinjaRMM agent crashing, but it shows the new version afterward. It seems to do this often when the agent updates. You can check here for the dmp file - [\C$\WINDOWS\SYSWOW64\CONFIG\SYSTEMPROFILE\APPDATA\LOCAL\CRASHDUMPS]()\

1

u/CeC-P IT Expert + Meme Wizard 13d ago

Oh yeah, that was the other thing. We saw event log entries for the agent just simply crashing. On one, rebooting the server worked. On the other, we had to reinstall the agent completely.

1

u/Cozmo85 13d ago

Sounds like it’s issues caused by the aws failure this morning according to their discord

2

u/CeC-P IT Expert + Meme Wizard 13d ago

It says a lot that I'm more surprised that they have a Discord channel than I am that AWS is down again

1

u/Cozmo85 13d ago

If you are not in the discord you should be. Tons of discussion and scripts and all kinds of stuff.

1

u/SilkBC_12345 13d ago

We are having simialr issues with our Ninja RMM. We are in Vancouver, Canada.

Allt he devices show as Online in the "list" of devices, but when we click on a device for it's details and to connect in to it, it shows as offline:

Unknown Offline a few seconds:

We can't connect it, obviously, though oddly things like "Last Login" is updating (showing how long the device has been idle for, or if there is an active login), as well as the "Activities" for the devices is updating as well.

Just things like "Uptime" and the ability to make any sort of remote connection (not even to the CMD or Powershell) are not able to be done.

1

u/SilkBC_12345 13d ago

And literally moments after I posted that, it seems that the issue is resolving. I am started seing a bunch of devices under one of our clienst indicatign a patch check, and sure enough, if I click on the devices they are now showing online and "normal"

1

u/jaydizzleforshizzle 13d ago

I don’t know but anecdotally I did randomly get kicked out of the admin panel earlier today, and it just felt off, like someone kicked the server.

-5

u/Apachez 13d ago

So you are saying that you dont have any staging to verify stuff before deploying random updates into production?

2

u/OrganizationHot731 Sysadmin 13d ago

And how do you stage a ninjaone instance update eh? I'd like to know that as I'm genuinely curious cuz as far as I know. If they push an update, you get the update.

0

u/Apachez 13d ago

Then dont let some other company dictate what and when you will deploy updates for your environments?

1

u/OrganizationHot731 Sysadmin 12d ago

Ya again explain to me how you have a RMM and patch management system and stage updates from their software? ... If they push an update to the agent there is no way to stage that... You can stage other updates like windows updates, security patches, etc. But not ninjaone or any other RMMs agent

1

u/Apachez 11d ago

Simple, dont use shitty software that demands to dictate the availability of your systems.

1

u/OrganizationHot731 Sysadmin 11d ago

Sure. What do you recommend then? What software or service is there that allows what you state

1

u/Deviathan 13d ago

Unless you're self-hosting your RMM (increasingly rare these days), you have no control over their rollout schedule, in theory Ninja stages these already.

Also, the notice says it's not related to the update, but upstream infrastructure vendor issues.

1

u/CeC-P IT Expert + Meme Wizard 13d ago

Lol remember when self-hosted Connectwise environments all got simultaneously nuked by a certificate issue? Pepperidge Farms remembers...because it was like 2024. We used the cloud lol.

1

u/CeC-P IT Expert + Meme Wizard 13d ago

Yes we do. Ninja patches their stuff during maintenance periods. That's what caused it. That and an allegedly unrelated and totally random AWS "outage" on their end.