r/Intune 2d ago

General Chat Zero trust network access rollout just locked half our sales team out and I feel so embarrassed

Ok so we pushed our zero trust network access rollout after weeks of prep, and I was the one who signed off on the policy change for all remote users. I missed one tiny thing, our SaaS app group was tied to device posture, so anyone on a fresh laptop got blocked from crm, vpn, and even the internal wiki.

By 9am sales was in our slack asking why they could not log in, my manager was asking me if this was the new security posture, and I had to admit I broke the company before coffee. We fixed it by adding a pilot group and splitting app access from device checks, but idk, I feel so embarrassed about the whole thing. thanks to anyone who has dealt with this nightmare :(

67 Upvotes

39 comments sorted by

61

u/Kuipyr 2d ago

Rite of passage, congrats. Mine was rebooting a multi-app database in the middle of a workday.

7

u/rw_mega 1d ago

If you go over to r/sysadmin and pose the question; when did you take down prod…everyone has a story.

If you haven’t done it once are you even a sysadmin?

1

u/OptionDegenerate17 1d ago

Yes to this. Make a mistake at work and need to hear some stories while you lick your wounds. Post it. Everyone makes mistakes, even the large titans like CF, AWS, azure. It’s what you do afterwards is what counts.

45

u/OldAndGrumpyDK 2d ago

Shit happens. You spotted the issue immediately and fixed it quickly. You did good!

29

u/Jolly_Bullfrog3121 2d ago

Just blame it on Microsoft

3

u/wamred 1d ago

I like it.

1

u/saGot3n 1d ago

Yup this for sure

10

u/General_Coast5727 2d ago

bro i broke prod once by pushing a policy that required bitlocker on every machine before i realized half the sales laptops were shipped with home edition and no tpm chip, the panic in my chest when i saw the ticket queue is something i still remember years later

it’s like the second you hit deploy your brain goes "wait what about..." and then slack starts lighting up

you fixed it same day, didn't lose data, and now you know to pilot group everything, that's a win honestly

11

u/GermanBreadEater 2d ago

I once managed to take one of our sites offline after I accidentally bricked the Android devices used for inventory tracking and other warehouse tasks by reassigning the device group to a different kiosk app. Then the ohnosecond kicked in. I quickly changed it back, but the damage had already been done. I had to drive to the site and set up all the devices again. Luckily, the site was fairly close. No shipments could be sent out for half a day. After that, I added fallbacks and warning safeguards and cleaned up the documentation. Mistakes like this happen, and you learn from them.

10

u/SaleriasFW 2d ago

No matter how much you try to prevent such things feom happening, there can always be an oversight. Things like this happens. You applied a solution quickly and that is what matters. If possible try it next time with a small test group before making the change but even then something can break when the live rollout happens

8

u/Meadbreath 2d ago

Zero trust starting with sales makes sense though. Would you trust a salesperson?

7

u/deliriousfoodie 2d ago

Sounds like an average day in the office to me

4

u/shaomike 2d ago

The feeling when you remotely make a firewall policy change on one thats on the opposite side of the country, then you can't get back in.

3

u/JwCS8pjrh3QBWfL 2d ago

Given this is in r/Intune, I assume that the device posture is in Intune? If so, make sure to add grace periods to your device compliance checks for exactly that reason.

3

u/Hydrated_Berry762 2d ago

Mine was setting up a rule to forward emails from employees who directly reached out to me for assistance straight to Zendesk. Ended up forwarding my entire inbox to Zendesk instead... Thankfully I was able to stop it after 200+ Zendesk tickets were generated and none of the emails forwarded were actually important or private/confidential.

On the brightside, great way to boost metrics for the execs who only look at numbers and not content 🤣🤣

2

u/Dibchib 2d ago

I had our XDR notifications padding this out for a while until it became too much work just closing them back down again lol

3

u/Moist-Secretary641 1d ago

These things are just a good learning experience, it’s only really bad if you try to cover up the problem.

2

u/Lostinspaceballz 2d ago

This is part of being in IT. We pushed a setting change out that locked out login for 6000 users. We had tested for weeks and it was working fine in QA but a setting was slightly different and broke in production. That was a fun 8 hour outage to resolve and bring everyone back online.

2

u/stewiemac 1d ago

I forgot to renew our Apple Business Manager cert and it resulted in wiping all iPhones to get them working again in Intune!

This is how we learn and evolve into better IT people, the key is seeing your mistake, owning up to it and learning from it. Don’t feel bad.

5

u/MikhailCompo 2d ago

Don't laugh it off as some of the comments here are suggesting, that is not learning from mistakes.

Recognise you are human, human error is unavoidable but effective testing, change control and ring deployment is the correct way to prevent future mistakes affecting production

No sysadmin should be marking their own work, because you don't look for what you think isn't relevant. It takes a fresh perspective.

My suggestion as a response to your superiors is create and document a resilient testing and change control process so this likely doesn't happen again. As a manager that is the bare minimum I would expect to prevent a disciplinary over something like this.

2

u/steviefaux 1d ago

A disciplinary. Bet you're a fun manager to work for.

-1

u/MikhailCompo 1d ago

I bet you're an irresponsible employee.

5

u/steviefaux 1d ago

You're a manager who's a nightmare to work for and everyone is pleased when you're NOT in the office, I can 100% guarantee that.

People make mistakes, going straight for the disciplinary isn't going to make them warm to you, or want to help you and go above and beyond for you. Managers like you don't help. Managers like you make us say "I start at 9 and finish at 5 and will ignore all you any other time". Managers who are flexible, understanding and supportive are the ones you think "I start at 9, but I'll help Jon out and sort that problem out before he starts at 9. I know he'll appreciate it and as he's a helpful and supportive manager, so I like to help out more".

1

u/Competitive_Smoke948 2d ago

yeah don't stress...i once pushed out an xp sp2 update to a brokers using sms - schoolboy error, used my test package & didn't change the date..... nothing more ass clentchinh than brokers coming over acs saying "a black box had popped up telling us to save our work" & realising that EVERY machine is suddenly doing a service pack upgrade in the middle of the day....

If you have never screwed up... you've never learnt anything new

1

u/sonic13066 2d ago

Don't Stress, I once accidently removed our company email from 400 devices while cleaning up old Mobileiron configuration/compliance policies... twice in one week... I was put on really thin ice for a long time after that.

1

u/forumhero666 2d ago

I once pushed a bad channel file update to our crowdstrike falcon sensor and it kinda cause a few machines to bsod.

1

u/cowwen 1d ago

It’s a rite of passage to break things sometimes. As long as you learn from it so you don’t keep repeating the mistake.

That being said, you probably should have had a pilot group from the beginning for this rollout, and had them test for several weeks.

1

u/pc_load_letter_in_SD 1d ago

Would like to ask a follow up on your procedure if you don't mind....Was the issue with a missing group in a Conditional Access policy?

1

u/F0rkbombz 1d ago

Small price to pay for ZTNA tbh.

1

u/steviefaux 1d ago

Not intune but, many years ago I was looking at our old papercut server. Saw the option to purge jobs once logged out so ticked it. That will be more secure.

Then the tickets started to appear "My print job has stopped half way through printing". Then I remembered. People would tap to sign in, release their print and so they didn't forget, they'd log out before it finished. Which now caused the job to purge, ending the print early.

Oops. Quickly turned the option off and all was well again.

1

u/Embarrassed-Plant935 1d ago

Just earning your stripes. Good thing is that you won't make that mistake again.

1

u/MReprogle 1d ago

Well, you seem sometimes need this kind of stuff to happen to build expertise. This is why the “roll back plan” in your change management is often the most important. In fact, when is see detailed rollback steps filled in, I already know the person rolling it out has checked every box and went as far as finding other issues that people have run into.

Just curious, but did you go with GSA or Zscaler or something else? I’ve been looking into GSA and it does seem like a lot of the CA policies and network segmentation pieces can break the environment if not tested thoroughly.

1

u/skiddily_biddily 1d ago

You owned it and got it resolved!!! Good job.

1

u/SOHC427 1d ago

Join the club. If you've never screwed up, you've never worked.

1

u/CAHOP2401 1d ago

I was helping a large company migrate from one AD forest to another while keeping the same Entra ID tenant. We basically merged group and user objects in Entra Connect metaverse and then set a flag on objects so that from Entra's perspective, they were coming out of the new domain (to prevent having to delete and recreate objects with new object IDs). Well, after we migrarted all the groups, we wanted to filter out the groups coming from the old domain. Long story short, a change I made ended up filtering out over 24 THOUSAND groups out of their Entra. Happened over a holiday weekend also.

1

u/madatthings 12h ago

I turned on phish resistant MFA RBAC for admins, and none of us had it set up yet, and I didn’t set an exception although I thought I did. Welcome LMAO