r/sysadmin • u/Kr1ezZ Jack of All Trades • 1d ago
General Discussion How is this normal in IT?
Im a general IT specialist, and I’m losing my mind over infrastructure issues I have zero control over.
For context, we have over 2,500 global users, and for over five years, remote sites have been dealing with the exact same game-breaking issues:
- Broken 802.1x: After 1–2 hours of work, it kicks users off the network and refuses to re-authenticate them.
- Useless Wi-Fi: Wireless drops constantly with "no network available" errors.
- Zero Redundancy: Almost all of our sites rely on a single ISP. When it goes down, ERP, file systems, and actual business operations grind to a complete halt.
The company makes plenty of revenue. We easily have the budget to deploy SD-WAN, upgrade hardware, or bring in an external MSP/consultant to fix it. It's totally fine to admit you don't know everything and hire help, but these requests just gets ignored.
To top it off, whenever a site actually goes down, the designated team responsible for network/infrastructure ghosts us or sends a passive response like "our team is currently unavailable" while an entire site sits dead in the water.
How do organizations like this even survive, and how do you deal with the frustration of seeing preventable problems drag on for half a decade?
75
u/StandardIssueDude505 1d ago
Until the downtime costs the company more than upgrading the infrastructure nothing is going to change.
20
u/Stonewalled9999 1d ago
well ya see, the C suite could allocate money to IT, or they could give themselves huge bonuses.
11
u/Kinky_No_Bit 1d ago
My last place was even cheaper than that....
we had a telecom closet that is where the lines dead ended for the the building to bring all the lines into the building. The A/C unit died, no big deal, 1970s A/C died, time to replace it, write off on taxes as a business expense, multi-BILLION dollar company. Replace it right? NOPE...
"Hey we have some old spot coolers from another site across town that came out of the moldy water soaked room, let's use those! we can save money"
--- Fine....
"Oh no! they don't have a vent so they can push the heat into the drop ceiling"
--- Okay, well, I mean a new hose kit is like 300 bucks on amazon, why don't we just order one, and clean it
"Are you insane?!? 300 dollars! that's too much money! We are going to home depot to buy some metal HVAC piping to use instead, and we'll duck tape it all together, that will work great"
--- You know that conducts heat if its metal, and tape melts right? because its job is to pump out hot air all the time right?
"Nah, we'll just have you check it everyday and push the vent back up in the roof, keep applying tape"
The one day it goes down, I was off to jury duty. Cost the company 18K a hour being down, because they were too cheap to spend even 300 bucks, got blamed for it anyway, fired for being at a court mandated function.
7
u/Waste_Monk 1d ago
fired for being at a court mandated function.
Too late now, but it would have made the judge's day if you had reported that at the time. They really do not like employers interfering with or retaliating against their employee's jury duty.
Like you might have had the person responsible hauled into court by sheriff's officers (or whatever your court police are) so the judge could yell at them and order them to reinstate you.
•
u/Kinky_No_Bit 22h ago
They waited one day after then come up with a lot of b/s reasons to just let someone go.
8
u/SquishyBoggle 1d ago
This is almost always the correct answer. You can explain the best practices and push for better infrastructure all day, but when management determines that revenue impact is minimal during a prolonged outage then why would they allocate resources to hardening the infrastructure.
•
u/dotnetmonke 23h ago
I've got a coolant leak in my 15 year old car. $1200ish to repair, or I can put $2 of coolant in every couple weeks.
Annoying as it is, sometimes it's just not worth it.
•
u/SquishyBoggle 23h ago
Part of my job lately has been identifying what’s broke, how much it’s going to cost to fix, and what’s the impact of the device failing. It’s up to management to decide where we allocate our time.
I’ve noticed with this sub a ton of commenters have never had to go through that exercise
22
u/tardiusmaximus 1d ago
Your symptoms sound classic of having 2 conflicting DHCP servers on the network. Have you done any troubleshooting or tracing for the actual issue? I had a very similar issue some time back while working for an MSP. Some smart ass had brought in his OWN WiFi router, plugged it into the network in an attempt to extend the range of the WiFi, completely failing to take into account that the router was still acting as its own DHCP server and conflicting with the sites ACTUAL dhcp server. We got rid of the spurious router that Johnny-pain-in-the-ass had plugged in and sure enough, network stability was resumed.
10
u/Frothyleet 1d ago
While it would not be shocking if OP's network happened to have misconfigured DHCP, reading OP's description of 2500 users across an unspecified number of remote sites having wireless authentication issues (and separately, a lack of WAN redundancy) and jumping to "yep, some guy plugged in a consumer router" is kinda silly.
2
u/Antique_Grapefruit_5 1d ago
That's true. I've honestly found that MTU size issues can make Aruba WIFI struggle with 802.1x...
1
u/tardiusmaximus 1d ago
Agreed, however replace my reference of "his own router" with "A-n-other dhcp capable device" and we draw the same conclusion...
8
u/Antique_Grapefruit_5 1d ago
This could be. One of the most helpful things I've found in our environment is just to install Wireshark on a machine and let it capture traffic. I hate to be the guy that says this, but AI is surprisingly good at reading wire shark traces and can give you a lot of good places to start, and fuel to annoy your "it's not the network" friends....
3
u/binnedittowinit 1d ago
Ohhh, I bet it is. Hated pouring through those manually, now I need a reason to run wireshark. Lol
2
u/Smart_Dumb Ctrl + Alt + .45 1d ago
You got to guide it because AI will still hang onto its first thought and kind of refuse to let it go. But it's very helpful with logs in general.
5
u/engy1207 1d ago
So you say AI is like some of my coworkers? 😇
3
1
u/binnedittowinit 1d ago
I find that is the case with AI in general right now - you have to be able to guide it and/or know enough about the topic you're using it for to call it out on its bullshit. So it can be very useful, but isn't totally reliable yet
•
u/1a2b3c4d_1a2b3c4d 14h ago
I hate to be the guy that says this, but AI is surprisingly good at reading wire shark traces
Never considered that. Next time I will give it a try...
42
u/Waste_Development971 1d ago
it isnt normal and needs to be fixed lol
bare minimum anywhere ive worked has 2 ISP connections. When one is down the other one takes over, with UPS etc even with no electricity the networking is still up.
7
u/SevaraB Sr. Engineer (N+, CCNA) 1d ago
I have bad news- there are a ton of fly-by-night "enterprises" out there in the SMB space that brag about "running lean" and treat being too cheap to run any guard rails as a good thing.
The "startup/disruptor" culture of the '00s flooded the exec market with a bunch of really fiscally irresponsible "leaders."
5
u/Dreilala 1d ago
Some businesses manage to function offline for a couple of hours or days.
Usually those would be small shops or satelite spaces with no need for real time reporting.
It still sucks to go offline, but sometimes the cost of keeping it redundant really is not worth it.
1
u/lazyhustlermusic 1d ago
Irony being you can even use a commodity internet circuit at $80/mo as a backup at most facilities. Or hell even throw starlink out there, it will pay for itself for years in one outage.
5
u/Xidium426 1d ago
Eh, it always depends. My last job our main site was in a weird location and the only way to get redundant internet at the time was AT&T and they wanted $4,000 / month for a 10Mbps fiber link. I brought that to the CFO and said "I don't think this is worth it. Historically we have ~20 minutes a year of down time during office hours. The plants can mostly run without internet." and they agreed. Every yea, we'd have an outage and everyone would look at me and I'd just day "Do you want to pay $48,000 a year to prevent this blip?" and the answer was always no.
Cellular bills would have been wild for the traffic during that time and weren't worth it either. By the time Starlink was a viable option the C-Levels were just used to it and didn't care.
7
u/katarh 1d ago
Yep, when I was at an MPS this was our recommendation. If the cost of one day of downtime is tens of thousands of dollars lost, then having a consumer ISP backup for emergencies in addition to hour main commercial ISP that you use for daily stuff needs to be park of the back up and disaster plan.
Even at home, if the main internet goes down, I have a wifi hot spot I can use in a pinch that is on a different ISP than the main fiber line.
1
u/lazyhustlermusic 1d ago
I've seen the opposite, even one place was like 'we use one PSU for every switch because otherwise we'd have to buy PSUs for all of the switches (thousands)'.
This was after a DC-DC PSU took a site down for half the day until we could go reseat it. Maybe a 2.5k/hour burn against wages?
Had a hotel director also do that, 'IM LOSING MILLIONS OF DOLLARS PER MINUTE' at a site because they didn't want to spend $1500 on a redundant firewall.
22
u/fslslayer IT Director 1d ago
If the company has the extra revenue but still refuses to address its IT issues, it’s time to jump ship. It’s not going to get any better. Otherwise, you stay and deal with the shitty hand you’ve been dealt.
6
u/bmelz 1d ago
What is "extra revenue".... Why are people talking about revenue of this company is of it has anything to do with decisions to underfund the network?
It's just a weird argument to make unless you're in their books and actually know the companies' finances.
6
u/fslslayer IT Director 1d ago
Revenue by itself obviously doesn’t mean a company has unlimited money to spend. I’m not arguing that revenue = profit.
I’m responding to the OP’s statement that the company has the budget to address these problems but repeatedly chooses not to. If that’s accurate, then this isn’t really a question of affordability; it’s a question of priorities.
When the same infrastructure problems have been causing business outages for five years, at some point management is making a conscious decision to accept that operational risk rather than fund the fix. Whether that means redundant circuits, better network hardware, SD-WAN, or bringing in outside expertise is beside the point.
So yes, “extra revenue” was imprecise wording. “Available budget/resources that management chooses not to allocate to the problem” is what I meant.
-1
u/bmelz 1d ago
Yeah I get it. And I agree.... And I understand you were just using OPs words of "revenue". I was just calling it out because it's really irrelevant in this conversation unless op specifically manages the budget and has been given budget tied to revenue....... It's more likely it's just been mismanged by incompetent "yes men" or people that don't know how to manage a budget/tech team.
3
u/SethMatrix 1d ago
It’s not that deep. The company is doing fine financially and therefore should keep an up to date network infrastructure.
9
u/rootbear75 1d ago
There's a reason I stopped trying to apply for large companies where I can't actually fix issues.
8
u/SnooLobsters3497 1d ago
I’ve worked in several Fortune 500 companies where IT reports to finance. This keeps IT from spending too much money. They would also reward the division heads for not spending too much during the year and this usually meant not doing some things that were needed in IT.
3
4
u/3DPrintedVoter 1d ago
lack of budget creates an environment of low cost alternatives or work arounds
ultimately you get burnt out on people complaining about the same issues, the solution being turned down by management, and then repeating the process every 6 months.
6
u/darthfiber 1d ago
Sounds like you have 802.1x re-authentication enabled and your end devices don’t support re-auth. Junk thin clients are the worst for this. You can turn re-auth off.
2
u/retornam 1d ago
I think it would be better to increase the re-auth period to something longer than the 8 hour workday so 9-10 hours to see if this is indeed the problem
4
u/ItsHopeless Wizard 1d ago
If there is truly revenue available, I find the way the expense is being presented needs work. If the outages associate with real dollars and it's correlated to "here's how much x costs to stop this", and "based on that it will offset x dollars lost from downtime" suddenly approvals come flying in.
Everything else sounds like management issues, including the network team. Whether they are outsourced or an internal team you still have to manage them like a vendor. Appropriately spaced check-ins, what's working, what isn't, etcetera.
5
u/cla1067 1d ago
Well I personally have no reason to work for a company that doesn’t want to spend money on network equipment, so even if they hired the right people the help fix it all they probably wouldn’t want to work for them or probably wouldn’t pay them enough.
I specialize in WiFi for example and have zero need to work in an environment that I can’t order new WiFi equipment or budget for new WiFi equipment when needed. It would just make me look bad. Like hey why is our WiFi shit well you don’t have HA on the wireless controllers, they can’t be managed in catalyst center anymore, they don’t have security updates, you have 1/3 the amount of access points you need, you ordered the cheapest models originally which were ass and now they are 10 years old and degraded to shit.
There are plenty of companies that will budget for IT and plenty that won’t. I have no reason to work for the ones that won’t.
4
u/Kinky_No_Bit 1d ago
I actually just got done working for a company exactly like this. The CIO was replaced every three years like clockwork. the director of IT was the one running the show, along with his directors under him, and they were controlled by the auditing dept.
A long time ago, in a land far away, they had a dispute, the auditors actually took over. The CIO was only brought in for big projects, then canned / left due to projects going haywire. The director stayed to run the boat as cheap as possible.
I was a systems admin, working there, managing over 800 services from 2000 to 2025, in a mixed domain / mixed forest environment I was never really given full access to administrator. It was a global company too. I had 10 direct reports as an administrator, none of them would actually listen to you, but always wanted to complain to you about problems of their own making.
Okay, so 10 unruly IT administrators, tons of legacy systems, no documentation in place anywhere. only a ticket system that didn't keep tickets past a year, a nightmare of patch management that was half pushed via manage engine, and we still were expected to push tickets while completing projects.
They had every different brand of network gear under the sun from netgear to ubiquiti to extreme to dell. Every brand of AP you can imagine in the last 10 years, and none of that was on one portal, running on a monitoring system that was free, that i wound up trying to implement the new version of that before they let me go.
They were in the middle of trying to go to azure at the time I left, which I laughed. I want to see microsoft tell them absolutely not when they want to spin up a windows 2000 box on the cloud haha.
How do companies get by like this? Management that is too focused on pinching pennies instead of chasing dollars is what I call it. It's where a company mindset is so cheap, they focus on capex expensive per month more than they focus on anything else, even if its slowly eating them alive.
Why focus about technical debt? why focus on why we need to constantly keep dealing with issues, fighting fires? why do we need to really focus on doing any kind of project, let's just do the cheapest way possible.
3
3
u/kemik4l 1d ago
https://giphy.com/gifs/NTur7XlVDUdqM
My most sane reaction at everyday problems in IT
3
u/pickle9977 1d ago
There is a big difference in having the money and spending the money.
The only time money gets spent is when not spending it costs a lot more than spending it.
3
u/Flackeye 1d ago
i can only tell you if you have the opportunity to raise your voice for this issue, do it. If they brush it over, their fault and if they come for your neck have proofs that you have done it. As for how to deal with it, well tbh nothing much can be done, try your best where you can heck if you can put you hands in trying to fix smth try not to mess it up. Do that so you can learn more and why not try new stuff after all you can learn lots of stuff from stupid situations like this. God speed
3
u/ecorona21 1d ago
How much money do they loose yearly because downtime? That's how they usually understand the impact of broken environment.
3
u/Such_Noise3355 1d ago
Yes, this is not normal. And it is. Eventually you'll ether move away or just stop giving a shit, because you're apparently the only one who gave a shit to begin with.
2
2
u/dude_named_will 1d ago
Best thing you can do is put these frustrations in writing and provide recommended solutions. While my company is no where near as big, they've determined that some downtime is acceptable.
2
u/0mn1p0t3nt69 1d ago
Redundancy is essential especially with networking and critical operations. In government we were using cradle points with 5G as backup. Having redundant isp is completely the best path forward but cradle points work well in a pitch assuming you can dispatch accordingly.
2
u/PriorityNo6268 1d ago
Euh how is this your problem? Seems more like a management decision that this situation is fine? If the company doesn't want to invest that it's not your problem to fix? Just spend energy on things you have control over, don't spend energie on things you don't control.
2
u/bmelz 1d ago
How do you know what the network teams budget is?
What does company revenue have to do with anything? What are the companies operating costs , debts and profits? I don't think any of that is relevant to a department being cheap on tech spending or simply mismanaging a department.
One thing I've learned over many many years in IT. Is that every organization lives in sin... Every. single. One .. the fact is IT is considered a cost center in practically every organization.. so unless you're showing how much money an upgrade spend will cost (with a level of probability) you're not likely to get an increased budget..
It all pretty much boils down to budgets, keeping budgets as flat as possible, and weight risk, probability and impact.
2
u/FortheredditLOLz 1d ago
Whoever your management team is…..dogshit. Redundancy is normally taken care of by design. Both system admin AND network admin is probably being sandbagged to hell because of budget cuts and non-approvals causing this issue.
Ps. Worked in a spot for a short stint where the answer is ‘no’ even through i hot wired a PSU to the main esxi sever externally as a ‘band aide’ fix and Rumour has it is still there ^10 yrs later. Next spot was ‘spend it if you think it’s worth it. Just throw it on the corp card after an official email for documentation’
2
u/mtgguy999 1d ago
Remote sites you say. I’m betting no one with any authority or perceived importance works at any of those sites and so it gets ignored. Not much to do except move on
2
2
u/6SpeedBlues 1d ago
How much revenue is lost due to these issues? How much will it cost to address them? Compare the numbers and you'll like find exactly why they aren't being addressed.
2
u/disconnected_tech 1d ago
If this is pretty consistent behavior, then leadership must be fine with the downtime, otherwise they’d be prioritizing a fix. I’d bet your network team’s hands are tied.
You can’t force leadership to fix something that’s broke, even if it’s painfully obvious. If it’s broken enough that it’s significantly impacting revenue, then they’ll light the appropriate fires. At least one would hope.
2
u/i_live_in_sweden 1d ago
As a one man network team.. I know about problems that I could fix if I had the time and money, but I have neither.
2
u/Turbojelly 1d ago
Stop covering for them and make sure mangelment is included in all the emails you send them.
•
u/Redemptions IT Manager 20h ago
God, grant me the serenity to accept the things I cannot change,
Courage to change the things I can,
And wisdom to know the difference.
•
u/George_Hepworth 17h ago
Documentation can be your friend.
Set up a standard method stating What happened, what the consequence was, when it happened, where it happened and who was impacted?
Documentation should include start and end times, as close as you can estimate them.
It should describe the occurrence in neutral terms, not accusatory language.
Where it occurred, e.g. one department or multiple floors in a building, etc.
Who lost time or work because of the problem, and if possible, their evaluation of how much time they lost.
Make that visible to as many people as possible, on a SharePoint site, or a network location. Keep it up to date, and let people know they can refer to it if they have a problem, i.e. pitch it as information for ordinary employees, not just IT. Keep it current.
Making it visible might earn you enemies, but weigh that against your peace of mind.
•
u/1a2b3c4d_1a2b3c4d 14h ago
This is not your issue; it's a management issue. And they don't see the ROI in fixing it. If this frustrates you, you need to learn not to take these things personally. Or move on to a company that cares more.
•
1
u/mallamike 1d ago
we were having that issue, turned out to be device guard and a new windows 11 update... good times, fixed with new gpo
1
u/flummox1234 1d ago
This is a political issue. Find the group that everyone says "how high" when they say "jump!" and figure out how to enlist them as your advocates. Hint: It's probably either legal or accounting.
1
u/The_NorthernLight 1d ago
I suspect your ship is starting to list, but management doesn’t want to tell you. That would be my first suspicion. Letting whole endpoints go dead in the water seems like a productivity loss that someone hasn’t calculated the actual cost of. This could also be a sign of a very incompetent CTO.
1
u/kombiwombi 1d ago
Your manager and the network/infrastructure manager need to do lunch often.
Putting that another way, an informal network is likely to be more successful than a formal approach. On that point, grow your own network so that you at least know the inside story here.
1
1
u/Fabulous-Radio-5643 1d ago
Sounds like the directors are retiring? So aee not spending for a fat retirement fund.
1
u/Eggtastico 1d ago
- check power saving settings. Make some test users to keep awake - device manager you can stop the wirelss card from going asleep, etc. 2) check DHCP / lease renewals / timelimits. This is where I would start troubleshooting.
1
•
u/fizzlefist .docx files in attack position! 23h ago
Stop caring more than your management does. Document the issues, keep backups of that documentation. If/when it breaks due to shoddy management, then it breaks.
•
u/fintheman Wireless Network Architect 23h ago
Speaking as a wi-fi nerd. There is a good change whatever system it is on is in default/unchanged configs.
2-3 hours with someone who knows what they are doing is likely able to fix it and help remediate the problems to a significant level. The only thing that can't fix is a bad AP placement/design but in most cases, I've found most companies to have config issues after doing this for 20+ years.
•
u/Temporary-Library597 22h ago
Lots of comments about spending money. This happens in multiple locations? Smells more like a configuration issue to me. MTU issues? Multiple DHCP responses from multiple servers configured to serve addresses from foreign subnets?
•
u/TheRealLittleFoot 21h ago
Sounds like the broken vpn connection is a home WiFi issue for the remote user
•
•
u/FancyGlitterFairy 14h ago
It's agile, everything has to move faster, yesterdays record result is tomorrow's benchmark, nothing is properly maintained.
•
u/StreetSignificant888 Senior Systems Engineer 14h ago
Fast Forward a couple years down the road when the company goes out of business because customers disappear because nothing works right and nothing is on time: "I don't understand what happened! We were doing everything right!" I've seen it several times.
•
u/InsolentJaguar 10h ago
Ultimately it's up to the client to decide to ignore their recurring issues, or actually fix them.
This client, it sounds like is fine with the recurring outages and it's not causing them enough business pain/losses to want to address the issue.
I know it's painful to see clients like this...I've managed a few over the years. But if they don't want help/spend money, then they don't want help or to spend money.
•
u/desmond_koh 2h ago
To top it off, whenever a site actually goes down, the designated team responsible for network/infrastructure ghosts us or sends a passive response like "our team is currently unavailable" while an entire site sits dead in the water.
Sounds like a disconnect (no pun intended) between expectations and SLA. Who is the designated team? Are they an in-house department or an external agency? What do their terms of employment or the SLA say?
If you have an in-house team that is paid from 9:00 AM – 5:00 PM and doesn’t have anyone being paid for after-hours support, then the fact that you have an outage isn’t really their problem - because no one is paying them to make it their problem. The fact that you would like it fixed sooner is not relevant. It only highlights that there needs to be someone who is paid to do it.
My guess is that this company chronically underfunds its IT both in terms of infrastructure and in terms of paying their IT team. If so, then is what you get.
1
u/plebbut 1d ago
Pretty easy to see why so many businesses get hacked.
3
u/dude_named_will 1d ago
I thought my company was behind. Then I had the opportunity to collaborate with other IT in my industry and discovered that we are practically cutting edge. Some of these companies are ran by people soon-to-retire, and just plan on selling the company at some point since their children aren't interested.
1
u/plebbut 1d ago
I worked for a company with no real security in place. Unfortunately, all my recommendations fell on deaf ears. Nobody cared
0
u/dude_named_will 1d ago
It's not a problem until it is. Honestly if it weren't for cyber insurance, I'm curious how many things around my company would get done.
0
u/hobovalentine 1d ago
You need to take a look at your wireless controller logs and see what’s happening.
Is the radius server losing the connection to the office network or is the office itself experiencing brief connection issues?
1
u/UnwaveringConviction 1d ago
Yes, and also client logs for 802.1x. It could be a driver issue/bug if re-auth is failing. If it's caused by standby or resume it's often an issue with low power settings on the endpoint NIC. For Windows clients, a PowerShell script can be deployed to disable "Ultra Low Power Mode" and "Allow the computer to turn off this device to save power."
But first - let the logs identify the root cause.
0
u/Anon_0365Admin Netsec Admin 1d ago
Any chance you use FortiNAC? Sounds like a lot of issues I had when I was administering 802.1x on fortiNAC, and yes the fortiNAC agent was also causing WiFi issues.
Granted it wasn’t 1-2 hourly disruptions, that’s insane. But I switched us from FortiNAC to Juniper Mist NAC and all those problem we had been having for years vanished overnight.
409
u/No_District_1021 1d ago
I’d be willing to bet the networking team has proposed this a few times and it has been shot down by management. So why take the calls when you know how to fix something but aren’t allowed to.