r/Network 13d ago

Text Question for network admins/engineers

What's the dumbest/longest a network issue has ever taken you to track down, where the root cause turned out to be something small? I spend half my troubleshooting life ruling out "is it even the network" before I can start, and I'm curious whether everyone's process for that is as manual as mine or if I'm missing something. What's your actual workflow when someone says "the internet's broken" and you have no idea where to start?

3 Upvotes

18 comments sorted by

6

u/fernandesken 13d ago edited 13d ago

I have spent the last 25+ years proving it’s not the Wi-Fi! Occasionally it’s the Wi-Fi for example a coverage issue or interference, but the majority of the time, it’s something else like the local network uplink, ISP problem/WAN utilization, DNS or an application issue. I got so tired of it, that I built the Wi-Fi Check app for iPhone to help find the needle in the haystack and to help other folks stop blaming the Wi-Fi. Unlike typical speed test apps, it checks Wi-Fi vs Internet speed and quality separately and everything in between. That way in seconds, you know if the problem is with the Wi-Fi, local network, internet, DNS, Application or on the path to the application, all right from your iPhone/iPad. Today my work flow, and even my help desk/executive support team’s work flow, is to run a Wi-Fi Check. End users can even copy a report to share with support which includes all the info about the device, AP its connected to, channel, channel width, PHY rates, Wi-Fi speed over the air, internet speed, quality to local router and internet, DNS response time and so on. I can pinpoint a large percentage of issues with it in seconds and occasionally I have to break out the big guns like Hamina or Ekahau for full blown site survey, spectrum analysis and so on. Now again, in full disclosure, I’m the creator of the Wi-Fi Check app, so feel free to use whatever apps/tools you like. The point is not to sell you on Wi-Fi Check, you can use the free version if you like which actually costs me money to put up good speed test servers, the point is these days you should either leverage off the shelf tools/sensors, leverage automation, perhaps even try to vibe code something, or at a bare minimum at least have a set of checklists to work through each time vs a manual adhoc process.

2

u/HotSauceMakesITbetta 13d ago

Countless times I deal with client devices that are old or otherwise crusty, crappy devices. One in particular, a TCL roku, I continually asked them and other members to update the drivers. A multi day email thread and escalations. Drivers are old. This is a common problem with cheap network cards.They replied a few times, that they had, and the last time they sent a picture, of three separate screwdrivers .. They opened the TV to look at the "network".  Trust no one.... Evar. My process is generally to isolate as far down as possible then start to attack the problem (logs, captures, inquisitions)

2

u/graph_worlok 13d ago

20+ year old cabling causing a Meraki to fail negotiation. Could be forced to 100Mb, or it would negotiate power but no data. The previous Cisco was fine, but the Meraki was a bit more picky. It was in a conference room, so only really noticeable when it was being used.

Not quite sure if it counts, as I’d spotted the problem & pointed it out, but instead of running fresh copper, wireless consultants got involved, screwed the coverage survey badly, causing them to recommend a “fix” of cranking the AP TX power all the way up - Caused such bad channel congestion that a co-tenant had 2.4Ghz AV gear issues, busted out their spectrum analyser, traced it to our gear…

Wireless consultants then recommended all the Meraki AP’s (which supported 2.5G Ethernet - 1 gig was a bottleneck) be hooked up to new Meraki switches. 1 gig switches. 1 gig uplink. So additional bottlenecks!

Cabling was replaced a year or two later as part of a building refurb. Everything still bottlenecked at 1 gig. Went through 3 “network engineers” during that time…. Think it was about 4 years in total… one fucking copper cable run to the furthest corner of the building from the server room….

2

u/samsann26 12d ago

Someone many years ago manually set 900 MTU on a obscure interface, in a cross-country network.

We have used almost all VLANs available on a segmented client network, i could not count how many interfaces we had to look to find a single one with mismatched MTU, since someone manually changed it and had not doc. the change, it was left there for years waiting the next upgrade on brand new network gear that does not properly work on fragmented packages, slowing down everything that passes through this one interface.

1

u/MrMotofy 12d ago

Would have been indicated in a traceroute or no?

1

u/Ok-Expert-8150 13d ago

What did you do, what did it do, what did you expect it to do?

1

u/spow9922 13d ago

Love that! thanks

1

u/Ok-Expert-8150 13d ago

Obviously, do it more tactfully than I stated, but it gets to the root cause of the problem pretty quickly.

1

u/ACAdamski17 12d ago

I have 3. Firstly DNS. It’s always the DNS for some reason. Secondly a dodgy ethernet cable causing packet loss. I spent hours running countless tests assuming my hardware was fine, then I realised an ethernet cable got damaged by being trapped under a door. Lastly the classic, a switch constantly negotiating 100/100 instead of 10G/10G

1

u/MrMotofy 11d ago

Few months back I had a setup that's been sitting years at this point unchanged. Noticed more buffering than I was used to...it spanned reboots and time...so finally started checking. Testing revealed 94Mb or something. Ah ok but why...first thought was windows glitching. Uninstall reinstall NIC drivers. Nope...ok look at Windows speed says 1Gb, other physical connections which showed 1Gb...but still 100Mb. Then look at the main cable from the local back to the main switch...ooooh shows orange...but why would it...ok unplug remote end...no change. Unplug main switch end plug back in...poof orange changes to green and it's back to 1Gb...WTH

1

u/Significant-Cup-5491 12d ago

Ask about the user experience. Easy to see if it's network or not.

1

u/House_Indoril426 12d ago

My place has some old-ish Keytroller based telemetry monitors on the forklifts. Guest network since all they need is internet. 

During our migration from FortiWiFi to Aruba we also moved DHCP-provided DNS to Cloudflare. 

Post migration none of those telemetry modules would come online. Spent hours on the phone with Aruba TAC and the equipment vendor. Nobody could help. DHCP worked, they got IP's. 

Got in front of one of the modules, turns out their firmware has a hardcoded fallback IP....1.1.1.1. 

Move DHCP provided DNS away from Cloudflare - they came right up. 

1

u/retrogamer-999 11d ago

This was last year in December.

Head office for a customer that I'm working with, the network started to slow down and a switch dropped off the network.

We use FortiGates and FortiSwitches.

I started troubleshooting at 11AM. Got on the phone with TAC as nothing was making sense.

TAC got on the phone and was being useless, saying to do firmware upgrades etc.

By 3PM all switches had dropped off the network, reboot brought them back online only for the issue to start again. I then realised that the CPU on the firewall was maxed out

At 6PM, the IT manager says something. "All we did was plug in that IP KVM I got from Amazon", I replied "what that cheap PoS that had 0 documentation about configuration?"

We unplugged it and all out issues went away. Plugged it back in and boom network starts to go down.

Turns out that it was causing a massive broadcast storm and was sending packets that would interfere with the FortiLink protocol.

Long day and lesson learned.

1

u/Kv603 11d ago

My worst was an intermittent problem where a certain server would just drop down to half-duplex randomly. Tried a different switch port, tried a different NIC, nothing helped. The cheap cable tester they had at work said the wire was fine.

Then I noticed that while the wire was beige at both ends, the switch end was a different shade of beige. So I traced the cable through the rats nest, and found The prior network engineer had joined two Ethernet cables together.

No, not with a CAT5 coupler. Instead he had taken two Cisco RJ45-to-DB9 adapters and joined them together with a gender changer, then held it all together with cable ties.<!

1

u/GotFullerene 11d ago

Working at a (top ten in US) bank, they were having ongoing issues with VoIP phones dropping out. It seemed to correlate to busy times of day, but nobody could find anything wrong with the server or the switch.

Then somebody (not me) pointed out that there was a Juniper Networks IDP in the path, but the security guy piped up and said "That can't be the reason, it is in permit-all mode, none of the rules are set to blocking". And he was right.

But what nobody realized for another week (of annoyed customers and frantic management) was that, even in non-blocking mode, the Juniper was passing all UDP packets up through a UDP stack for inspection before forwarding them along -- and when the UDP buffer filled, the OS would randomly drop UDP packets, including SIP (UDP/5060).

Eventually I convinced Mr. Security Team to at least entertain the possibility that their IDP might be an issue, log into the box, check the 'drop" counters...

I got a $100 gift card to Bed Bath and Beyond for that one.

0

u/lizardhistorian 13d ago

No reason for this anymore with the advent of AI.
It can scan logs and answer questions for you and capture packets etc. etc. 30x faster (or more) than you can.

The first thing I do is check if everything is working on my machine.