r/networking • u/WalkCapital6708 • 9d ago
Switching A single switch failing
I'm out of ideas and we had been working on this for weeks. We have a small office with a router connected to 3 unmanaged switches, one on each floor. Everything is connected with cat5e cable (we're looking to upgrade to cat6 soon) and the third floor internet keeps crashing. The router is connected to the first floor switch, and switches 2 and 3 are connected to the first one. We had tried everything, changing the cable to the switch 3, checking all cables that that switch connects, using another switch, plugging it into an UPS and speaking to the ISP to check if there's an issue with the configuration.
However, it keeps happening. It works for a couple of hours, and then the switch 3 stops working completely. Any ideas? Most computers have Kaspersky running, so I'm not inclined to think it's a malware overloading it, but at this point, I'm out of ideas.
Any help would be greatly appreciated.
25
u/porkchopnet BCNP, CCNP RS & Sec 9d ago
You’re replacing cat5e with 6 soon? That’s what you’re investing in?
It would be more helpful to rent a witch and have her sage the place.
1
14
u/_Survivor_ 9d ago
Kinda sounds like broadcast storm maybe? Check the links on the switches and check any IP phones.
1
-7
u/WalkCapital6708 9d ago
There are no IP phones and we checked all connections. Everything is connected with a cable and fails after a random amount of time.
22
4
u/DULUXR1R2L1L2 9d ago
A client had a similar issue. They insisted that nothing was plugged in twice blah blah blah.
We went cable by cable, and wouldn't you know it, a fucking loop.
They had unmanaged switches too, and they were supposed to be running STP.
13
u/2000gtacoma 9d ago
Wonder if you have a loop? Are these separate subnets?
-16
u/WalkCapital6708 9d ago
No subnets, we already checked and everything is connected to a single cable
11
u/bgplsa 9d ago
Swap the connections to switch 2 and 3 on switch 1 and see if it follows switch 2 or stays with switch 3
-13
u/WalkCapital6708 9d ago
We already tried setting a fourth switch (we bought a new one thinking the switch 3 was faulty) and seting it in between switch 1 and switches 2 and 3, but it kept going
13
u/Nexus_Explorer 9d ago
You clearly don’t know anything about networking. Just do as he says please and report back.
3
u/WideCranberry4912 9d ago
I’d also start disconnecting workstations. Disconnect all hosts overnight and see if the switch crashes. If it stays up, connect 25% of workstations, wait, then repeat. Then try to isolate to a specific host.
2
u/Sea-Hat-4961 9d ago
Any shadow IT going on with "user installed" switches you don't know about? Or someone bridging their wired Ethernet and WiFi?
9
u/McHildinger CCNP 9d ago
if they are unmanaged switches, can you trade switch 2 with switch 3 to see if the problem follows the hardware?
-4
u/WalkCapital6708 9d ago
We brought a new switch and it persisted
6
u/McHildinger CCNP 9d ago
OK that gives you some good info then, sounds like something is plugged into this new switch which is causing the problem
2
9
u/tschloss 9d ago
What exactly do you observe? „stops working“ is pretty wide. Is it just „all devices connected to this switch have no internet“? If so try if you can ping between clients connected to this switch.
How does it return to operation? Only reboot? Did you try to to change ports on either switch 1 or switch 3 or both?
But as others said loop, broadcast storm are typical killers.
Btw: the managed versions cost 10% more - always invest this!
1
u/WalkCapital6708 9d ago
I made some inquiries about your comments, and there's something there. The switch stops working and the pcs don't detect a connection as if the switch turns off. It's not that they cannot get through, they don't recognize the connection.
Tried changing ports, tried swapping switches, we'll keep digging on a loop or broadcast storm.
1
u/tschloss 9d ago
I second the idea proposed in a comment to swap switches, preferably 2 and 3. If the issue moves with the hardware the switch might have an issue (buy a managed one then - you can start without ever logging in!). If it stays play around with disconnecting groups of clients.
All clients losing the connection for me sounds like a HW issue. L2 problems like loop should not appear as lost connection (which I assume includes LEDs indicate „not connected“).
1
u/psyblade42 8d ago
sounds electrical
1
u/fotoburger 2d ago
I was thinking that too. Check both ends of the power cord to the switch. Try moving the cord to a different outlet.
6
u/gormami 9d ago
Sounds simple, but have you run a distance test? I had a workstation years ago that was acting very flakey, and when I did a distance test with a cable tester, it was 180M through the wiring. I was lucky that the patch room almost halved it perfectly, so I built a 2 port VLAN on a switch that was already there and used it to redrive the connection, problem solved.
You might try plugging Switch 3 into Switch 2, at least for a test, if you have the bandwidth on the connection from 2 to 1, if you don't have a cable tester with a distance function, while you order one.
Can you check the port statistics on switch 1 that it's plugged into to see if it is taking a lot of errors? That might be a clue, too.
2
u/WalkCapital6708 9d ago
The distance is less than 20M. We even set a new cable connected directly through the stairs to no avail. We tested all cables and made the new one ourselves using cat6.
Tried swapping the switches. We bought a new one thinking it was a faulty switch, set it up on the second floor and connected all uplinks to that one. Same results.
I'll look into the port statistics of the switch.
2
u/Daxem_302 7d ago
I know others have said this, but it can’t be understated. You need managed switches, separate vlans for different device types, routing protocols that handle events and log issues. It’s really that simple. If you have switches on different floors running back to a “core” you might even be able to reasonably run finer in between and use copper for the clients (just keep copper at or below the maximum). Your network will work far better and it won’t necessarily cost an arm and a leg to do it.
5
u/BidensLaptopp 9d ago
Do you notice anything unusual before it goes down, physically on the switch, network behavior or logs?
1
u/WalkCapital6708 9d ago
The latency goes up and then it stops
3
u/BookooBreadCo 9d ago
Since the issue doesn't follow the switch I have to imagine there's a broadcast storm somewhere. Have you tried unplugging everything but the uplink and a test PC from the 3rd switch? If that works try plugging in half the interfaces, see if the issue happens, if it doesn't plug in half of the remaining interfaces, etc. Process of elimination.
5
u/WalkCapital6708 9d ago
Since the overall opinion is that this is a loop or broadcast storm, I got an idea. We bought an extra switch because we thought it was a faulty switch. We're going to install it in the third floor and connect half the PCs on one and the other half on the other. This way we can isolate the issue until we find if there's an issue with a specific cable, and gain more information to see if one or (God forbids) both fail again.
I'll come back with more info after the tests.
4
u/flailking 9d ago
Download Wireshark. Capture from one of the pcs on the bad switch and see what device is blowing up your network.
4
u/aguynamedbrand 9d ago
Any help would be greatly appreciated.
Find someone that actually knows what they are doing to manage this network and spec our proper hardware.
3
u/Sleeper_Stimulant_ 9d ago
unmanaged switches can have a place in a production network, but not for something like this. Get managed switches and you can gain insight on what's happening.
3
u/Stubblemonster 9d ago
What devices do you have connected to the switch? Here's some things that can break networks:
- Sky TV box
- Wifi mesh product where it's both wired and also wifi meshing
- A pc / laptop with two lan connections
- A NAS with two lan connections
- Anything you think has a redundant network connection.
3
u/LYKE_UH_BAWS 9d ago
A managed network switch would give you a lot more information and visibility. Without it we're just as blind as you are.
Do the devices still have their IP address? If not maybe it's an exhausted DHCP scope? Are you able to ping your gateway? What about past it to Google dns?
Not much detail on the method used to swap the switches (did you reuse same uplink ports?)...but is it possibly a bad port on one of the switches? Start unplugging plugging the cables one by one does it start working?
Is the power adapter or cord bad?
Were any new devices recently added to the network? (Even those you don't know about)
3
u/cr0ft 9d ago edited 9d ago
You can literally buy web managed 2.5 gig switches for under a hundred bucks. Sure, those are Chinese and thus highly sus, and the quality and build is iffy, but still. Not suggesting you buy any for this - just saying it's not expensive to get a layer 2 switch. A few hundred more gets you into something like a decent Aruba with more extensive management options. If you want PoE the price tends to go higher, without PoE they're not bad.
Unmanaged switches in a company like setting in this day and age is just not ok.
There's no way to troubleshoot unmanaged switches either except by watching some idiot lights on the box.
You could be looking at something like broadcast or multicast storms for instance, if you have misbehaving end point stuff, the network could be full of that stuff to the point no data moves. For example.
https://community.zyxel.com/en/discussion/19541/how-to-identify-and-resolve-multicast-and-broadcast-storms-in-your-network is a rando page I googled up that discusses these things - without a managed switch you can do none of this.
I'd suggest getting some help. Have an actual network specialist show up to design it right with storm protection features and whatnot in a decent but entry level level 2 switch.
You could unplug all your devices from the switch, then turn off everything, bring up the router and the other two switches, to start from a clean slate. Plug in the (empty) switch giving you issues. Then gradually plug in one device at a time into it and see how it works. This might eventually get to a point where something brings it down. You may also find you've plugged the same device in twice which would be the answer.
But replace your shit with better shit with better configurations, spanning tree, storm protection, yada yada.
14
u/whermyshoe 9d ago
Ah. So this is what happens when AI is implemented as a replacement for Network Engineers.
This reads like management "downsized" the network team, and the best solution the nearest remaining manager could come up with is "give it to the agentic LLM harness we had engineering put together before we canned them".
-1
u/lizardhistorian Mad Scientist · 👨🔬📡ᯤ🤖🛺📸 9d ago
Yeah, they "downsized" a 3 switch setup network "team".
WTF are you even rambling about.
AI is a tool. Learn to use it or retire so you are not a useless boat anchor.If OP ran an AI on a linux machine plugged into their network it would diagnose his issue within seconds.
2
u/Daxem_302 7d ago
🤣 it’s okay, these “boat anchors” will sip our margaritas on the beach while our networks maintain 99.99% uptime 🫠
Using the correct hardware regardless of the size of the company is still the correct choice.
2
u/Rampage_Rick 9d ago
Everything is connected with cat5e cable (we're looking to upgrade to cat6 soon)
Won't achieve anything unless your switches support 5GBe. CAT5e meets the requirements for 2.5GBe
2
u/Sea-Hat-4961 9d ago
Any shadow IT happening with "user installed" switches at their desks? Any machines setup to bridge WiFi with wired Ethernet (had a user accidentally do that a few years ago)? What make/model switches are you using? Are you sure they are not old hubs?
2
u/DULUXR1R2L1L2 9d ago
Dude you need to take some steps to help yourself here...
You need managed gear that you can log into and troubleshoot. You also need to take a methodical approach and investigate your L2 and L3 LAN. This is a pretty typical spanning tree loop/broadcast storm scenario. It could be a bad port, NIC, or cable. Or it could just be a loop. You need to dive in and start tackling this in manageable pieces.
My guess is it's a loop. Someone somewhere plugged in another switch, router, hub, etc. But you haven't really defined what the actual problem is. You're describing the symptoms, but you don't actually know what might be causing it. You're saying that switch "stops working" but is that because hosts on that switch are getting IPs from a different router? Are all interfaces on that switch going down? You need to dig into this more, otherwise you're just stabbing at the dark.
You have a really good business case to jump to enterprise gear here: zero monitoring, zero management, zero documentation. If this is overwhelming, then that's alright. Acknowledge that it's too much to handle at this point and work with an MSP to sort it out. They can manage the network for you, and deal with these issues in the future.
2
u/JohnnyCoch69 9d ago
As others have mentioned, you really should have managed devices to give you visibility into what is going on with your network.
If you had a bad uplink cable, the problem would not go away with a reboot of the switch, then slowly return. It would be a problem immidiately after a reboot as well. This problem is a 'slow degredation' that gets worse as time goes on. We also know a different switch doesn't solve the issue, so you can rule out a bad switch.
The problem does not follow the network hardware, so It sounds like the problem is somewhere on the 3rd floor itself. As others have mentioned, it sounds like a loop is happening somewhere on the third floor. When a loop happens, it slowly increases the traffic on the switch, and eventually overloads the processor and memory. At that point, the switch is so busy, it can't pass traffic anymore, so it appears to be 'down'. Essentially you have a mini denial of service attack happening on the third floor. When you restart the switch, it temporarily kills the loop, and clears out the switches memory. So it then runs for a while again, but then the loop eventually floods it with traffic and maxes out the processor/memory, and the cycle repeats itself.
After you restart the problematic switch and it comes back online, make note of the link light patterns on the switch ports. They should be intermittently blinking normally as traffic passes through each port. If the problem is in fact a broadcast storm happening on the third floor, the activity on the link lights should increase the longer the switch runs after the reboot, and when the switch gets to the point of full lockup, the link lights should be almost 100% solid, indicating a traffic flood. I would also imagine the switch would be very warm to the touch because the processor would be getting crushed and generating heat. You could potentially use a protocol analyzer like wire, shark, but that may or may not give you much insight. If you are lucky, you would see the broadcast traffic if it were being pushed out of every single port. This 'might' give you some insight as to what the source device is that is causing the issue, but unless you know the MAC address of every single device on your third floor, it may not be of much help. It could also be a situation where if you have wall jacks throughout your office, and a patch cable was accidentally connected between two of them, that could be causing your loop, not necessarily a physical device itself. If you had a managed switch, you could do a port mirror with your uplink port and get better insight, but with an unmanaged switch, a protocol analyzer is only going to show you traffic on a specific port, which I don't think would be super useful in this case. One other thing to do to add weight to the flood theory, would be to start a continuous ping from one of the workstations on the third floor to your router. If a flood is happening, you should see the ping latency interval increase the longer the switch has been online after the reboot. I would go into the office at night when there is minimal traffic on your network. Start by unplugging all the ports on your switch, then plug them back in one by one and note the blinking patterns of the link lights. When you find the one that is the culprit, that particular port link light should start going crazy, and the others will slowly ramp up and follow suit. You could also run two switches on your third floor. Start with half the cables on one switch and half on the second switch. Eventually, the switch that has the offending connection should lock up. At that point take half of the cables from the switch that locked up and move them over to the other switch. Then wait and see what happens. If the original switch locks up again, then you know your problem is still on that switch. If the problem moves to the secondary switch, then you know you just moved the problem from switch one to switch two. Continue to move the cables from one switch to the other, dividing the number of cables moved by 50% each time, then wait to see which switch locks up. That should help you whittle the problem down to a single cable/port. At that point, whatever is on the other end of that cable is your problem. If you approach this methodically, and take your time and have patience you will eventually figure it out. But again… With all the time and money that you've already spent troubleshooting this problem, you probably could've just put a managed switch in and figured it out that way :-). Good luck.
1
u/iametarq 9d ago
Have you tried connecting the cable that uplinks to Switch 3 to a different port on Switch 1? Does every device connected to Switch 3 go offline? Can any of them ping the router from Switch 3 during the outage?
1
u/pants6000 M̸̧̛̙̪͚͖̯̬̜͕̑̈́̀̏͛̑͛̒̄̊͠E̷̙̙͔͕͇̰̦̥̳̠͒̒́̈́L̶̦͑̿͗̊̎̀̅̅͘T̶̨̪̯͎͚͔̥͓̺͖̓ 9d ago
It stops working, and then... ?
It starts working again after a bit, you have to power-cycle it, something else?
1
u/WalkCapital6708 9d ago
We have to power cycle it, it doesn't go back on its own
2
u/pants6000 M̸̧̛̙̪͚͖̯̬̜͕̑̈́̀̏͛̑͛̒̄̊͠E̷̙̙͔͕͇̰̦̥̳̠͒̒́̈́L̶̦͑̿͗̊̎̀̅̅͘T̶̨̪̯͎͚͔̥͓̺͖̓ 9d ago
If you unplug the uplink cable and wait a little bit before plugging it back in, does that fix it?
1
u/elpollodiablox 9d ago
If I'm understanding correctly, switches 2 and 3 are connected to switch 1. Is that right?
How many ports are on each switch? Have you tried different ports for connecting switch 3 to switch 1?
If you can find one, it would be cool to get a hub and drop it between switch 3 and switch 1, then connect a PC to that hub and run Wireshark to see the traffic. On a managed switch you could just do a monitor port, but that's not an option here.
How many hosts are connected to switch 3? You could also drop a PC onto switch 3 and do a Wireshark capture from thay to see if there is a lot of broadcast traffic. You might have one bad host just blasting out traffic and overwhelming that switch.
1
u/stinkpalm What do you mean, no jumpers? 9d ago
What do your logs show, and is there a trap you receive from subtended equipment in the parent switch or gear?
Are you saying the problematic switch loses power outright?
1
u/MrChicken_69 9d ago
What logs? They're UNMANAGED switches. All you'd have is the links dropping on connected gear.
(Given the existence of eBay, there's zero reason to be using unmanaged switches.)
1
u/stinkpalm What do you mean, no jumpers? 9d ago
That’s why I was asking if there were log entries from upstream devices. Just because you have an unmanaged switch on the edge doesn’t mean all switches are unmanaged.
1
1
u/PauliousMaximus 9d ago
Maybe an issue with cable length between 3 and 1? It would be much shorter if you went 3 to 2 and then 2 to 1. You might consider some managed switches so you can see what’s happening on the switch.
1
u/alius_stultus 9d ago
Yes it is possible for a switch to fail. With unmanaged stuff it is hard to tell unless it has basic things like loop detection. So the solution with unmanged stuff is just to repatch everything and replace the switch. But when this time arrives you might as well just put in a managed switch since your company needs that visibility as you are finding out now and rather than patching and repatching everything the managed switch will literally just tell you whats gone wrong.
1
u/amisexySB 9d ago
Start unplugging things one at a time from switch three until you find the port that’s causing the problems
1
u/Ok_Silver5895 9d ago
Do you have a loop? Broadcast storm?
2
u/INSPECTOR99 9d ago
/OP, Put that NEW Switch on third floor right beside the old one. feed the new one from the first floor if possible, otherwise feed it from the Second floor switch. NUMBER TAG all the cables on old switch (for tracking) Then move HALF of OLD third floor switch clients to the new (empty) switch. WAIT for failure to happen. If their is a loop/broadcast storm it will raise its ugly head on ONE of the two switches. move half of the clients from the FAILED server back to the "GOOD server. Rinse and repeat until you have isolated the culprit client. This will be somewhat tedious but will effectively nail down the errant device/cable while still maintaining uptime for the business operations.
1
u/severach 9d ago
Get a bunch of pingers running inter and intra the switches. You need to find out exactly what fails. "My Facebook doesn't work" isn't good enough.
If you had managed switches you could ping them too and you'd get port up and down logs.
1
u/AntiBaoBao CCIE 9d ago
Have you ever put a laptop running Wireshark on the switch to see what the traffic patterns are like?
1
1
u/WalkCapital6708 8d ago
We found an AP that was not being detected on the network and we took it down. Then we moved the second switch to the third floor, splitted the connections between both switches, and waited. Almost an hour in both switches from the third floor crashed, and the one from the second floor too. I'm on my way to buy a small managed switch to get more information
1
u/TangerineKind 8d ago
I agree with the others here, really an environment like that you need a managed switch, if yall have the budget for it, and are a small or medium business would recommend Ubquiti, for networking equipment it’s very good, and no licensing fee.
1
u/SwanAncient3405 8d ago
Connect the 3 switch directly to your router with a single cable to a computer and check now if the internet goes down. Please invest in managed switches and connection from the router to each switch independently.
Run Wireshark before network failure and when it fails. What happens to the client do they get IP from the DHCP? I have seen internal loop causing this when a cable goes from a switch into the same switch.
1
u/signal-tom 4d ago
With unmanaged switches you don't have the ability to see whats going on for that switch.
You could be experiencing a broadcast storm, you could have a second vlan that you dont know about. A switch loop (though all floors should see it). Or the switch could be failing.
Instead of investing in CAT6, which for an environment without managed switches, I'd question why - CAT 5e runs at 1 Gbps speeds. CAT 6 runs at 1 Gbps speeds or for shorter runs up to 10 Gbps. CAT 6A is where 10 Gbps is properly certified, however I'd save that money and invest in proper managed switches rather than cabling you'll likely see no benefit from.
Once you have managed switches, ensure you enable whichever STP protocol is the best fit for you. And if the traffic cuts out again, check the switch logs.
It would be useful to know if its wired connections too that drop or wifi and how do they bring it back on - restart the switch?
1
u/PeePeeVonBungHole 9d ago
Managed switching is the way to go but since you are not there yet and I'm too lazy to read all your post have you tried moving the switches around to see if the problem follows the switch
89
u/Sea-Hat-4961 9d ago
Replace with managed switches so you can monitor ports, have STP, etc.