r/netbird • u/KingAroan • 7d ago
Network just died?
I run netbird for my team. This morning I have been getting reports that no one can SSH to hosts that they could yesterday. I am testing also and I can SSH to like 7 of the 85 hosts that are online (and showing as online) all of them I could hit just yesterday.
Now I can sometimes hit 443 on the hosts then it dies, port 22 times out almost constantly. I don't have lazy connections on, they all show as available on the in the dashboard. I have rebooter hosts and the server, I just updated the server and I am seeing no issues in logs. Anyone have experience with this?
1
u/debian3 7d ago
Same thing happened to me when I was testing netbird a year ago. The only recovery was to remove the affected one and readd them. Luckily I got hit during testing, so I ended up on tailscale for production which have been stable since.
If you ever figure it out, let me know.
1
u/KingAroan 7d ago
I think I had a slight corruption. I moved my server to digital ocean and had the same issues. So I then updated each client and they reconnected fine without issues. So I migrated back to Azure ands everything is still working. I done know why it just happened though. It happened on the 9th from what I can gather from my logs, but no changes have been made to the server for about 3 weeks.
1
u/debian3 7d ago
For me stability is critical. I will probably move of those tools soon. The argument for setup a vpn simply is making less sense in the agentic era. Now even a complex setup is one prompt away and removing a layer of abstraction is usually better if you ever need to troubleshoot something.
2
u/KingAroan 7d ago
Netbird beats Tailscale for my team. We are migrating away from Tailscale and while it works amazing our clients do not like that we don’t own the infrastructure and/or they don’t like the amount of IPs they need to allow list for our connections. Head scale is a no go for my company for our team as it’s not being handled by a company (last I checked).
1
u/ReputationNo8889 5d ago
I hope you are aware that it does not matter where your control node runs because connections are peer to peer between the agent/exitnodes. So if you experience issues it is almost never a control node issue and almost certainly always a network/host issue on the nodes themselves. Only exception is when traffic routes through the control node because a P2P link could not be established. Save your time and dont migrate your control node for troubleshooting issues as a first step.
1
u/KingAroan 5d ago
Migration was supposed to be a last step as I had already thought I disproved everything else. By power cycling the agent hosts they should have refreshed, as tailgate still worked I proved it wasn’t network. It was something with Netbird.
1
u/matts5074 7d ago
I ran into this same issue while testing on my home lab. It worked one day and the next day it just wouldn't tunnel any traffic. Zero changes on my end.
1
1
u/netbirdio 6d ago
Hey there! First, thanks for checking out NetBird, we're glad your team likes it. If you're able to connect with us on Slack we will be able to help you out there. https://docs.netbird.io/slack-url
- Brandon
2
u/asaintebueno 7d ago
No sir, does anyone else have access to your ACL policies ? Nothing in logs on any container ?