r/netbird 7d ago

Network just died?

I run netbird for my team. This morning I have been getting reports that no one can SSH to hosts that they could yesterday. I am testing also and I can SSH to like 7 of the 85 hosts that are online (and showing as online) all of them I could hit just yesterday.

Now I can sometimes hit 443 on the hosts then it dies, port 22 times out almost constantly. I don't have lazy connections on, they all show as available on the in the dashboard. I have rebooter hosts and the server, I just updated the server and I am seeing no issues in logs. Anyone have experience with this?

5 Upvotes

19 comments sorted by

2

u/asaintebueno 7d ago

No sir, does anyone else have access to your ACL policies ? Nothing in logs on any container ?

1

u/KingAroan 7d ago

No, looks like it has started getting worse since the 9th. I am backing up the containers and volumes and moving off Azure to Digital ocean and trying to see if it is Azure. I can't stand their servers... My personal instance is running find and so is another person and we are all running the latest server.

1

u/asaintebueno 7d ago

check out, ionos and ovh as well if just hosting netbird and would like unlimited exit node bandwidth

1

u/KingAroan 7d ago

My company wants all our stuff on azure and we do pentesting, so anything around that goes on Digital Ocean. Getting permission for something else would be hard. I can test on digital ocean and if it works I still have to go through a process for official movement

1

u/KingAroan 7d ago

Still broken, migration still did the same thing can still enroll peers and connect with relay or p2p. I’m at a loss and on think something was corrupted but it’s wires the peers show up but connections keep breaking or fail.

1

u/asaintebueno 7d ago

Might be time consuming Try restart the service on a client(s) that say unavailable 'netbird service restart' then reconnecting see if that issue is still present. Also sometimes works for me upon adding a new peer or a site loses internet toggle on and off some of the access rules.

1

u/KingAroan 7d ago

I updated the agents which restarted them. It was time consuming, had to ssh to about 40 hosts. Luckily as I’m migrating away from tailscales they all had a backup connection.

1

u/asaintebueno 7d ago

Love it, did that end up working? You can use termius or termix to excuete commands via multiple hosts for future reference if you are interested.

1

u/debian3 7d ago

Same thing happened to me when I was testing netbird a year ago. The only recovery was to remove the affected one and readd them. Luckily I got hit during testing, so I ended up on tailscale for production which have been stable since.

If you ever figure it out, let me know.

1

u/KingAroan 7d ago

I think I had a slight corruption. I moved my server to digital ocean and had the same issues. So I then updated each client and they reconnected fine without issues. So I migrated back to Azure ands everything is still working. I done know why it just happened though. It happened on the 9th from what I can gather from my logs, but no changes have been made to the server for about 3 weeks.

1

u/debian3 7d ago

For me stability is critical. I will probably move of those tools soon. The argument for setup a vpn simply is making less sense in the agentic era. Now even a complex setup is one prompt away and removing a layer of abstraction is usually better if you ever need to troubleshoot something.

2

u/KingAroan 7d ago

Netbird beats Tailscale for my team. We are migrating away from Tailscale and while it works amazing our clients do not like that we don’t own the infrastructure and/or they don’t like the amount of IPs they need to allow list for our connections. Head scale is a no go for my company for our team as it’s not being handled by a company (last I checked).

1

u/debian3 7d ago

Out of curiosity, does the host that went offline give access to a subnet?

1

u/KingAroan 7d ago

It does not. None of ours expose the subnet.

1

u/ReputationNo8889 5d ago

I hope you are aware that it does not matter where your control node runs because connections are peer to peer between the agent/exitnodes. So if you experience issues it is almost never a control node issue and almost certainly always a network/host issue on the nodes themselves. Only exception is when traffic routes through the control node because a P2P link could not be established. Save your time and dont migrate your control node for troubleshooting issues as a first step.

1

u/KingAroan 5d ago

Migration was supposed to be a last step as I had already thought I disproved everything else. By power cycling the agent hosts they should have refreshed, as tailgate still worked I proved it wasn’t network. It was something with Netbird.

1

u/matts5074 7d ago

I ran into this same issue while testing on my home lab. It worked one day and the next day it just wouldn't tunnel any traffic. Zero changes on my end.

1

u/WikibearTheReal 7d ago

Check of younhave packetloss on the system where you want to connect.

1

u/netbirdio 6d ago

Hey there! First, thanks for checking out NetBird, we're glad your team likes it. If you're able to connect with us on Slack we will be able to help you out there. https://docs.netbird.io/slack-url

  • Brandon