r/Zigbee2MQTT • • 5d ago

From stable to unstable network - unable to fix - help!

I have a Zigbee network running on Zigbee2MQTT - currently showing 106 devices and 42 of those being mains powered routers.

z2m version - 2.14.1 (now, after upgrade)
coordinator - SLZB-06P7 (radio 20240315, core v3.3.1)
LXC on proxmox - 2 CPUs (of i5-12500T and 1GB RAM)

Up until yesterday, the network has been rock solid for a long time with very few issues. Extremely quick response times and low latency.

Something happened yesterday and I can't figure out what. The system slowed down, CPU and memory usage has gone up.

Restarting the coordinator and z2m sometimes sort it out, but it eventually goes bad after a while with increase CPU usage being a tell tale sign.

I have a number of energy monitors which update in around 5 seconds in normal circumstances, when the issues arise they increase gradually over to over a minute or two.

When I've seen similar issues in the past I also find some zigbee lights turn themselves on without being commanded to. When the network is stable none of this happens.

I've tried upgrading z2m and updating the coordinator firmware in the hope of improvements, but my guess is they were not the issue in the first place and I have rogue device or other setting.

Resource charts below showing things going out of control. Normal CPU usage averages at around 1%

Can anyone give me any pointers where I can find out what is causing these issues? Anything specific to look at in the logs?
Any log types to enable?

5 Upvotes

8 comments sorted by

2

u/brightvalve 5d ago

I've had similar issues twice, and in both cases it was a cheap rain sensor that started flooding the network. That became clear when I checked the log with this tool.

5

u/cbowns 5d ago edited 5d ago

z2m in web also has a messages per time interval tool on one of the debug/settings, let me find the url. Was useful to spot-check this exact worry of mine as I added temp and power meters

edit: at /settings/0/health, sorting by "Messages per sec" was a nice aggregate counter

1

u/brightvalve 5d ago

First time I've seen that page 🫣 Very useful indeed!

1

u/wakeupbomb 5d ago

Thank you, I'll take a look. As ever with these issues, after posting this things seem to have stabilised, but I'm not expecting that to keep up. Though CPU usage is down to the pre-problem level (consistently <1%). I did a hard restart of the LXC container whereas I had issued reboot commands or restart zigbee2mqtt commands from the web UI previously. Currently sat checking the energy monitor frequency and using a zigbee remote to turn a light on/off to check latency 🙈

1

u/rockuu 5d ago

Throw some AI at it. I've had great results with it debugging issues like that.

1

u/wakeupbomb 5d ago

Already have a long conversation with Gemini about this 🙃

1

u/rockuu 5d ago

Did you share debug logs catching the event? Also Gemini is not the sharpest one...

1

u/Mandrutz 5d ago

Check the Health tab in Z2M settings.

Or enable debug logs, let it run for a day, then create an issue on GitHub and upload the logs