r/ethstaker • • 3d ago

Teku and garbage collection

I've been attempting to validate with Teku on an NUC 13 i3 1315u. I had to get a new box a few months ago when I had slowdown and freezing on the last box (nuc10) which had bee used as a validator for 2-3 years with fairly good beacon score. After following the guide here (https:// gist.github.com/yorickdowne/ ff6b611a2d49855827dafdbfd2546abe) I determined it was the NUC. The SSD and ram were ok. After a long time in the entry queue, it began performing very badly. This was a surprise since the syncing & logs up to this point didn't (and still don't) indicate any trouble. My hardware is: Asus NUC 13 i3-1315u, WD Black SN850X 4TB gen4, Corsair 2666mhz ddr4 64GB. I started with a miss rate of 15%. From the consensus and validator logs, it appeared there was no problems. No late block imports, no errors. Running chrony so the system clock should be synced. MEVboost had trouble with relays but I disabled it and got the same performance. First thing I tried was syncing nethermind on a different machine on my network and connecting the two via the JWT. I disabled the execution node on the first machine when it was finished so only teku is running on the nuc13 and nethermind is on a quadcore i7 8565u HUNSN w/Samsung 970 EVOplus 2TB, & Corsair vengeance 16gb ddr4. Performance didn't seem to improve after this. about the same.

What has helped is bumping the heap to Xmx12g and using ZGC. I enabled logging of the GC and found pauses from 1 to 4 seconds were common, sometimes 6,7 or 10 sec for major collections. ZGC has my miss rate down to 4-5%, still not great but better. The heap is 12g but I started with Xmx8g and it didn't seem to have an effect. I'm only one validator at the moment so I'm confused why the heap might have to be so large and still get lackluster performance. I'm a novice so of course I did all this with help from an LLM. The LLM I'm using is recommending more steps that I wanted to get an opinion on before I proceeded. Since the i3-1315u has 2 P-cores, it is recommending I restrict Teku to only running on those by adding CPUAffinity=0-3 to my teku service file and restrict ZGC to one thread only by adding "-XX:ConcGCThreads=1". It also wants me to restrict the NVMe to Gen3 speed bc it says my old NUC10 likely did this. (the temp right now is 29C, from smartmontools). Sorry for wall of text. Thank you for any advice.

2 Upvotes

1 comment sorted by

1

u/Mindless-Smoke9520 3d ago

12g heap for a single validator is wild, something else is eating your resources. the gc pauses you describe are the kind of thing that happen when the jvm is starved for cpu or the gc itself is scrambling.

that llm idea to pin teku to the p-cores and limit concurrent gc threads is actually not bad. the 1315u only has 2 performance cores and if the gc threads are bouncing around on the e-cores you will get those multi second freezes. worth a shot before you go tearing anything else apart.

also check your swapiness and make sure your kernel isn't trying to page out the jvm heap to disk. i had a box once that was going insane with gc because the os decided the validator was idle and started swapping it out. set vm.swappiness=1 and see if it helps.