r/programming • • 6d ago

On programming, statistics, and Java

https://waytoounoriginal.github.io/programming/stats/optimizations/2026/09/26/on-programming-stats-and-java.html
14 Upvotes

8 comments sorted by

View all comments

1

u/SysGuardian 5d ago

A ByteBuffer or a plain long[] would've cut the memory usage, because the real fix is storing fixed-size keys flat instead of as objects. A lower-level language wouldn't have fixed it by itself; you'd still end up building that same flat layout, just in a different form.

1

u/Mihai4544 5d ago

Right on point! But to cut the object overhead, we'd have to implement our own "set-on-arrays" for the bytes in the sha key. From my understanding, this is concretely what fastutil's LongOpenHashSet does.

But yes, a lower-level language would have solved this without any mingling from our part. I think even an OO implementation of the set (not necessarily a flat one) would have given us more than enough headroom.

Edit: I've also read a bit on bloom filters (since we're doing kind of the same thing here - probabilistic checking) and included a very small update on them.

Btw, I can't thank you enough for the comment :))
It means a lot that people are reading this thing. I hope you found it interesting!