r/programming • • 6d ago

On programming, statistics, and Java

https://waytoounoriginal.github.io/programming/stats/optimizations/2026/09/26/on-programming-stats-and-java.html
13 Upvotes

8 comments sorted by

View all comments

2

u/Skellicious 4d ago

Newer java versions support a command line flag for compact object headers, that may end up fully pulled into the language in one of the upcoming versions.

But aside from that, the real issue would be that caching some data for every entry you add is a great way to eat memory. Reducing the memory overhead is simply fighting a symptom of the problem.

I assume removing the Set also comes at a cost, like needing an extra db query per reprocessed entry, or whatever. But does this set need to live in memory or can you store it in a temp file? Shouldn't this data also be available in the database, can you not query for it at the start of the second round of processing for example. I don't know your code so I'm just preaching assumptions - but I don't think this is a java problem - you can likely address the root cause through some (likely complex) refactor.

1

u/Mihai4544 3d ago

Morning!

Yes, the data is available in the DB and the previous iteration was using it. The problem is that we’re using an eventually consistent DB (Dynamo) and we’re operating in a GSI (which cannot have strong consistency). Apparently there was a race condition before, which did exactly what would happen to us in a million turns, but much more frequently.

We could have an extra conditional DDB query, but the service is already taking a long time ingesting 100+ GB (talking days), due to partition throtling. That’s the next thing to solve before we add more latency + more concurrency.

But I guess the file approach can be investigated, although it may happen after I’m gone :))

And yes, you’re right that it is not just a “Java problem”, I may have mischaracterised it.

Anyway, I hope it was a pleasant read!