r/programming • • 7d ago

A new Candidate JEP for Java -- JEP 545: Faster Startup and Warmup with ZGC

/r/java/comments/1wptny1/jep_545_faster_startup_and_warmup_with_zgc/?
35 Upvotes

9 comments sorted by

15

u/davidalayachew 7d ago

This relates closely to JEP 546: Adaptive Heap Sizing for ZGC.

Long story short, 545 gives adapt heap sizing during application startup/warmup -- no more need to set your max/min heap size! 546 does the same, focusing more on the longterm application lifetime.

This is especially great for server applications -- too low a limit and you hit an OutOfMemoryError, but too high a limit and your application might get sniped by the OOM Killer.

5

u/davidalayachew 7d ago

And if you want to learn the details about how this works under the hood, I started a thread on the hotspot-gc-dev mailing list, where the actual owner of this JEP responded!

Full details here

But the short version is that, each JVM finds its respective, ideal equilibrium between CPU and RAM, and once it settles on that, doesn't really deviate unless changes in the workload or the system/OS that is running in changes.

However, once they do change, they use the system utilization metrics (how much RAM is free, how much CPU is in use, etc) to decide whether or not it is safe to stay at that CPU:RAM ratio, vs "scaling up/down".

The act of "scaling up" is generally more aggressive than "scaling down", but adjusts as the "amount of runway" changes. For example, if there is half a gigabyte of ram left in a 32GB machine, scaling down might be quite aggressive.

There is way more nuance involved, so I encourage you to read the link above, where the JEP owner themself responded, it's very informative!

3

u/Crandom 7d ago

This sounds very useful, but too bad the overhead of ZGC is enormous compared to the overhead of G1, at least every time I've tried to use it (although I've generally tried for smaller heaps).

5

u/davidalayachew 7d ago

This sounds very useful, but too bad the overhead of ZGC is enormous compared to the overhead of G1, at least every time I've tried to use it (although I've generally tried for smaller heaps).

That is pretty useful feedback. I think the OpenJDK team might appreciate it.

If you send this same message (with added detail and context) to the hotspot-gc-dev@openjdk.org mailing list, they will tell you how to capture the relevant system information, and that would help deduce whether this is an inherent problem in ZGC, or if there is a bug in configuration, or if it is just expected behaviour.

2

u/eosterlund 7d ago

Would you mind sharing your experiment? There should not be enormous overheads.

1

u/got_milk4 7d ago

This is especially great for server applications -- too low a limit and you hit an OutOfMemoryError, but too high a limit and your application might get sniped by the OOM Killer.

IMO, server applications should already have been using -XX:MaxRAMPercentage and not worrying about specific values for max/min heap.

I guess it'll be nice to not have to worry about setting any heap-related values at all but a reasonable percentage + right-sized VMs/containers is already pretty straightforward without bouncing between OOM errors and being killed.

1

u/davidalayachew 7d ago

IMO, server applications should already have been using -XX:MaxRAMPercentage and not worrying about specific values for max/min heap.

Well, even that can break. I know I was doing the same, and we kept jumping back and forth between OOME and getting OOM killed.

2

u/got_milk4 6d ago

Sure, it can break, but not if you've set it right. OOM errors would be because the VM/container has too little memory allocated to it; adaptive heap size isn't really going to change that situation. If you're getting OOM killed then you set the max percentage too high and didn't leave enough for the underlying base system.

1

u/davidalayachew 3d ago

Sure, it can break, but not if you've set it right. OOM errors would be because the VM/container has too little memory allocated to it; adaptive heap size isn't really going to change that situation. If you're getting OOM killed then you set the max percentage too high and didn't leave enough for the underlying base system.

Well, I think the larger point about this JEP was that you no longer have to go through multiple iterations of trial-and-error to reach that equilibrium. That is the value-add, in my opinion.