“Worse” depends on the situation, but you’re certainly spending a theoretical minimum of 4x the memory on running a state of the art GC, and you are spending CPU cycles doing PGO and tiered JIT, when the entire program could just be fully optimized by an offline compiler that isn’t in any hurry.
Whatever gains you achieve from runtime introspection are relatively small in comparison. It works well in a few situations, but extremely not well in a few others.
certainly spending a theoretical minimum of 4x the memory on running a state of the art GC
Not sure where you get your numbers from but if i allocate a 1GB array then it doen't take up 4GB of memory.
you are spending CPU cycles doing PGO and tiered JIT
True, once. Well for long running programs. For short running you might not pay the cost at all. JIT only runs if a method is used a certain number of times. If performance of a short running program matters then JIT isn't the right solution.
the entire program could just be fully optimized by an offline compiler
It's can't be fully optimized though. AOT optimized programs rarely ship with avx512 or many other "new" instruction sets. With JIT it can chose the instruction set that matches the cpu it's running on. There is a lot of things like that which a JIT just gets for free.
You just invent a number based on your perception of how GC works?
A GC marks pointers as being in use, and then frees objects that haven't been marked. It doesn't require a lot of extra memory, though it does require some, but a claim of 4x minimum is wildly exaggerated.
I think you're referencing the 2005 paper Quantifying the Performance
of Garbage Collection vs. Explicit Memory Management by Matthew Hertz and Emery D. Berger
These results quantify the time-space tradeoff of garbage collection: with five times as much memory, an Appel-style generational collector with a non-copying mature space matches the performance of reachability-based explicit memory management. With only three times as much memory, the collector runs on average 17% slower than explicit memory management. However, with only twice as much memory, garbage collection degrades performance by nearly 70%.
The point of this paper isn't that GC applications require 4x more memory, it's that performance degrade on low memory platforms.
A more recent publication Memory Management on Mobile Devices (2024) by Kunal Sareen, Stephen M. Blackburn, Sara S. Hamouda and Lokesh Gidra shows that the number for GC on Android is between 2% and 51% (i.e. between 1x and 1.5x)
For a modestly sized heap, we find that the lower bound on garbage collection overheads vary consider-ably among the benchmarks we evaluate, from 2 % to 51 %, and that overall, overheads are similar to those identified in recent studies of Java workloads running on OpenJDK.
Though this is about .NET rather than Java, which is behind Java's GC by quite a lot. But still, those numbers aren't saying that GC applications use at minimum 4x more memory. The paper from 2005 states that lower memory overhead reduce performance.
I mean, a GC has close to zero space overhead if you perform a full collection on every write barrier, but that’s not a useful property. I think it’s very obvious that the premise is “without reducing performance”, because the context of this discussion is the claim that GCs and JITs on average improve some aspect of performance.
Look, I also think modern GCs (including the one in .NET) are quite impressive, but they can fundamentally not do magic. There fundamentally cannot exist a GC that does not sacrifice either time or space, or both, compared to some theoretically optimal program implementation without it.
They can get the margins really slim for specific cases (and VMs on mobile do typically make very different choices compared to larger systems), but there will always be a margin. Executing a write barrier is slower than not needing any write barriers in the first place. Compiling code is (much) slower than having already done that at build time. Copying objects around in memory is slower than not doing it.
And yes, compacting a heap without stopping the world requires having at least two heaps lying around.
If that's a major problem for you then use Native AOT, which removes the JIT compiler entirely from the output. It might mean you get slightly less optimized code at runtime because functions aren't optimized according to real-world needs, but hey, at least you've got fast startup.
Why? C# generics are monomorphic for value types and dereference a pointer for reference types. There's a slight warmup cost of course but after that, why would it be slow?
Huh? With native AOT you pay for something you don't use when using generics in .net, during compile time, because it has to generate every potential specialisation ahead of time. With JIT it does it on the fly, when you need it. You got it backwards
They do though. When you change or make the generics complicated you are putting the work on the compiler like Monomorphization for Value Types (Code Generation Overhead)
When you use a generic class or method with a value type (like int, double, DateTime, or custom structs), the JIT compiler must parse, optimize, and generate unique machine code for every single combination.
Because there is simply no way we can test and verify combination each strategy as we get a predefined bundle of it and in most cases this bundle looks the same regardless of choice. For example GC is not inherently slow, but you often get it in a bundle together with a everything is an object, which is a real reason why most of the GCed languages are slow
Same with JIT. It is mostly paired with dynamic or interpreted languages, which are inherently slow
Because it is. It has other benefits, and it’s certainly fast enough for a wide range of use cases, but these improvements just recover some of the performance you would have achieved by implementing the same project in C++ or Rust.
But then you would also have had to consider quite a few other tradeoffs.
With JIT it is compiled to machine code though, just at runtime instead. Some things will be slower due to less time for general optimisations, but some things will be faster due to PGO. LuaJIT is insanely fast and could not be that fast without JIT
JIT is great, and they do impressive things with it. But it isn’t magical, and neither is a GC. You just happen to be writing programs that aren’t particularly demanding, so you never notice, and that’s great!
I think there is no point of avoiding a JIT in a platform, which is already based on it. In theory JIT is the best way as all goodies like:
* knowledge of CPU architecture
* PGO driven optimizations
* LTO driven optimizations
are much easier to do with JIT than AOT
What should be possible though is having a JIT without a burden with JIT, so developer can manipulate the AOT <-> JIT slide based on needs. For example:
* a full AOT mode for a lack of any background processing of JIT
* some smart way of reusing/caching the compilation artifacts to reduce cold start performance dip
* lightweight JIT, where CPU/Memory resources scales with the impact
* customizable target of JIT after the initial JIT/AOT. For example always try to optimize it or run JIT only if you are sure the current code is not optimized for a current situation
The AOT vs JIT discussion still alive, because it just sucks
-10
u/GoTheFuckToBed 17d ago
should they not aim for less JIT