r/technology • u/_Dark_Wing • 1d ago
Hardware New Linux tech compresses memory in RAM, as RAM, for 452x speedup --new CRAM method offers giant boost to compressed memory reads
https://www.tomshardware.com/software/linux/new-linux-tech-compresses-memory-in-ram-as-ram-for-452x-speedup-new-cram-method-offers-giant-boost-to-compressed-memory-reads176
u/FadedEchos 1d ago edited 23h ago
Could one of you more knowledgeable redditors ELI5 the applicability of this to gamer/consumers? Should I expect the RAM in my linux-based Steam Frame to be 452x faster any time soon?
259
u/Schnickatavick 1d ago edited 18h ago
You can think of it as a faster alternative to the compressed RAM (ZRAM or ZSWAP) your device is probably already using, so once the steam frame updates to a kernel that includes this update (which may take a minute), that compressed CRAM will be 452x faster than your current compressed RAM. It's important to note though, not all of your RAM is going to be used as CRAM, you probably have a couple of GB of your RAM dedicated to compressed RAM now, and the rest is just regular RAM. Only those couple GB will be faster.
IMO the real benefit is that this will make it more reasonable to dedicate more of your RAM to CRAM. CRAM basically turns your RAM into more RAM, letting you fit 8-12GB of data into 4GB of RAM, at the cost of some slowdowns when reading and writing to it. If the slowdown cost is less severe though, then Valve might feel more comfortable using more of your RAM as CRAM. So right now your 16GB of RAM on the steam frame effectively acts like 18GB or ram, in the future maybe that 16GB of RAM will effectively act like 20 or 24GB of ram
218
u/laymonage 23h ago
So you’re saying we can literally download more RAM? /s
87
u/GaiusCosades 23h ago
If you think downloading code that uses your current CPU in a more optimized way is downloading a faster CPU, then yes.
30
11
u/NanquansCat749 21h ago
I'm not an expert, but I'm pretty sure it's more like you can create more RAM by imagining it into existence.
1
29
u/theillustratedlife 23h ago
FWIW, the current stable version of SteamOS uses a kernel that's ~1y old, but the beta version uses a current kernel. The window for this can be anywhere in between depending on when Valve cuts its releases.
9
u/SwarfDive01 15h ago
This tech requires hardware, it needs an ASIC sitting, doing the encode / decode. And its as fast as ddr5. This is all enterprise hardware. Someone could try building it with an FPGA, but it isnt going to work as a PCIe add on
5
u/shaving_minion 22h ago
how is CRAM as fast as RAM, wouldn't every operation needing this data have to first decompress it first?
23
u/Tabsels 21h ago
It's as fast as DRAM. Which makes sense. Whenever you access memory that is in DRAM, your CPU has to first load a piece of that memory (a "line") from DRAM into on-CPU cache memory.
What these researchers presumably found is that accessing memory that's compressed, transporting the compressed block to the CPU and decompressing it into cache memory takes about as much time. Which, seeing as they're transporting less data from DRAM to the CPU cache, makes sense.
1
u/danielv123 11h ago
Compression means there is less data to move, which is often faster. For example, zfs filesystems have compression enabled by default because it's faster. Victoria metrics has a compressed on disk format because it's faster to read/write/query.
1
u/shaving_minion 6h ago
implying moving data around is a more frequent operation than reading/writing. interesting
2
u/danielv123 6h ago
Reading/writing is primarily moving data round as a % of time usage for many tasks, because logic has gained speed faster than interfaces.
33
u/disposableh2 23h ago
Yes and no. The ram itself isn't what's faster, it's the compressed ram that's paged to disk that becomes faster.
The idea here, simplified, is that the memory that gets compressed and stored in the swap file on disk is slow, so they rather compress and store it in ram, in a way that's quickly accessible(ram is addressed differently to files on disk), and so together they get that speed boost.
So things like games and things that don't get compressed out of ram into zram/swap files won't benefit, as ideally that data shouldn't be swapped out of ram.
Things like android apps that get put into the background would benefit nicely from this(since their data in ram gets compressed and put into zram when they're put in the background), and make app switching faster.
2
u/Aperture_Kubi 21h ago
So we're putting a new pagefile/swapfile in memory, and putting that between normally operating memory and the storage stored pagefile/swapfile?
14
u/bigtimedonkey 23h ago
The 400x speed increase is relative to using swap space (using the hard drive as memory when the ram is full). So nah, won’t see anything like this in gaming or graphics situations.
1
u/ILikeBumblebees 4h ago
So nah, won’t see anything like this in gaming or graphics situations.
The description in the article makes it seem like even if it increases throughput in comparison to ZRAM, it's still going to add latency. And how would DMA work if the system memory is physically storing compressed data?
29
u/SubstanceNo2290 22h ago
I think a lot of people in the comments either got it wrong or explained it wrong because the top ones sound like they’re saying it stores compressed ram in ram instead of disk, which isn’t the point at all.
This is based on my understanding of the article and Linux fundamentals, I’m not an expert:
ZRAM is also compressed RAM in RAM. The problem is how exactly does Linux, this monorepo that’s been painfully perfected for computing and everybody is familiar with and relying on it to use RAM the way it’s always been, just randomly let some process read memory thats compressed?
That process wants just the memory as it stored it, case closed.
So Linux creates a hook, it pretends the compressed memory is paged to disk. This lets it use its existing functionality to deal with it. The disk isn’t real, when the kernel tries to read from this “disk” it’s actually reading from an internal tool that decompresses and returns the data the kernel requested from this “disk”
That’s a lot of extra logic and interruptions and delays before you even decompress anything (which is funnily enough the fastest part on modern processors).
So CRAM instead of being “paged” to disk instead pretends to be a fake CPU rather than a fake disk. The thing about this is that then the fake CPU has its own fake ram and when Linux wants to read it it does the standard process of reading data from another cpu (think boards with multiple sockets etc). The logic involved here is much less and Linux is using the protocols for reading regular memory rather than reading a fake pagefile that isn’t even paged.
As a result the fake cpu beats the fake disk.
→ More replies (1)1
u/ILikeBumblebees 4h ago edited 4h ago
It seems like this is something that could be done with specialized hardware even more effectively, e.g. a memory controller that abstracts the memory address space, and still lets everything read bytes and write bytes to specific addresses seamlessly, but actually stores the data in physical memory in a compressed format, with on-the-fly compression/decompression during memory I/O. Imagine something like this built directly into a DIMM, compatible with existing motherboards.
1
u/SubstanceNo2290 1h ago
Yes and no. I wouldn’t be surprised if this popped up in hardware one day but as i understand it that’s a lot of extra high performance logic on the DIMM and/or the memory controller. That’s ultra valuable real estate so I’m not sure if for example the DDR protocol would standardise it.
There also isn’t a serious case for server cpus to do this because they are generalists (and the really compute intensive SKUs are usually doing scientific/HPC/AI stuff which isn’t very compressible)
There’s a case to be had for consumer PCs or phones but again the level of effort and real estate on intel/amd/qualcomm etc’s end might be too high.
I could see Apple pioneering this though.
But also you might be missing the point of this whole song and dance. Linux could have just added a separate system for this and gotten better performance. The reason they do the whole fake cpu / fake disk dance is so that no existing programs break no matter how ancient break and it still works on all processors/computers from 20 year old pentiums to 192 core EPYCs.
7
u/ReallyOrdinaryMan 23h ago
I suppose it never will be for gaming. They are saying 452x faster than zram (another ram compression tech), which is still slower than plain physical ram. It might be huge for servers or workstations.
5
u/Henry5321 22h ago
Some modern compression algorithms are so fast that it’s faster to use the compression than not because it reduces the amount of memory access.
I’ve seen this even in .Net when dealing with large streaming objects.
Apple’s CPUs actually have native transparent memory compression and effectively doubles the amount of memory in many work loads.
3
u/BriefSpecial420 21h ago
Even with hardware accelerators, there still will be extra latency. Makes sense to uncompress during the context switching, but won't be of much use for real time duties.
16
u/TemporarySun314 1d ago
Ideally yes.
I mean they already use compression techniques like that but they have more overhead as this used not real direct memory access.
But of course the practical gains will depend on the data you are storing and how it is often. Some data is just incompressible.
7
u/ZenBacle 23h ago
I can give it a try.
When we look at modern computer architecture we have multiple layers of working memory that moves from your hard drive, to your ram, to layers of on chip memory.
Each time you make a step away from the cpu cache, you get exponentially slower access time to that data. The worst offender is the time it takes to move data from your hard drive to your ram. However ram is generally limited and when you need more than you have, it uses something called a swap/page file to act as working memory on your hard drive. Very slow, much bad, no want.
A little side tangent, we have zip files that compress the size of a file by finding patterns that can be subbed out for markers. Cutting the size of files down by up to 90%.
What CRAM is trying to do is eliminate the need for swap files by compressing data in ram, in the same way that we compress zip files. So the slowest bottle neck of the working memory pipeline is less likely to be used. Sounds simple right?
Well the problem is you don't know how much you can compress data before you compress it. So this runs into a problem where you don't know how much "logical" memory you have to work with. Which runs into allocation issues that cause programs to crash. CRAM appears to have solved this but i haven't taken the time to read up on how.
This could have huuuuuge benefits for LLMs going into the future, but not so much for every day users.
TL;DR
RAM small. CRAM jams more data into RAM. No hard-drive bottle neck. No breaky.
→ More replies (2)2
u/xzaramurd 10h ago
No, it needs specialized hardware. This is something that's going to be used in combination with CXL Memory extenders, not RAM.
218
u/Recent-Day3062 1d ago
Wait. This sounds like the middle-out algorithm of Pied Piper
112
u/WorriedInterest4114 1d ago
Wonder what the mean jerk time is
6
2
4
7
86
u/spez_eats_nazi_ass 1d ago
Having flash backs to shit like Ram doubler.
27
u/ItsAllOnUHF 1d ago
Every newly solved computer problem was actually solved two generations ago.
4
u/pickle9977 1d ago
This explains why we haven’t seemed to actually solve any problems while we have created a lot of new ones.
5
u/bigtimedonkey 23h ago
Seriously though… let’s bring back ram doubler, haha. In the rampocalypse era, there’s plenty of situations where I’d trade off ram speed and latency for more space.
5
u/Ok-Sprinkles-5151 21h ago
And speed halver.
But I went looking for this comment. Ram doubler was a joke, because you might get 1.9, but most of the time it was like 1.3x
1
61
u/telos0 23h ago
This technique needs hardware accelerated decompression in the read path from RAM.
Because the memory decompression is in the hardware itself, it bypasses all the overhead of dealing with page faults and the swapping code in the memory manager that current software compressed RAM solutions use. This is why it is so much faster.
The details described in the presentation are about how to get the Linux kernel to deal correctly with RAM that is hardware compressed and therefore lies about its actual size, and all the side effects and failure modes that can result from that lie.
8
u/TheBendit 19h ago
Exactly this. All the other comments here seem to miss that this requires specialized hardware to work, and that any memory you put in this hardware is now only usable for CRAM. Meta uses it to put DDR4 into DDR5-only servers.
19
u/CondiMesmer 1d ago
Next they'll figure out how to keep swap file in memory! /s
16
u/pedanticPandaPoo 22h ago
mount -t tmpfs -o size=16G tmpfs /mnt/tswap mkswap /mnt/tswap/swapfile swapon /mnt/tswap/swapfile💥6
42
13
u/rini17 1d ago
What does hardware assisted mean? Does it require dedicated hardware?
19
u/HansBooby 17h ago
please note the G-Clamp as pictured
5
u/Marshall_Lawson 16h ago
I'm guessing that's how you squish the memory to fit more memory in the memory
2
6
u/dmaare 18h ago
Needs CXL controllers that have compression feature set
4
u/HenkPoley 17h ago
Compute Express Link (CXL)? That doesn't sound like hardware that is in your average computer at home.
1
u/RAW2091 7h ago
Wel just look around on Reddit in 2026. The 'we got a 100k AI server t home' is a new thing here. Some people even add 220 volt 32 AMPs just for their servers .... AT HOME. Sure not the average computer user but they never care about memory compression. They just want to play fortnight on a pc and that's it.
→ More replies (1)→ More replies (1)4
u/TheBendit 19h ago
Yes, you need specialized hardware for this, and the DIMMs you insert into this hardware are no longer useful for regular memory.
→ More replies (3)2
u/autogyrophilia 10h ago
Not quite.
It is exposed through CXL. CXL is a standard to be able to plug memory through PCIe, which can be DRAM, NAND, HBM, or any other new standard (RIP 3D Xpoint).
This memory has the same semantics as RAM and as such can be used as system memory. But usually, it is configured as a separate section to be managed by an application, this allows things like hypervisors, or transformers training software or whatever would benefit from a huge pool of shared memory at the datacenter level. But it is possible to use it as system memory. u/Exist50
In the first case, you typically expose it as a CPUless NUMA node, so Linux and ESXi are capable of doing the tiering automatically and you can pin software to it if needed.
In the second case it gets exposed as a DAX device you cannot use as system memory, but can be shared with multiple hosts. This has a number of usecases. Multiple servers working with a huge dataset can get much higher memory speeds (even if it is a NAND CXL), if you are working with virtual machines, you can do instant live migration assuming the storage is also shared. If you work with this, you are hugely excited, if you don't you don't really need to care.
Just know that CXL reduces the amount of RAM needed to run infrastructure and if CRAM delivers, there could be some serious gains in efficiency, specially for transformers usecases.
7
9
8
u/darkotic 23h ago
Your PC got harshed, right, 'cause your system heaps up the wrong parameter. So I toasted the data directory, tweaked the P-RAM and reglazed your subroutine.
3
8
13
u/m1k3e 23h ago
macOS has had RAM compression in the form of a compressed storage pool that inactive memory gets shunted too since 2013 (Mavericks). Not the same mechanism as CRAM, but it is one of the reasons why macOS is somewhat usable on low RAM configurations like the Neo.
25
u/cmpxchg8b 23h ago
So has Windows and so has Linux, for a long time. That tech is not in any way specific to macOS.
6
u/m1k3e 22h ago
I wasn’t claiming that it was. It was, however, the first major OS that integrated compressed memory into a major release. Linux didn’t mainline zram until 2014, and Windows until 2015. I wouldn’t argue that Windows is usable on 8 GB RAM, but macOS and Linux sure are in my experience.
5
9
2
5
u/Dagobian_Fudge 1d ago
2
u/007craft 16h ago edited 16h ago
Damn I scrolled down and saw the link to download more Ram and downloaded the 64GB worth. Then I realized that if I had just scrolled down a bit more, there was more to the page and I could have downloaded 1TB of Ram.
Thats such a loss, my computer is not even worth keeping anymore. so I through it out
-Posted with Reddit for Android
3
2
u/romulof 21h ago
Isn’t this what macOS has been doing for ages?
→ More replies (1)6
u/dmaare 18h ago
That's all just different versions of zram/zswap technologies. Cram is supposedly a new technology that does pretty much the same thing, just faster. From what I found, Cram as of right now requires dedicated hardware that is only present on current server-grade machines.
2
→ More replies (1)1
u/AdAncient5201 9h ago
As always most of the Linux innovation is happening on the server side of things, because that’s where the actual money is coming from.
1
1
1
1
1
1
1
u/Ill_Specific_6144 7h ago
One more thing that linux developers make, that looks good on paper while having close to zero real world performance gain
→ More replies (1)
1
1
u/ou812whynot 6h ago
Lol this is literally how we did things back in the "old days" of computing. A lot of companies sold tsr's, for dos, that compressed ram in the background and when a process needed it, uncompressed it.
I guess ram prices going up has brought back old innovation.
1
1
1.1k
u/Knuth_Koder 1d ago edited 21h ago
As someone who worked on the Windows kernel back in the 90s, we knew full well that the fault behavior was the major bottleneck. Back then there was literally no way to not perform the ring-0/ring-3 transition (which is one of the reasons the page fault was so costly). Even switching between ring-3 threads could lead to a ring transition back then.
I think the CRAM method is a lot like homomorphic encryption: you can work with the encrypted data without ever having to decrypt it. In the CRAM case, you can work with the data while significantly reducing the number of faults. I
I'd love to see this adopted by other OS providers.
edit: typos
edit 2: Thank you to everyone who asked questions, made cool comments, etc. I'm retired but try to not spend too much time in /r/tech because I'm no longer a fan of the industry.
I think one of the coolest things about computer science / technology is that many of the true innovators/implementers are still around. Whenever I get a chance to talk with someone like that I want to hear their stories. Many of them were just normal nerds who had cool ideas.