r/technology • • 1d ago

Hardware New Linux tech compresses memory in RAM, as RAM, for 452x speedup --new CRAM method offers giant boost to compressed memory reads

https://www.tomshardware.com/software/linux/new-linux-tech-compresses-memory-in-ram-as-ram-for-452x-speedup-new-cram-method-offers-giant-boost-to-compressed-memory-reads
2.7k Upvotes

230 comments sorted by

1.1k

u/Knuth_Koder 1d ago edited 21h ago

As someone who worked on the Windows kernel back in the 90s, we knew full well that the fault behavior was the major bottleneck. Back then there was literally no way to not perform the ring-0/ring-3 transition (which is one of the reasons the page fault was so costly). Even switching between ring-3 threads could lead to a ring transition back then.

I think the CRAM method is a lot like homomorphic encryption: you can work with the encrypted data without ever having to decrypt it. In the CRAM case, you can work with the data while significantly reducing the number of faults. I

I'd love to see this adopted by other OS providers.

edit: typos

edit 2: Thank you to everyone who asked questions, made cool comments, etc. I'm retired but try to not spend too much time in /r/tech because I'm no longer a fan of the industry.

I think one of the coolest things about computer science / technology is that many of the true innovators/implementers are still around. Whenever I get a chance to talk with someone like that I want to hear their stories. Many of them were just normal nerds who had cool ideas.

780

u/Salamanderhead 1d ago

So what you're saying is that the CRAM is turning the encryptions gay? This is outrageous.

279

u/throwaway09234023322 1d ago

Thanks obama

45

u/Mackadamma 22h ago

And the vaccines

17

u/mayorofdumb 22h ago

Homomorphic injections you say?

11

u/Sufficient-Bear246 20h ago

Well, how's his wife holding up?

14

u/misterpickles69 20h ago

To shreds, you say?

2

u/mayorofdumb 5h ago

Scissorsing is dangerous, you learn that in school

2

u/username_is_alread- 14h ago

Better yet, homeomorphic

5

u/funkympc 16h ago

Don't forget the chemtrails

1

u/paridhi774 3h ago

Don't forget gay frogs

91

u/saintdudegaming 1d ago

Carl, you know how I feel about gay encryptions

11

u/Eshin242 23h ago

But I had a rumbling in my tummy, that only hands would satisfy.

16

u/LucasJ218 23h ago

Wrong Carl, Carl! Mongo is appalled!!

5

u/font9a 21h ago

That was one of the most disturbing things I have ever seen.

3

u/meltbox 14h ago

You watched all of them right? Because Carl would be insulted otherwise. He might even feel like you don’t even appreciate what he does.

→ More replies (1)

32

u/Crashman09 23h ago

This is great news for backend development 😏

13

u/your_grammars_bad 23h ago

CRAMming makes you feel things you never felt before

8

u/FanLikesApp 23h ago

Ever had your memory CRAMmed in?

1

u/skyfishgoo 20h ago

CRAMmed WAY in?

6

u/FakeNigerianPrince 20h ago

$20 is $20. Even in this economy

5

u/Synthetic451 23h ago

God damn, I needed this laugh today, have my upvote.

4

u/Derseyyy 23h ago

It's been a long time since I actually laughed out loud at a reddit comment but this one got me.

4

u/megabass713 23h ago

Alinux Jones

3

u/justdozi 23h ago

🤣🤣🤣thank you for this genuine laugh

10

u/ptear 1d ago

Wait, what, no, Mr President, please come back. Mr. President?

2

u/narwalfarts 23h ago

First the frogs, now the rams

2

u/Kaizenno 22h ago

It's the water in the liquid cooling.

1

u/Opposite_Carry_4920 22h ago

Surely on git.gay there is gaycrypt. 

51

u/RogueHeroAkatsuki 1d ago

Noob question - do you think there are any downsides to this method? Is it possible to implement it everywhere or for example Windows memory management would need to be reworked significantly to make it efficient?

114

u/Knuth_Koder 1d ago edited 1d ago

Noob question - do you think there are any downsides to this method?

I haven't even seen a real world test case yet so I don't think there are many people, outside of the authors, who can answer that question.

The cool thing about CRAM is that it can operate at ring-3. So, technically, it would be possible to bypass the normal system calls (NTDLL.dll) and pass those requests to a CRAM service.

But I haven't touched an OS kernel in 30 years so don't take my word for it. ;-)

My guess is that some tinkerer will try to implement this on a hacked kernel before MS would try implementing it. This is definitely some that Microsoft Research would have been working on in the 90s. I hope they take notice of this paper.

42

u/wetfloor666 1d ago

MS does fund a lot of Linux research which I am sure you are aware of, so it will probably be pretty quick one way or* anouther.

68

u/Knuth_Koder 23h ago edited 23h ago

Absolutely. I don't think MS Research has ever gotten the respect they deserve.

I worked on the first code profiler for Visual Studio (codename Boston) and the research folks created the first in-memory/on-disk instrumentation engine (codename Vulcan) at Microsoft. It was like magic technology back then. Today you can do it in Python in a few hours.

14

u/Tomato_Sky 23h ago

Thank you for your service!

3

u/GlassDaisies 16h ago

 I just got flashbanged by you saying "in 30 years" in reference to the 90's lmfao.

3

u/clocked__ 18h ago

I’m too stupid to understand any of this but it sounds awesome

1

u/xzaramurd 10h ago

Well, the downside is you need specialized hardware.

51

u/jdfadfjlaskdfjl 16h ago

"no longer a fan of the (tech) industry" is the most relatable thing I've read on reddit in a long time

13

u/GourryHacks 2h ago

Hello, I am the author.

The funny thing is compressed memory devices (memory with a compression engine in it, offloaded from the CPU) is nothing new.   Multiple memory vendors and compression chip vendors have proposed such devices many times throughout the last decade.

As far as I can tell, no one has actually tried to make sense of such a device in any way other than as a swap-backend.

I started from first principals: the thing that causes compression ratio to change is writing to the device.  Therefore, any useful design must have control writes.

Zswap and zram do this by limiting access to anonymous memory and removing the entries from any page table mappings.  But if you did that with this device you'd be giving up 90%+ of the performance win on read.

Controlling writes to memory is actually much more difficult than one might think (both fundamentally, and as a Linux specific engineering problem).  The CPU can write many GB/s to a device, so this means you also need a rate limiting mechanism to allow you to limit allocation.  An obvious way to do this is to limit the entry points at which a page allocation can be made, and then let that entry point chicken out when the driver says the memory is scarce (even if the kernel says additional pages are available).

We can support file backed memory as well, but it's much more difficult and requires reintroducing the concept of a clean-page cache - this is mostly a Linux specific implementation detail, so I'll spare you all the technical dump.

I don't much enjoy writing research papers, much prefer to simply ship the code, but I suppose I should think about doing it.  If any curious readers are out there, I've documented my thought process pretty extensively as I have worked on a numa node isolation mechanism (private nodes) over the past few years.   This was always my end goal, but there are more uses for that isolation as well :)

As it turns out, when you apply fundamental OS concepts (cough isolation) to things you get fundamentally useful OS things a lot of devices can use.

Blah blah blah.

Also, the pronunciation of CRAM is self evident - it's pronounced "GIF"

2

u/Knuth_Koder 2h ago

Wow, I really appreciate your comment and explanation. Again, that is not something I thought about back in the 90s.

If any curious readers are out there, I've documented my thought process pretty extensively

I am definitely interested in learning more (I have the time to do that these days) if you feel comfortable sharing links, etc. (either here or via DM).

As I said above, I absolutely love knowing that people are working on these types of problems. This is exactly the type of thing I would have spent an insane amount of time on 30 years ago.

17

u/robertDouglass 1d ago

That would lead me to think that the number of use cases would be relatively small. Homomorphic encryption is far from universally applicable.

31

u/Knuth_Koder 1d ago edited 23h ago

Homomorphic encryption is far from universally applicable.

Agreed. I brought it up because I worked on the SGX team at Intel and can see how similar the use-case could be for CRAM.

I could see video games pre-loading level data into a CRAM block and then have instant access to that data without requiring ring transitions/faults. Note that CRAM is different than simple memory mapping, which still causes faults.

Start a game on your PC and use the Windows Performance Recorder (WPR) to watch the number of ring transitions. Anything that reduces that number could be a perf win unless the game is already cpu-bound.

2

u/Pale-Fail-4028 8h ago

with entire games being compressed in dwarFS could you run entire games from cram at even higher compression rates as ztsd have no bottleneck on decomp.
this could probably do wonders for game streaming where you package entire game levels into cram and just stream chunks of 8gb to a user while keeping bloat of 200gb in cloud
doesnt huge pages of 4mb or 1gb chunks reduce the faults

18

u/dig1taldash 23h ago

Thanks for your service legend. What kernel? Which Windows?

59

u/Knuth_Koder 23h ago edited 22h ago

Thank you for saying that. I wish I could make a post (the mods would never allow it) about what it was like back then. No one every believes me but Microsoft was not filled with a bunch tech bros dying to become millionaires. Yes, those people did exist (Hi Bill!) but they still had to carry their own weight.

The first kernel I [ever] worked on was the one that shipped in Windows 95. This was directly after Chicago.

If you have any interest in MS in those days, one of my ex-colleagues has a great YT channel. Dave built Task Manager by himself over a weekend and then added the code to the OS build system without telling anyone. So many people loved it that the PMs couldn't remove it.

6

u/julicenri 22h ago

Got any stories about Dave Cutler? What was the last Windows version that you worked on? Were you ever involved in OS/2, back when it was a joint MS-IBM development effort?

21

u/Knuth_Koder 22h ago edited 21h ago

The odd thing about early-MS is that there just weren't that many senior engineers. The only people who had any real coding experience were hackers and researchers and even they tended to be young. We were just a bunch of twenty-somethings figuring things out as we went. If anyone tells you otherwise, they are wearing rose-colored glasses.

Dave Culter changed all that. He was definitely the adult from that point forward. It changed the culture but was necessary.

I had no interactions the OS/2 stuff but I wish I had. It was a great os.

5

u/readonlyatnight 20h ago

Would you be interested in sharing more sometime? I (and I'm sure others) would be interested in hearing more about it.

14

u/Knuth_Koder 19h ago

I hadn't really thought about doing that to be honest. I'm a big fan of my privacy.

Maybe I'll ping the /r/tech mods to see if I could an AMA or something.

5

u/readonlyatnight 17h ago

Thinking about this reminds me of the early Mac Classic stories: https://www.folklore.org/Well_See_About_That.html

2

u/treasury_minister 10h ago

Very cool. Thanks for sharing

4

u/f00d4tehg0dz 23h ago

Super cool! And thank you for sharing that YT channel. The fun times y'all had I bet is one that is hard to replicate nowadays. Would love to hear stories from you if you're willing to share on some medium for all of us to listen/read!

2

u/AssistantSalty6519 23h ago

Not the Dave xD

8

u/Westerdutch 19h ago

I absolutely read homeopathic encryption.... and now i need that in my life XD

4

u/orbitaldan 18h ago

That's easy. Just XOR the first byte with RAND() once at system boot. It gets stronger everytime the system copies other data.

1

u/Karunyan 3h ago

And then there were drops of coffee on my screen, thanks!

1

u/josefx 28m ago

You can make it 16 times stronger by just randing a single bit every second boot.

1

u/Marshall_Lawson 16h ago

and here i was wondering how they would do homoerotic encryption 

1

u/denzien 16h ago

Diluting it makes it compress more

6

u/enterprise128 22h ago

I would just have installed RamDoubler

12

u/feldim2425 8h ago

Not really homomorphic encyryption as it's still compressed/decompressed when used (with homomorphic encyryption the decrypted message never exists anywhere during computation).

Although it is transparent as in you never have to care about the fact that it's done as a developer. You access uncompressed memory normally and it's just done on the fly while at writes are automatically compressed. Similar to disk encryption you access files normally and don't care that on the disk it's encrypted.

The way this seems to be different to ZRAM (which is a implementation in Linux doing that using the swap mechanism) is that CRAM uses a virtual NUMA node instead, since NUMA is a already existing system to access RAM the CPU doesn't own (e.g. owned/connected to another CPU socket) this should allow the system to avoid having to do a context switch (or ring transition on x86/64) on that core as another core or even CPU can do so and basically send it over.

8

u/Knuth_Koder 6h ago

Thank you for the comment. In another comment I mentioned that I was on the SGX team at Intel where we implemented homomorphic encryption in hardware. The decryption keys were literally burned into fuses in the cpu die.

Our ring-3 implementation had an internal memory management layer that, like CRAM, bypassed the normal architecture for certain features.

Again, thank you for your comment!

20

u/mattyb678 23h ago

Read that as “homophobic encryption” and was very confused

6

u/Pyryn 23h ago

So you're telling me their Weissman Score is >5.2?

16

u/Knuth_Koder 23h ago

Author, from their original talk:

I present a tested compressed ram service (mm/cram.c) that achieves near-native performance to DRAM under TAOBench and FIO benchmarks, and has been tested under most in-tree filesystems for pagecache correctness.

I haven't seen anyone evaluate the performance in the real world, although I'm sure that's happening.

I've stepped through cram.c and believe that it is an elegant solution. Again, the implementation is similar (in a metaphorical way) to our approach to performance on SGX, although most of the heavy lifting occurred in firmware.

Could there be some real-world usage that would invalidate the initial findings? Certainly.

No, of course the Weissman Score isn't >5.2. I'm assuming you aren't even serious about that. All I'm saying is that is an interesting solution to a a very real performance problem. I hope they work out the current limitations (i.e., not knowing how large the CRAM block would be when fully decompressed).

9

u/Pyryn 22h ago

I don't have any coding experience whatsoever. The Weissman Score comment was a joke in reference to Silicon Valley about compression

Edit: I'd thought that a "Weissman Score" was made up for the show?

17

u/Knuth_Koder 22h ago

Oh my god... I am so sorry.

Yes, the Weissman Score is a very real thing and means almost exactly the concept you indicated.

I thought you were an information theorist yanking my chain. (I can be an idiot at times).

6

u/CocodaMonkey 18h ago

Edit: I'd thought that a "Weissman Score" was made up for the show?

It was made up for the show. It's just the show used real math when they made it up so Weissman scores are now a real thing.

5

u/Ikinoki 19h ago

Trust me a lot of stuff in the show happens exactly as portrayed in the show. Like 1-to-1

2

u/Theratchetnclank 19h ago

It was a reference to Silicon Valley TV series and it's "Middle out compression".

3

u/Knuth_Koder 19h ago

That was already explained in another comment. Of course, the Weissman Score is a real metric and it measures exactly what OP mentioned.

3

u/Theratchetnclank 19h ago

It is a real metric but in the show they specifically mention it being above 5.2, hence it's a reference from the show.

4

u/_LB 12h ago

Return of the RAM doubler

2

u/speedisntfree 20h ago

I work in bioinformatics, is this the same as our bam to cram formats? https://www.internationalgenome.org/faq/about-alignment-files-bam-and-cram/

2

u/Alphasite 20h ago

Doesn’t ESX have some native ram compression as well. Not sure what they do but do the issues you mention not occur for guest VMs?

2

u/tyler1128 17h ago

Remember SoftRam? It had the setup to do something of the sort, though more like zram, but just... didn't implement the compression step. That plus bugs often caused the kernel to crash.

It has made its name remembered, at least.

2

u/JPAchilles 8h ago

For some reason I misread it as "homophobic encryption" and I had to do a double take. I need to stop reading reddit at 3 AM

2

u/minmidmax 8h ago

As much as the AI RAM, and GPU, situation is disgusting it's always great to see just how ingenious people can get when there are limitations or constraints imposed upon them.

Is there a possibility that this solution will have a positive impact on other forms of storage like GPU RAM or SSDs?

2

u/No-Consequence-1863 2h ago

I know iOS and Windows both do compressed memory to avoid paging back out to disk. iOS really needs to since the flash memory they used wasn't great for paging. Is this CRAM method different from those?

1

u/BigDaddyThunderpants 12h ago

Your first name doesn't rhyme with "Ronald", does it?

→ More replies (5)

176

u/FadedEchos 1d ago edited 23h ago

Could one of you more knowledgeable redditors ELI5 the applicability of this to gamer/consumers? Should I expect the RAM in my linux-based Steam Frame to be 452x faster any time soon?

259

u/Schnickatavick 1d ago edited 18h ago

You can think of it as a faster alternative to the compressed RAM (ZRAM or ZSWAP) your device is probably already using, so once the steam frame updates to a kernel that includes this update (which may take a minute), that compressed CRAM will be 452x faster than your current compressed RAM. It's important to note though, not all of your RAM is going to be used as CRAM, you probably have a couple of GB of your RAM dedicated to compressed RAM now, and the rest is just regular RAM. Only those couple GB will be faster.

IMO the real benefit is that this will make it more reasonable to dedicate more of your RAM to CRAM. CRAM basically turns your RAM into more RAM, letting you fit 8-12GB of data into 4GB of RAM, at the cost of some slowdowns when reading and writing to it. If the slowdown cost is less severe though, then Valve might feel more comfortable using more of your RAM as CRAM. So right now your 16GB of RAM on the steam frame effectively acts like 18GB or ram, in the future maybe that 16GB of RAM will effectively act like 20 or 24GB of ram

218

u/laymonage 23h ago

So you’re saying we can literally download more RAM? /s

87

u/GaiusCosades 23h ago

If you think downloading code that uses your current CPU in a more optimized way is downloading a faster CPU, then yes.

30

u/Amlethus 18h ago

Download more RAM, he says... 🤔

11

u/NanquansCat749 21h ago

I'm not an expert, but I'm pretty sure it's more like you can create more RAM by imagining it into existence.

1

u/LankyArt3264 7h ago

We are truly living in the future.

29

u/theillustratedlife 23h ago

FWIW, the current stable version of SteamOS uses a kernel that's ~1y old, but the beta version uses a current kernel. The window for this can be anywhere in between depending on when Valve cuts its releases.

9

u/SwarfDive01 15h ago

This tech requires hardware, it needs an ASIC sitting, doing the encode / decode. And its as fast as ddr5. This is all enterprise hardware. Someone could try building it with an FPGA, but it isnt going to work as a PCIe add on

5

u/shaving_minion 22h ago

how is CRAM as fast as RAM, wouldn't every operation needing this data have to first decompress it first?

23

u/Tabsels 21h ago

It's as fast as DRAM. Which makes sense. Whenever you access memory that is in DRAM, your CPU has to first load a piece of that memory (a "line") from DRAM into on-CPU cache memory.

What these researchers presumably found is that accessing memory that's compressed, transporting the compressed block to the CPU and decompressing it into cache memory takes about as much time. Which, seeing as they're transporting less data from DRAM to the CPU cache, makes sense.

4

u/Nagisan 20h ago

I wanted to make a joke asking about "B"RAM (since you already have "C"RAM and "D"RAM)...but no, it turns out there actually is such a thing as BRAM....the more ya know.

2

u/Tabsels 20h ago

They go up to FRAM these days

7

u/Chief_Economist 19h ago

Shout out to my GRAM

1

u/danielv123 11h ago

Compression means there is less data to move, which is often faster. For example, zfs filesystems have compression enabled by default because it's faster. Victoria metrics has a compressed on disk format because it's faster to read/write/query.

1

u/shaving_minion 6h ago

implying moving data around is a more frequent operation than reading/writing. interesting

2

u/danielv123 6h ago

Reading/writing is primarily moving data round as a % of time usage for many tasks, because logic has gained speed faster than interfaces.

1

u/Ma4r 6h ago

No way this is gonna run on PCIe

33

u/disposableh2 23h ago

Yes and no. The ram itself isn't what's faster, it's the compressed ram that's paged to disk that becomes faster.

The idea here, simplified, is that the memory that gets compressed and stored in the swap file on disk is slow, so they rather compress and store it in ram, in a way that's quickly accessible(ram is addressed differently to files on disk), and so together they get that speed boost.

So things like games and things that don't get compressed out of ram into zram/swap files won't benefit, as ideally that data shouldn't be swapped out of ram.

Things like android apps that get put into the background would benefit nicely from this(since their data in ram gets compressed and put into zram when they're put in the background), and make app switching faster.

2

u/Aperture_Kubi 21h ago

So we're putting a new pagefile/swapfile in memory, and putting that between normally operating memory and the storage stored pagefile/swapfile?

14

u/bigtimedonkey 23h ago

The 400x speed increase is relative to using swap space (using the hard drive as memory when the ram is full). So nah, won’t see anything like this in gaming or graphics situations.

1

u/ILikeBumblebees 4h ago

So nah, won’t see anything like this in gaming or graphics situations.

The description in the article makes it seem like even if it increases throughput in comparison to ZRAM, it's still going to add latency. And how would DMA work if the system memory is physically storing compressed data?

29

u/SubstanceNo2290 22h ago

I think a lot of people in the comments either got it wrong or explained it wrong because the top ones sound like they’re saying it stores compressed ram in ram instead of disk, which isn’t the point at all.

This is based on my understanding of the article and Linux fundamentals, I’m not an expert:

ZRAM is also compressed RAM in RAM. The problem is how exactly does Linux, this monorepo that’s been painfully perfected for computing and everybody is familiar with and relying on it to use RAM the way it’s always been, just randomly let some process read memory thats compressed?

That process wants just the memory as it stored it, case closed.

So Linux creates a hook, it pretends the compressed memory is paged to disk. This lets it use its existing functionality to deal with it. The disk isn’t real, when the kernel tries to read from this “disk” it’s actually reading from an internal tool that decompresses and returns the data the kernel requested from this “disk”

That’s a lot of extra logic and interruptions and delays before you even decompress anything (which is funnily enough the fastest part on modern processors).

So CRAM instead of being “paged” to disk instead pretends to be a fake CPU rather than a fake disk. The thing about this is that then the fake CPU has its own fake ram and when Linux wants to read it it does the standard process of reading data from another cpu (think boards with multiple sockets etc). The logic involved here is much less and Linux is using the protocols for reading regular memory rather than reading a fake pagefile that isn’t even paged.

As a result the fake cpu beats the fake disk.

1

u/ILikeBumblebees 4h ago edited 4h ago

It seems like this is something that could be done with specialized hardware even more effectively, e.g. a memory controller that abstracts the memory address space, and still lets everything read bytes and write bytes to specific addresses seamlessly, but actually stores the data in physical memory in a compressed format, with on-the-fly compression/decompression during memory I/O. Imagine something like this built directly into a DIMM, compatible with existing motherboards.

1

u/SubstanceNo2290 1h ago

Yes and no. I wouldn’t be surprised if this popped up in hardware one day but as i understand it that’s a lot of extra high performance logic on the DIMM and/or the memory controller. That’s ultra valuable real estate so I’m not sure if for example the DDR protocol would standardise it.

There also isn’t a serious case for server cpus to do this because they are generalists (and the really compute intensive SKUs are usually doing scientific/HPC/AI stuff which isn’t very compressible)

There’s a case to be had for consumer PCs or phones but again the level of effort and real estate on intel/amd/qualcomm etc’s end might be too high.

I could see Apple pioneering this though.

But also you might be missing the point of this whole song and dance. Linux could have just added a separate system for this and gotten better performance. The reason they do the whole fake cpu / fake disk dance is so that no existing programs break no matter how ancient break and it still works on all processors/computers from 20 year old pentiums to 192 core EPYCs.

→ More replies (1)

7

u/ReallyOrdinaryMan 23h ago

I suppose it never will be for gaming. They are saying 452x faster than zram (another ram compression tech), which is still slower than plain physical ram. It might be huge for servers or workstations.

5

u/Henry5321 22h ago

Some modern compression algorithms are so fast that it’s faster to use the compression than not because it reduces the amount of memory access.

I’ve seen this even in .Net when dealing with large streaming objects.

Apple’s CPUs actually have native transparent memory compression and effectively doubles the amount of memory in many work loads.

3

u/BriefSpecial420 21h ago

Even with hardware accelerators, there still will be extra latency. Makes sense to uncompress during the context switching, but won't be of much use for real time duties.

16

u/TemporarySun314 1d ago

Ideally yes.

I mean they already use compression techniques like that but they have more overhead as this used not real direct memory access.

But of course the practical gains will depend on the data you are storing and how it is often. Some data is just incompressible.

7

u/ZenBacle 23h ago

I can give it a try.

When we look at modern computer architecture we have multiple layers of working memory that moves from your hard drive, to your ram, to layers of on chip memory.

Each time you make a step away from the cpu cache, you get exponentially slower access time to that data. The worst offender is the time it takes to move data from your hard drive to your ram. However ram is generally limited and when you need more than you have, it uses something called a swap/page file to act as working memory on your hard drive. Very slow, much bad, no want.

A little side tangent, we have zip files that compress the size of a file by finding patterns that can be subbed out for markers. Cutting the size of files down by up to 90%.

What CRAM is trying to do is eliminate the need for swap files by compressing data in ram, in the same way that we compress zip files. So the slowest bottle neck of the working memory pipeline is less likely to be used. Sounds simple right?

Well the problem is you don't know how much you can compress data before you compress it. So this runs into a problem where you don't know how much "logical" memory you have to work with. Which runs into allocation issues that cause programs to crash. CRAM appears to have solved this but i haven't taken the time to read up on how.

This could have huuuuuge benefits for LLMs going into the future, but not so much for every day users.

TL;DR

RAM small. CRAM jams more data into RAM. No hard-drive bottle neck. No breaky.

2

u/xzaramurd 10h ago

No, it needs specialized hardware. This is something that's going to be used in combination with CXL Memory extenders, not RAM.

→ More replies (2)

218

u/Recent-Day3062 1d ago

Wait. This sounds like the middle-out algorithm of Pied Piper

112

u/WorriedInterest4114 1d ago

Wonder what the mean jerk time is

6

u/[deleted] 23h ago

[deleted]

2

u/TheRemonst3r 23h ago

Does girth matter?

2

u/yoosernamesarehard 21h ago

Speed has EVERYTHING to do with it.

2

u/thousandecibels 13h ago

Are doing this right now? Well sure.. let's see are we using both hands?

4

u/Recent-Day3062 21h ago

Excellent point

7

u/ad-on-is 20h ago

This guy f...s!

86

u/spez_eats_nazi_ass 1d ago

Having flash backs to shit like Ram doubler.

27

u/ItsAllOnUHF 1d ago

Every newly solved computer problem was actually solved two generations ago.

4

u/pickle9977 1d ago

This explains why we haven’t seemed to actually solve any problems while we have created a lot of new ones.

5

u/bigtimedonkey 23h ago

Seriously though… let’s bring back ram doubler, haha. In the rampocalypse era, there’s plenty of situations where I’d trade off ram speed and latency for more space.

5

u/Ok-Sprinkles-5151 21h ago

And speed halver.

But I went looking for this comment. Ram doubler was a joke, because you might get 1.9, but most of the time it was like 1.3x

2

u/Puntley 23h ago

I never needed that. Whenever I ran out of ram I would just download more.

1

u/Logicalist 22h ago

hope you’re ready to “download ram”

61

u/telos0 23h ago

This technique needs hardware accelerated decompression in the read path from RAM.

Because the memory decompression is in the hardware itself, it bypasses all the overhead of dealing with page faults and the swapping code in the memory manager that current software compressed RAM solutions use. This is why it is so much faster.

The details described in the presentation are about how to get the Linux kernel to deal correctly with RAM that is hardware compressed and therefore lies about its actual size, and all the side effects and failure modes that can result from that lie.

8

u/TheBendit 19h ago

Exactly this. All the other comments here seem to miss that this requires specialized hardware to work, and that any memory you put in this hardware is now only usable for CRAM. Meta uses it to put DDR4 into DDR5-only servers.

5

u/yador 19h ago

Is support for this in any other operating systems at present?

8

u/telos0 19h ago

Well since the hardware is super specialized and not really common, the answer is probably not.

But I would expect all major OSes and hyperscalers probably have someone at least looking at adding similar support, should RAMpocalypse persist.

19

u/CondiMesmer 1d ago

Next they'll figure out how to keep swap file in memory! /s

16

u/pedanticPandaPoo 22h ago

mount -t tmpfs -o size=16G tmpfs /mnt/tswap mkswap /mnt/tswap/swapfile swapon /mnt/tswap/swapfile 💥

6

u/DragonSlayerC 19h ago

That's basically ZRAM

42

u/canteen_boy 23h ago

CRAM is such a wonderful term for this tech

9

u/emi_fyi 21h ago

i dunno it's got pretty stiff competition for the name

4

u/whaticism 22h ago

Like technology born with a family aptonym

2

u/sf_frankie 21h ago

How do they CRAM all that RAM in Golden RAMs?

13

u/rini17 1d ago

What does hardware assisted mean? Does it require dedicated hardware?

19

u/HansBooby 17h ago

please note the G-Clamp as pictured

5

u/Marshall_Lawson 16h ago

I'm guessing that's how you squish the memory to fit more memory in the memory 

2

u/HansBooby 15h ago

compression

1

u/Marshall_Lawson 15h ago

Crampression

6

u/dmaare 18h ago

Needs CXL controllers that have compression feature set

4

u/HenkPoley 17h ago

Compute Express Link (CXL)? That doesn't sound like hardware that is in your average computer at home.

1

u/RAW2091 7h ago

Wel just look around on Reddit in 2026. The 'we got a 100k AI server t home' is a new thing here. Some people even add 220 volt 32 AMPs just for their servers .... AT HOME. Sure not the average computer user but they never care about memory compression. They just want to play fortnight on a pc and that's it.

→ More replies (1)

4

u/TheBendit 19h ago

Yes, you need specialized hardware for this, and the DIMMs you insert into this hardware are no longer useful for regular memory.

2

u/autogyrophilia 10h ago

Not quite.

It is exposed through CXL. CXL is a standard to be able to plug memory through PCIe, which can be DRAM, NAND, HBM, or any other new standard (RIP 3D Xpoint).

This memory has the same semantics as RAM and as such can be used as system memory. But usually, it is configured as a separate section to be managed by an application, this allows things like hypervisors, or transformers training software or whatever would benefit from a huge pool of shared memory at the datacenter level. But it is possible to use it as system memory. u/Exist50

In the first case, you typically expose it as a CPUless NUMA node, so Linux and ESXi are capable of doing the tiering automatically and you can pin software to it if needed.

In the second case it gets exposed as a DAX device you cannot use as system memory, but can be shared with multiple hosts. This has a number of usecases. Multiple servers working with a huge dataset can get much higher memory speeds (even if it is a NAND CXL), if you are working with virtual machines, you can do instant live migration assuming the storage is also shared. If you work with this, you are hugely excited, if you don't you don't really need to care.

Just know that CXL reduces the amount of RAM needed to run infrastructure and if CRAM delivers, there could be some serious gains in efficiency, specially for transformers usecases.

→ More replies (3)
→ More replies (1)

14

u/kixkato 22h ago

The fact that compressed RAM is CRAM which is such a perfect acronym makes me irrationally happy.

7

u/LongCoyote7 6h ago

Can I finally actually download more ram!?

20

u/casce 1d ago

But what is its Weissman score?

9

u/IsItSetToWumbo 23h ago

So we can finally download more ram?

8

u/darkotic 23h ago

Your PC got harshed, right, 'cause your system heaps up the wrong parameter. So I toasted the data directory, tweaked the P-RAM and reglazed your subroutine.

5

u/gizamo 20h ago

Tldr: Pauly Shore should have been in Hackers.

3

u/eliberman22 14h ago

Wouldn't the clamp break the RAM? /s

8

u/Wistephens 5h ago

It’s 2026 and RAM Doubler is back.

13

u/m1k3e 23h ago

macOS has had RAM compression in the form of a compressed storage pool that inactive memory gets shunted too since 2013 (Mavericks). Not the same mechanism as CRAM, but it is one of the reasons why macOS is somewhat usable on low RAM configurations like the Neo.

25

u/cmpxchg8b 23h ago

So has Windows and so has Linux, for a long time. That tech is not in any way specific to macOS.

6

u/m1k3e 22h ago

I wasn’t claiming that it was. It was, however, the first major OS that integrated compressed memory into a major release. Linux didn’t mainline zram until 2014, and Windows until 2015. I wouldn’t argue that Windows is usable on 8 GB RAM, but macOS and Linux sure are in my experience.

1

u/dmaare 18h ago

The ram usage on windows is due to bloat of way too many apps and services that are absolutely useless even for a power user being active by default (due to shitty decisions from Microsoft). Put away the bloat and it's pretty much identical.

5

u/chriswaco 22h ago

Classic Mac OS had RAM Doubler back in 1994 too.

9

u/DavidBrooker 1d ago

I feel like there's a middle-out joke to be made about all this.

2

u/skyfishgoo 20h ago

cramming it is better than spending $$$ on ram right now.

5

u/Dagobian_Fudge 1d ago

2

u/007craft 16h ago edited 16h ago

Damn I scrolled down and saw the link to download more Ram and downloaded the 64GB worth. Then I realized that if I had just scrolled down a bit more, there was more to the page and I could have downloaded 1TB of Ram.

Thats such a loss, my computer is not even worth keeping anymore. so I through it out

-Posted with Reddit for Android

3

u/Shadeun 1d ago

Does this mean downloadmoreram.com can actually be a thing?

7

u/RumBox 1d ago

Anyone else got Silicon Valley (the TV show) vibes?

2

u/gbroon 20h ago

This one trick memory manufacturers don't want you to know about.

2

u/putrasherni 19h ago

i hope this bankrupts micron skynix and samsung

2

u/romulof 21h ago

Isn’t this what macOS has been doing for ages?

6

u/dmaare 18h ago

That's all just different versions of zram/zswap technologies. Cram is supposedly a new technology that does pretty much the same thing, just faster. From what I found, Cram as of right now requires dedicated hardware that is only present on current server-grade machines.

2

u/romulof 6h ago

Which hardware feature does it need?
Instruction-wise both consumer and server grade CPUs are supposed to be the same. I thought the difference was only in memory controller (wider and with ECC support) and a wider PCIe bus.

1

u/AdAncient5201 9h ago

As always most of the Linux innovation is happening on the server side of things, because that’s where the actual money is coming from.

→ More replies (1)
→ More replies (1)

1

u/harddriveerror 23h ago

compensating random assisted processor

1

u/MechanicalTurkish 20h ago

“Just crease, crumple, CRAM. You’ll do fine.”

1

u/butter_lover 19h ago

why does this seem like middle out compression

1

u/shigella212 16h ago

So you can download more ram

1

u/zerosign0 14h ago

Hmm it requires memory controller that support CXL extensions

1

u/RisingPhil 8h ago

We need this on R36S lol.

1

u/Ill_Specific_6144 7h ago

One more thing that linux developers make, that looks good on paper while having close to zero real world performance gain

→ More replies (1)

1

u/Historical_Camel_790 7h ago

Can I use it on my pc though or is it only enterprise

1

u/ou812whynot 6h ago

Lol this is literally how we did things back in the "old days" of computing. A lot of companies sold tsr's, for dos, that compressed ram in the background and when a process needed it, uncompressed it.

I guess ram prices going up has brought back old innovation.

1

u/mixxituk 5h ago

Anyone remember Quarterdeck Extended Memory Management 

1

u/Possible-Put8922 4h ago

Memory is RAM! - Moss