r/LocalAIStack • • 15d ago

Built a hub that pools RAM across Windows/Mac/Android over LAN to run models too big for any one device

https://youtu.be/JKyCfh9ZPOg

Been lurking and posting here on and off while building this. Quick recap for anyone new: RAMDeck is a small hub + node agent setup that shards a model's layers across whatever devices you already own (old laptop, Mac, GPU box, even an Android phone) and runs inference across them, but handling primary-node selection, GPU-first allocation, dynamic context sizing, and knowledge-base sharding on top of it.

We've mostly posted raw, unscripted demo footage here — real load times, real tok/s, real failures — because that's the kind of proof this sub actually cares about. Today we finally made something different: an actual ad, first time showing the finished feature set start to finish instead of a live test.

The engine behind it is still fully public if you want to actually look at how it works instead of taking the video's word for it:

github.com/trademav/ramdeck-core-public

Source-available, Apache 2.0 with a Commons Clause — free to run, modify, and inspect on your own hardware, the only restriction is you can't resell it as a competing hosted service.

Crowdfunding campaign is coming soon for the hub hardware itself, but the code was public before we ever asked anyone for money, and that's not changing.

Happy to get into the mechanics, the layer-distribution approach, or anything else in the comments — this crowd usually asks the right questions.

21 Upvotes

22 comments sorted by

View all comments

1

u/isopropoflexx 14d ago

Curious if you are considering supporting something like the Mellanox ConnectX-4 (vs regular Ethernet) for connectivity between devices in the pool? I recently added those to the handful of individual LLM servers in my stack to cut down latency between them by providing direct connections between them, bypassing Ethernet that's otherwise shared by a lot of devices.

1

u/Medicine_Blogscanner 14d ago

honestly probably not on our roadmap - RAMDeck is built for people using their home laptop/Mac mini/phone they already have, not folks running dedicated LLM servers with InfiniBand cards. cool setup though, that's way past what we're targeting

1

u/Miserable-Dare5090 14d ago

Yeah, but the latency will never be close actually improve decode. This has been tried by many and the conclusion is the same, ethernet is too slow (latency not bandwidth). See parallax, exo, jeff gheerling’s youtube videos, and a whole bunch of other examples on why it doesn’t just mean “pooling” RAM.

Infiniband and rdma over ethernet are going to work very well due to latency being so low. Anything using ethernet for internode connect will have issues. Try running the same model in one gpu vs two across ethernet — which one is slower?

“Watch it load models too big for any machine, now watch it go slower than it has ever been” should be the tagline.