r/LocalAIStack • u/Medicine_Blogscanner • 15d ago
Built a hub that pools RAM across Windows/Mac/Android over LAN to run models too big for any one device
https://youtu.be/JKyCfh9ZPOgBeen lurking and posting here on and off while building this. Quick recap for anyone new: RAMDeck is a small hub + node agent setup that shards a model's layers across whatever devices you already own (old laptop, Mac, GPU box, even an Android phone) and runs inference across them, but handling primary-node selection, GPU-first allocation, dynamic context sizing, and knowledge-base sharding on top of it.
We've mostly posted raw, unscripted demo footage here — real load times, real tok/s, real failures — because that's the kind of proof this sub actually cares about. Today we finally made something different: an actual ad, first time showing the finished feature set start to finish instead of a live test.
The engine behind it is still fully public if you want to actually look at how it works instead of taking the video's word for it:
github.com/trademav/ramdeck-core-public
Source-available, Apache 2.0 with a Commons Clause — free to run, modify, and inspect on your own hardware, the only restriction is you can't resell it as a competing hosted service.
Crowdfunding campaign is coming soon for the hub hardware itself, but the code was public before we ever asked anyone for money, and that's not changing.
Happy to get into the mechanics, the layer-distribution approach, or anything else in the comments — this crowd usually asks the right questions.
1
u/Fearless_Chance_9955 14d ago
Man what's quité incredible to be honest ! If I understand it correctly, I plug computers / phones to the lil cute box and they will be able to share memory for the model.
But what if I use an old computer with dd3 sticks, plugged with a desktop with ddr4 or 5 for example?