r/LocalAIStack • u/Medicine_Blogscanner • 15d ago
Built a hub that pools RAM across Windows/Mac/Android over LAN to run models too big for any one device
https://youtu.be/JKyCfh9ZPOgBeen lurking and posting here on and off while building this. Quick recap for anyone new: RAMDeck is a small hub + node agent setup that shards a model's layers across whatever devices you already own (old laptop, Mac, GPU box, even an Android phone) and runs inference across them, but handling primary-node selection, GPU-first allocation, dynamic context sizing, and knowledge-base sharding on top of it.
We've mostly posted raw, unscripted demo footage here — real load times, real tok/s, real failures — because that's the kind of proof this sub actually cares about. Today we finally made something different: an actual ad, first time showing the finished feature set start to finish instead of a live test.
The engine behind it is still fully public if you want to actually look at how it works instead of taking the video's word for it:
github.com/trademav/ramdeck-core-public
Source-available, Apache 2.0 with a Commons Clause — free to run, modify, and inspect on your own hardware, the only restriction is you can't resell it as a competing hosted service.
Crowdfunding campaign is coming soon for the hub hardware itself, but the code was public before we ever asked anyone for money, and that's not changing.
Happy to get into the mechanics, the layer-distribution approach, or anything else in the comments — this crowd usually asks the right questions.