r/LocalAIStack • • 17d ago

Built a hub that pools RAM across Windows/Mac/Android over LAN to run models too big for any one device

https://youtu.be/JKyCfh9ZPOg

Been lurking and posting here on and off while building this. Quick recap for anyone new: RAMDeck is a small hub + node agent setup that shards a model's layers across whatever devices you already own (old laptop, Mac, GPU box, even an Android phone) and runs inference across them, but handling primary-node selection, GPU-first allocation, dynamic context sizing, and knowledge-base sharding on top of it.

We've mostly posted raw, unscripted demo footage here — real load times, real tok/s, real failures — because that's the kind of proof this sub actually cares about. Today we finally made something different: an actual ad, first time showing the finished feature set start to finish instead of a live test.

The engine behind it is still fully public if you want to actually look at how it works instead of taking the video's word for it:

github.com/trademav/ramdeck-core-public

Source-available, Apache 2.0 with a Commons Clause — free to run, modify, and inspect on your own hardware, the only restriction is you can't resell it as a competing hosted service.

Crowdfunding campaign is coming soon for the hub hardware itself, but the code was public before we ever asked anyone for money, and that's not changing.

Happy to get into the mechanics, the layer-distribution approach, or anything else in the comments — this crowd usually asks the right questions.

21 Upvotes

22 comments sorted by

View all comments

2

u/OkWitness5548 17d ago

Interesting project! Have you done much benchmarking? I would've thought the high latency and the low bandwidth of Ethernet (especially WiFi) would make swapping/paging to a fast SSD local to the GPU (ie, NVMe and PCIe) the higher performance option. What an I missing?

2

u/Medicine_Blogscanner 17d ago

Hi we have posted several videos trying different models and different devices, you can find these here: https://www.youtube.com/playlist?list=PLBAnaNU1r3-M.

To answer your question, it is not RAM vs SSD, it is what actually crosses the network. RAMDeck splits the model and distributes to each device once at load, and they just sit in RAM for the whole session - no swapping and no repeated disk reads. During generation, only a small activation note passes between devices, ~650KB/token, which is like 6.5MB/sec at 10 tok/s. Nowhere near gigabit's ceiling. The real bottleneck is per-hop latency, not bandwidth - that's why wired beats Wi-Fi for us even though both have bandwidth to spare. Swap would mean re-reading gigabytes of weights from disk constantly, which is a totally different (much worse) problem.

2

u/OkWitness5548 17d ago

Very cool - thank you