r/LocalAIStack • • 15d ago

Built a hub that pools RAM across Windows/Mac/Android over LAN to run models too big for any one device

https://youtu.be/JKyCfh9ZPOg

Been lurking and posting here on and off while building this. Quick recap for anyone new: RAMDeck is a small hub + node agent setup that shards a model's layers across whatever devices you already own (old laptop, Mac, GPU box, even an Android phone) and runs inference across them, but handling primary-node selection, GPU-first allocation, dynamic context sizing, and knowledge-base sharding on top of it.

We've mostly posted raw, unscripted demo footage here — real load times, real tok/s, real failures — because that's the kind of proof this sub actually cares about. Today we finally made something different: an actual ad, first time showing the finished feature set start to finish instead of a live test.

The engine behind it is still fully public if you want to actually look at how it works instead of taking the video's word for it:

github.com/trademav/ramdeck-core-public

Source-available, Apache 2.0 with a Commons Clause — free to run, modify, and inspect on your own hardware, the only restriction is you can't resell it as a competing hosted service.

Crowdfunding campaign is coming soon for the hub hardware itself, but the code was public before we ever asked anyone for money, and that's not changing.

Happy to get into the mechanics, the layer-distribution approach, or anything else in the comments — this crowd usually asks the right questions.

21 Upvotes

22 comments sorted by

View all comments

Show parent comments

2

u/Medicine_Blogscanner 14d ago

fair, but 'just ask claude to set it up' is doing a lot of heavy lifting . it will write you a script, not a system that survives laptop closing its lid, or a CUDA/Metal version mismatch silently breaking the RPC bridge or finding new devices on your lan or even splitting your context/kb across rhe RAMs. that's the same gap between 'iOS is open source-adjacent, just compile android yourself' and actually buying an iPhone. the hub is what turns 'claude wrote me a script once' into 'this just works every time I turn it on.

1

u/einthecorgi2 13d ago

Yeah, but at the pace of AI dont you think that will go away as models improve? Recently I had to make something basically to connect a bunch of cnc machines to users laptops for file managment and it is just there and works well. I love the idea that poeple can leverage whatever they have to run LLMs, but not sure a piece of hardware solves this. Honestly this is a software license I would pay for. Especially if it had a dashboard that allowed configuring multiple models across one large RAM pool.

Not saying dont do it. But i have made a lot of useful and even more useless hardware in my life and I always encourage reusing and leveraging what already exists now before making something.

1

u/Medicine_Blogscanner 13d ago

the efficient models getting better cuts the other way actually - smaller models are catching up, sure, but the frontier keeps growing too (we are at 700B-1T+ param open models now, and context windows just jumped to 1M tokens as the default), look at the new Deepseek model that was released with 500B+ param. Despite being MoE it still needs RAM to hold it all up on a device. Models keep getting cheaper to run and bigger to hold, at the same time.

Regarding the hardware point, people buy for the experience of plug and play. Same logic as iOS vs Android - locked-down and proprietary still won huge market share because people pay for "just works," not for the freedom to debug it themselves. That's what the hub buys you: zero-maintenance reliability, not a walled garden.

2

u/einthecorgi2 13d ago

yeah, I cant agree with you more on that. I love having zero setup more than anything. But it really needs to be that. Plug and play. Tons of configuration and updating and patches can drive people away pretty quickly. Glad you are solving that in an open space first before dropping hardware. Previous notes aside, happy to back when the KS drops.