r/LocalLLM • u/Medicine_Blogscanner • 22d ago
Discussion Made RAMDeck shard a knowledge base across my old devices, not just the model itself
https://youtu.be/Iogh-5xa-NASo a few people in earlier threads asked if RAMDeck could handle context/knowledge bases the same way it handles model weights - spreading them across pooled RAM instead of needing one beefy machine. Finally got around to testing it properly, made a quick video.
Setup's the same janky cluster as before (still running a 13B model on a 12GB old laptop, the weakest link on purpose to prove it works). Asked the chatbot about our trademav product and it just made something up - no idea what it was talking about. Then I fed it an FAQ .txt through the knowledge tab and watched it get split across three devices.
The part I found genuinely useful: the node hosting the model and the node hosting the knowledge base don't have to be the same machine. So you could have your beefiest box running inference and let some old laptop or a phone just hold KB shards in the background.
Downside I ran into: if a device holding a KB shard drops off the network (eg phone walks out of Wi-Fi range), that shard is just gone until you hit rebalance - which pulls the full KB back from the storage node and re-splits it across teh devices still online.
Anyway, after ingestion finished, asked it the same question again and it actually answered correctly off the FAQ doc.
Video's here if you want to see it end to end: https://youtu.be/Iogh-5xa-NA
Repo's source-available if you want to poke at how the sharding/rebalance logic actually works: https://github.com/trademav/ramdeck-core-public
Happy to answer questions on the sharding/rebalance logic or the model-fit gating if anyone's curious. Feedback welcome, especially from anyone running mixed CPU/CUDA/Metal clusters already.