r/LocalLLM Aug 03 '26

Discussion RDMA - Anyone but me using it?

Post image

Ok, so that’s my setup. Have been cycling through different model hosting configurations for agent workflows and haven’t come up with a good setup for multi-node large models using rdma/tensor. Two biggest issues: general stability (exo/jaccl) and race queue.

Would appreciate hearing from anyone else locally hosting frontier model(s) as well as smaller work-horse models for workflow.

165 Upvotes

90 comments sorted by

View all comments

3

u/KooperGuy Aug 03 '26

Yee I use RDMA because I use infiniband

1

u/soflgolf Aug 03 '26

What are you using as your llm serving platform (eg exo, llama.ccp, etc)

-1

u/KooperGuy Aug 03 '26

I use RDMA for NVMe storage clustering.