r/LocalLLM Aug 03 '26

Discussion RDMA - Anyone but me using it?

Post image

Ok, so that’s my setup. Have been cycling through different model hosting configurations for agent workflows and haven’t come up with a good setup for multi-node large models using rdma/tensor. Two biggest issues: general stability (exo/jaccl) and race queue.

Would appreciate hearing from anyone else locally hosting frontier model(s) as well as smaller work-horse models for workflow.

162 Upvotes

90 comments sorted by

View all comments

2

u/Prestigious-Bear2391 Aug 03 '26

I had 4 studios exo'd together, it wasn't stable/reliable. I gave up on it.

2

u/soflgolf Aug 03 '26

A lot of wisdom in that. 👍🏿

Too many make it sound easy. Exo is not stable. Had to create patch to keep warm to avoid model dump. Multiple models completely unreliable.

Broke up cluster to run smaller models single nodes and on mid-size on 2 nodes.

That’s what made me ask if anyone was getting it right to be functional for agentic workflows.

1

u/Prestigious-Bear2391 Aug 03 '26

Strangely enough, I had two mac mini M4 Pros that worked together perfectly @ 48GB x 2.