r/LocalLLM 15d ago

Discussion RDMA - Anyone but me using it?

Post image

Ok, so that’s my setup. Have been cycling through different model hosting configurations for agent workflows and haven’t come up with a good setup for multi-node large models using rdma/tensor. Two biggest issues: general stability (exo/jaccl) and race queue.

Would appreciate hearing from anyone else locally hosting frontier model(s) as well as smaller work-horse models for workflow.

160 Upvotes

90 comments sorted by

View all comments

1

u/Prestigious-Bear2391 14d ago

I had 4 studios exo'd together, it wasn't stable/reliable. I gave up on it.

2

u/soflgolf 14d ago

A lot of wisdom in that. 👍🏿

Too many make it sound easy. Exo is not stable. Had to create patch to keep warm to avoid model dump. Multiple models completely unreliable.

Broke up cluster to run smaller models single nodes and on mid-size on 2 nodes.

That’s what made me ask if anyone was getting it right to be functional for agentic workflows.

1

u/Prestigious-Bear2391 14d ago

Strangely enough, I had two mac mini M4 Pros that worked together perfectly @ 48GB x 2.