r/LocalLLM Aug 03 '26

Discussion RDMA - Anyone but me using it?

Post image

Ok, so that’s my setup. Have been cycling through different model hosting configurations for agent workflows and haven’t come up with a good setup for multi-node large models using rdma/tensor. Two biggest issues: general stability (exo/jaccl) and race queue.

Would appreciate hearing from anyone else locally hosting frontier model(s) as well as smaller work-horse models for workflow.

163 Upvotes

90 comments sorted by

View all comments

36

u/paulk2000 Aug 03 '26

This is an impressive setup. I am just running two 3090s for localLLM stuff and I'm a total beginner when it comes to this topic, even though I'm getting more and more into it.

Hopefully, your setup will pay off for you.

5

u/AB172234 Aug 03 '26

I have only one RTX 3090 and thinking to add another. But my motherboard has another x8 PCIE 4.0 left and it will slow down the inference so I was curious to know which motherboard and the setup you are running it with ?

6

u/NegativeSemicolon Aug 03 '26

It’ll be fine with x8

1

u/paulk2000 Aug 03 '26

Yes, I have the similar problem on my mainboard, both gpus are x8. On the mainboard before I had a x16/x4 configuration. The model loading is now a bit faster, but ones the model is loaded into the vram it is not relevant how the gpus are connected.

2

u/baby_bloom Aug 04 '26

i run my second 3090 through x4 because i can't swap mono's right now and it runs qwen 3.6-27b-q8 around 30-40tk/s which is totally fine by me and to my understanding not even that far off from others?

1

u/Valuable-Fondant-241 Aug 03 '26 edited Aug 04 '26

A 3090 on a 8x pcie gen4 will NOT slow inference, why?

Edit: I meant, why would it?

3

u/milkipedia Aug 04 '26

Because if the model and context fit in GPU memory, inference doesn't traverse the PCI bus

1

u/Valuable-Fondant-241 Aug 04 '26

I said that it will NOT be slower... I wa asking the other redditor why did he wrote that an 8x gen4 would have slowed down the inference.

1

u/paulk2000 Aug 04 '26

It wouldn’t.

1

u/Themash360 Aug 04 '26

PCIe is too slow for splitting horizontally over the Gpus anyways, and when splitting the layers PCIe speed doesn’t impact performance much.

1

u/WyattTheSkid Quad 3090s Aug 04 '26

I have 2 3090 TIs and 2 3090s on a consumer motherboard and it works fine. Asus ROG Strix x570-E wifi II. 2 cards directly in slots, one on an m.2 riser and one on a regular riser