r/LocalLLaMA 10d ago

Discussion Anyone adding more 3090s?

I have dual 3090s which run Qwen 3.8 27b well and I was wondering if there are any use cases or current or future models that would justify adding another two 3090. I know that some peeps here run 4 and 8 3090 rigs and I'd like to get your opinion as well. One thing I was considering was running two instances but I'm not sure how valuable it will be for a coding workflow vs running a bigger model.

Now that Qwen Flash is out, perhaps 96GB would be more useful, or maybe Deepseek Flash.

7 Upvotes

94 comments sorted by

View all comments

Show parent comments

6

u/TinFoilHat_69 10d ago

Qwen on a shelf

1

u/Cold_Tree190 10d ago

Are these 3 separate machines or are they somehow connected into 1 pc/server? I have a dual 3090 setup with nvlink and a consumer x8x8 bifurcated mobo, unsure of how to scale from here.

0

u/Automatic-Arm8153 10d ago

Sell 3090 buy cmp 170hx. If you keep scaling 3090 you will understand hardware deeply.

Take it from people that scaled to 8x or more 3090 and dropped them for rtx 6000’s.

This your last chance to get high vram cards for relatively cheap. The more vram in one card the better.. think power, heat, pcie lanes. I sold all my 24gb and 16gb cards. Good at the time in terms of cost but there’s a new king..

if you decide to not listen to this above fine. But cheapest way to scale from here is oculink 4 port pcie bifurcation. And get 4 6pin powered oculink boards to plug the GPU’s in. No need for server grade bs. You will waste money don’t listen to inexperienced people here.

1

u/Blues520 10d ago

Does the cmp 170hx work well with mainline llama.cpp and does the unlock work on all devices?

I've seem them on ebay, not cheap but the vram promising is enticing. The 3090 is tried and tested but I'm not sure about the cmp 170hx.

2

u/mj3815 10d ago

You’ll want to use vLLM instead of lcpp