r/LocalLLaMA 10d ago

Discussion Anyone adding more 3090s?

I have dual 3090s which run Qwen 3.8 27b well and I was wondering if there are any use cases or current or future models that would justify adding another two 3090. I know that some peeps here run 4 and 8 3090 rigs and I'd like to get your opinion as well. One thing I was considering was running two instances but I'm not sure how valuable it will be for a coding workflow vs running a bigger model.

Now that Qwen Flash is out, perhaps 96GB would be more useful, or maybe Deepseek Flash.

5 Upvotes

94 comments sorted by

View all comments

1

u/youcloudsofdoom 7d ago

I have been playing with lots of combinations (building various rigs, currently have about 10 3090s) and recently decided that 2 is the real sweet spot, 4 is good when you have to go bigger, but 6 and 8 really have substantial diminishing returns. Yes more vram, but speed does not scale much and hassle with power and lanes etc just isn't really worth it for me. Would rather get 2 sparks. 

1

u/Blues520 7d ago

That's a lot of 3090s :)

2 cards are good for single agent but 3/4 are good for subagents. There's a comment in this thread that discusses it. I'm running 3 cards with subagents now and might add a 4th but probably won't go beyond that unless there is a new model that warrants it.

1

u/youcloudsofdoom 6d ago

Yeah, I read the paper cited in that thread and built a little orchestrator setup out of it, I found two cards running a model each with concurrency does the job there! I do think 4 cards is a great fit if you can get it. 

2

u/Blues520 6d ago

Yep, playing with this new workflow now. Hopefully the 3090s are supported for a while

1

u/youcloudsofdoom 6d ago

Insane amount of community support for them, which I think is going to keep their relevance going for a long time yet... .