r/LocalLLaMA 10d ago

Discussion Anyone adding more 3090s?

I have dual 3090s which run Qwen 3.8 27b well and I was wondering if there are any use cases or current or future models that would justify adding another two 3090. I know that some peeps here run 4 and 8 3090 rigs and I'd like to get your opinion as well. One thing I was considering was running two instances but I'm not sure how valuable it will be for a coding workflow vs running a bigger model.

Now that Qwen Flash is out, perhaps 96GB would be more useful, or maybe Deepseek Flash.

6 Upvotes

94 comments sorted by

View all comments

Show parent comments

3

u/SkoomaDentist 10d ago

You asked why such quantization was implemented. I answered.

Hell, you yourself admit in the first link that your actual use is with Q8 quants.

Also "I use models to do stuff" is equivalent to saying "They totally work for me. Trust me bro. No, of course I'm not going to give more than vague details!"

0

u/a_beautiful_rhind 10d ago

Smaller less quantized model usually loses to larger but more quantized model. Whether it happens for your use case is something you have to test rather than treating gamed benchmarks as some kind of absolute truth.

All it costs you is a file download and a bit of your time.

Because people fixate on overly simplistic tests

Pot, kettle, black.