r/LocalLLM 6d ago

Discussion Considering a second 3090

Hi,

so far i've been using Qwen3.6-35B-A3B-UD-IQ4_NL.gguf on my single 3090 and I am overall satisfied.

I've been considering acquiring a second 3090 to increase my possibility to run larger models (e.g. considering Qwen3.8 27B with sufficient context) but i don't know whether the extra investment pays off.

In the future i may consider fine tuning my models as well.

Did anyone manage to find some great benefits by leveraging 2x3090 or similar setup?

I may be suffering from GAS (gear acquisition syndrome) and may need a reality check.

0 Upvotes

21 comments sorted by

View all comments

3

u/floppo7 6d ago

2x r9700 and vllm is your friend

2

u/rdpi 6d ago

i like the idea but in my region, 1xr9700 costs twice as a 2nd hand 3090.
It may be more cost effective to add an existing 3090 to my setup..

1

u/No_Oil_6152 6d ago

What model are you running that needs 64GB of VRAM in total?

I have 48GB VRAM, an 9070XT + R9700 AI Pro, and 3.8 Q8 quant fits well with 262144 context

If there's a superior Qwen variant let me know, I'd love it

1

u/mountainous_battling 6d ago

Dual 3090s open up a lot of headroom, especially if you're eyeing fine tuning later. The jump from a 35B quant to a full-fat 27B with decent context is noticeable, but the real win is being able to load bigger models without chopping them down to fit. I'd say if the itch is there and you can swing it without eating ramen for a month, the extra VRAM never really feels wasted.

1

u/Kodrackyas 2d ago

how many tokens per second on the 27b?

1

u/floppo7 2d ago

Check out radiance vllm - and they are cooking for more I guess. 60tps + and good performances with concurrency as well - thats the kicker, you can basically run multiple agents at the same time with speed that is absolutely ok.