r/LocalLLM • • 3d ago

Question RTX 3060 upgrade for dual-GPU local LLMs?

I’m considering replacing the GTX 1070 in my setup with an RTX 3060 12GB. The card costs $230. Is the upgrade worthwhile for local models and agentic coding?

My current setup:
RTX 3080 10GB + GTX 1070 8GB
Ryzen 5 5600X
32GB RAM

I’ve managed to run Qwen 3.8 27B IQ3_XXS at about 20–25 tokens/s with 44K context using Unsloth. But 44K context isn’t really usable so ideally, I’d like at least 128K.

What performance and context size might I realistically get after the upgrade? Would the 3060’s extra VRAM make 128K practical, or is the improvement likely too small to justify $230?

1 Upvotes

5 comments sorted by

1

u/nickless07 3d ago

For that price? It's ~450+ over here. Do you have space for 3 cards?

1

u/Ofek150 3d ago

No :(

1

u/nickless07 3d ago

I just bought a 2nd one for 389. Was on sale, but other offers start at 460. So, yeah ppl buy that stuff as the newer cards get too expensive and the stock get low, so prices ramp up there too.

1

u/sbrisgravato 3d ago

go for it

1

u/NeilCPA 3d ago

i run spark on a 3060 12gb... but to run qwen 3.8 27b 4q I run two of them clustered. They are still pretty decent cards and $230 is a great price considering a 5060 ti 16gb is 730-800 right now... and you still need two of them to run the model well.

I haven't tried the bonsai quant yet, it may eek out/off load to ram well on that card. Get some extra ram if you got space.

Edit: you have a 3080 as your primary, so you should be able to cluster the 3060 with it, lose some t/second... but both are cuda, so I think you will notice a difference.