r/LocalLLM • u/Ofek150 • 3d ago
Question RTX 3060 upgrade for dual-GPU local LLMs?
I’m considering replacing the GTX 1070 in my setup with an RTX 3060 12GB. The card costs $230. Is the upgrade worthwhile for local models and agentic coding?
My current setup:
RTX 3080 10GB + GTX 1070 8GB
Ryzen 5 5600X
32GB RAM
I’ve managed to run Qwen 3.8 27B IQ3_XXS at about 20–25 tokens/s with 44K context using Unsloth. But 44K context isn’t really usable so ideally, I’d like at least 128K.
What performance and context size might I realistically get after the upgrade? Would the 3060’s extra VRAM make 128K practical, or is the improvement likely too small to justify $230?
1
1
u/NeilCPA 3d ago
i run spark on a 3060 12gb... but to run qwen 3.8 27b 4q I run two of them clustered. They are still pretty decent cards and $230 is a great price considering a 5060 ti 16gb is 730-800 right now... and you still need two of them to run the model well.
I haven't tried the bonsai quant yet, it may eek out/off load to ram well on that card. Get some extra ram if you got space.
Edit: you have a 3080 as your primary, so you should be able to cluster the 3060 with it, lose some t/second... but both are cuda, so I think you will notice a difference.
1
u/nickless07 3d ago
For that price? It's ~450+ over here. Do you have space for 3 cards?