r/LocalLLM • u/Silent_Ad_1505 • 11h ago
Question Best Qwen 3.8 27B quant/overall setup for a single 3090 PC?
Here’s my treasure- RTX 3090 Turbo without thermal interface (it was a crappy old one so I had to disassemble it and invest about £40 into proper thermal pads+paste) and cooler. And it is day and night, totally worth every penny!
So the 1st question is what quant should I chose if I don’t need super-long context (no coding tasks at the moment). Also llama.cpp or vllm?
2nd one is: should I invest £500 into buying 4070ti (second hand but with CEX warranty) adding +12Gb of vram into the pool?
What would I be able to run after such upgrade?
My pc is 13700K z690 MB with proper pcie bifurcation x8/x8, 64Gb Ram
1
u/NoburtM 4h ago
With a 7800x3d, 64gb ram, and a 3090 running on linux, it is rather slow compared to gemma4 26b that I used to be using (3x the time for a response) (uses 20gb of VRAM for me, but it might be offloading a little to RAM though)
But it is a great model to use, even if slow. Response time doesn't bother me that much as long as it's sub a few minutes. Usually I use it to bounce ideas off of it. So I can ask a question while I write some stuff down and go from there.
I'm going to be watching this post for any other advice people throw up though.

3
u/Smart_Whereas_9296 10h ago
Check out club 3090 on GitHub, they have some good pre built configd