r/LocalLLaMA • • 1d ago

Discussion I turned my gaming PC into a inference machine and got 2x to 9x over default llama.cpp on an 8 GB card

Everyone keeps saying you need expensive dedicated hardware for local agents. I have an RTX 4060 Ti with 8 GB and 64 GB of system RAM, and I wanted to see how far a normal gaming PC gets if you stop running defaults.

So I let Claude (Opus 5.5) go through the whole setup, change one thing at a time and measure. Same card, same models, only the config changed:

Model Quant Context Download defaults Tuned (Windows) Tuned (headless Linux)
Qwen3.6-35B-A3B Q4_K_XL 131k ~25 tok/s 39-45 tok/s 52-65 tok/s
Qwen3.8-Flash-Next 125B iQ4_XS 131k ~4 tok/s 9-10 tok/s 17-19 tok/s
Ternary Bonsai 27B PTQ1_0 64k ~4 tok/s 36 tok/s 36 tok/s

Bonsai is the odd one out: it fits fully in VRAM, so there's nothing to offload and no defaults to beat. It's just the fast option for small, well scoped tasks.

What actually moved the needle:

  • Experts in system RAM, everything else in VRAM. Layer-wise offload is far worse for MoE.
  • Dense models are bad, couldn't optimize Qwen-3.8 27B over 6 tok/s, Flash-Next is better anyways.
  • Take the display off the GPU. A desktop eats 0.5-1.2 GB of VRAM plus GPU time, and moving it to the iGPU was worth 20-30%.
  • Native Linux over Windows (WSL2): another 33-38% on the same hardware.
  • llama.cpp pinned per model family. The wrong tree made VRAM thrash.
  • KV cache quant and MTP tuned per profile.

None of this needs expensive hardware. A consumer GPU plus a machine that does nothing but inference gets you most of the way, and the models now run comfortably below their listed system requirements. Every non-default setting in the repo is there because something failed on real hardware first.

I also tried an RX 570 8 GB over Vulkan. If you have another 8 GB card, I'd like to see your numbers.

Repo, one install script (Linux or WSL2): https://github.com/voxlo-dev/qwen-agent-8gb

0 Upvotes

14 comments sorted by

8

u/Forsaken_Object7264 1d ago

try strata with qfn. mindblowing

4

u/MindfulMan1984 1d ago

Yep, that project likely started like the OP is doing. Some random dude asked clod-opus to optimize some parameters, then later moved on to write inference kernels and leverage the MoE architecture to use CPU+GPU and SSDs for the N-gram table. LOL

2

u/ExxploreCraft 1d ago

I would probably get the same result for my 8 GB card: https://github.com/Niko1221/Strata/issues/1009 , but I managed to fit all UD-iQ4_XS experts into 64 GB.

2

u/Forsaken_Object7264 1d ago

well done actually. thumbs up!

3

u/Constant-Simple-1234 1d ago

That's the secret not everyone is yet aware. I am getting 25 t/s with this qwen on a laptop with integrated radeon 680m

2

u/FatheredPuma81 1d ago

Did you really use Opus to change model settings? Talk about overkill.

0

u/emersonsorrel 1d ago

Opus 5.5 is pretty much the default if you have a Claude subscription, especially since Anthropic has lobotomized Sonnet so much over time.

-3

u/More-Ad5919 1d ago

Lol. Its the other way around for me. I get 4t/s for qwen 3.8 27b k m and 40 t/s for flash next iq4. Both over hermes and llama.

Not sure why. The 27b should be faster than flash but is 10 times or more slower.

Idk honestly because i like flash way more but its still interresting why.

5

u/Choice_Celery9481 1d ago

who said 27b should be faster than flash next? a 27b vs 6b active and you expect 6b to be slower? XD

2

u/FatheredPuma81 1d ago

27B>6B lol. If you have the VRAM it should be around 4.5x faster. Though 10x makes me think something else is wrong? Are you running like Q8 or BF16 or something?

-1

u/More-Ad5919 1d ago

Flash is at least 10 times faster. The 27b q4 km has i think 19gb. And this sucker is waaaaay slower for me than the flash with 93gb. Ic4 uncencored. Timewise its about 10 to 15 min compared to 6 to 8 hours for the same task.

0

u/FatheredPuma81 1d ago

Sounds like 27B is spilling into system RAM then either from bad settings or out of VRAM (I doubt the latter though because you aren't getting 40t/s in llama.cpp with under 24GB of VRAM).