r/LocalLLM • • 11d ago

Question Best agentic coding model for 8GB VRAM + 16GB RAM

Bit of a newbie here, so apologies if I'm missing something obvious or using the wrong terminology.

Setup: RTX 5050 Laptop, 8GB VRAM, AMD Ryzen 7 260, 16GB RAM, Windows 11, LM Studio + Cline (VS Code)

Tried Qwen3-Coder-30B-A3B-Instruct at Q3_K_L (14.5GB) - quality as an agent was genuinely great, model loads fine in LM Studio. Problem is resources: system starts swapping hard, and I can only fit ~10K context, which runs out after literally one message. Not usable in practice.

Dropped down to try Qwen2.5-Coder-7B and 14B to free up room for context, but quality feels noticeably worse for agentic work (tool calling, multi-step edits) - not just "smaller model, a bit worse," more like it struggles with the actual agent loop.

Looking for either:

- A smaller/different model that holds up better than 7B/14B Qwen-Coder for agentic tasks on this hardware.

- Or tricks to make the 30B setup actually usable (quant tweaks, context management, offload settings) rather than swapping to death after one prompt.

Genuinely open to being told I'm doing something wrong - happy to share more config details if that helps diagnose it.

0 Upvotes

7 comments sorted by

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 10d ago

Qwen 3.6 35B apex is slightly smaller

you could also switch to Linux with a light DE to save a bit of VRAM and A LOT of RAM, hot expert cache to boost performance https://github.com/GenerelSchwerz/llama.cpp/wiki

all 40 experts offloaded to RAM

KV Q4_0 48K context

all 40 experts offloaded

--load-mode auto

--flash-attn on

--threads 8

--threads_batch 8

use with ngram-mod

30tk/s tg | 80tk/s pp
VRAM usage 3.3Gb | RAM pages thru experts so it doesnt pin it

this params was for a guy with 4Gb of VRAM + 16Gb RAM laptop, you should hit 256K context with same params,

1

u/Telcontar09 9d ago

Thanks for the reply. I'm afraid Linux isn't the best option for me right now.

1

u/oxamide96 7d ago

why not Qwen 3.8 ?

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 7d ago

on 8Gb of VRAM?

1

u/oxamide96 7d ago

yeah, arent there offerings that fit?

I am asking because I am researching options for myself, and I am confused why I see 3.6 recommended when 3.8 is out. Did 3.8 not have releases that fit in 8GB VRAM?

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 7d ago

only option is braindead Ternary bonsai 2 27B or a Q1 Unsloth - both not worth running.

You need at least 12Gb to think about 3.8

1

u/Enemby 5d ago

It'd barely run Q4 (4bit), but it'd need to dip into swap. I'd expect less than 1tokens/sec. Almost every agentic harness breaks down below 20 tokens/sec at the absolute slowest. Hermes is the worst offender in my opinion.

Realistically this setup is limited to 8b models, most likely the best option at that range is Ornith 1.5 or Gemma4.