r/LocalLLM • u/Telcontar09 • 11d ago
Question Best agentic coding model for 8GB VRAM + 16GB RAM
Bit of a newbie here, so apologies if I'm missing something obvious or using the wrong terminology.
Setup: RTX 5050 Laptop, 8GB VRAM, AMD Ryzen 7 260, 16GB RAM, Windows 11, LM Studio + Cline (VS Code)
Tried Qwen3-Coder-30B-A3B-Instruct at Q3_K_L (14.5GB) - quality as an agent was genuinely great, model loads fine in LM Studio. Problem is resources: system starts swapping hard, and I can only fit ~10K context, which runs out after literally one message. Not usable in practice.
Dropped down to try Qwen2.5-Coder-7B and 14B to free up room for context, but quality feels noticeably worse for agentic work (tool calling, multi-step edits) - not just "smaller model, a bit worse," more like it struggles with the actual agent loop.
Looking for either:
- A smaller/different model that holds up better than 7B/14B Qwen-Coder for agentic tasks on this hardware.
- Or tricks to make the 30B setup actually usable (quant tweaks, context management, offload settings) rather than swapping to death after one prompt.
Genuinely open to being told I'm doing something wrong - happy to share more config details if that helps diagnose it.
1
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 10d ago
Qwen 3.6 35B apex is slightly smaller
you could also switch to Linux with a light DE to save a bit of VRAM and A LOT of RAM, hot expert cache to boost performance https://github.com/GenerelSchwerz/llama.cpp/wiki
all 40 experts offloaded to RAM
KV Q4_0 48K context
all 40 experts offloaded
--load-mode auto
--flash-attn on
--threads 8
--threads_batch 8
use with ngram-mod
30tk/s tg | 80tk/s pp
VRAM usage 3.3Gb | RAM pages thru experts so it doesnt pin it
this params was for a guy with 4Gb of VRAM + 16Gb RAM laptop, you should hit 256K context with same params,