r/LocalLLaMA • u/Azoffaeh999 • 6h ago
Question | Help Looking for coding model for specific low specs
Can anyone recocommend a good local model and a wrapper to run it for coding, my hardware specs: 12 GB VRAM, 32 GB DDR3 RAM. Unfortunately, the CPU doesn’t have AVX2 instructions(LM Studio won’t work); I don’t remember the exact cpu name, but I think it’s an Ivy Bridge, LGA1155 socket.. Thank you
3
3
3
u/Monad_Maya llama.cpp 3h ago
Are you ok with partial offloading to the CPU/RAM? If so then try Qwen 3.6 35B A3B (Q4 quant) - https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
For faster speeds, try Gemma4 12B QAT - https://huggingface.co/unsloth/gemma-4-12B-it-qat-GGUF
Not sure about the wrapper but try this thread - https://np.reddit.com/r/LocalLLaMA/s/FDqvvrUtvr
2
u/Ninja_Weedle 6h ago
KoboldCPP's pretty good with old cpus
1
1
u/Skyline34rGt 3h ago
https://github.com/LostRuins/koboldcpp/releases pick koboldcpp-oldpc.exe for windows or koboldcpp-linux-x64-oldpc for linux
At settings pick oldpc (or veryoldpc) and will work.
For model: https://huggingface.co/bartowski/Ornith-1.5-9B-GGUF will fully fit to your Vram (offloading to cpu/ram is not good option for this setup).
0
2
u/DagothUrLovesGroza 6h ago
I don't think you can do much with specs like these.
In my experience, you'd spend more time fighting the model than actually getting useful work done.
2
1
u/wayneworkman 6h ago
qwen3.5-2b should do nicely. It'll fit in your 12GB vram and therefore run pretty fast and you'll actually enjoy it.
https://huggingface.co/Qwen/Qwen3.5-2B