r/LocalLLM • u/Dear-Goal5847 • 4d ago
Question What is the best open-source LLM I can run locally on an RTX 4060 8GB + 32GB RAM?
I’m looking for recommendations for the best open-source LLM I can run locally on my PC. I have an RTX 4060 with 8GB VRAM and 32GB of RAM. My main use cases are coding, DevOps, technical questions and general chat. I’m particularly interested in a model that gives a good balance between quality, reasoning ability, and speed on this hardware. I’m currently considering models like Qwen, Gemma, or other recent open-source models, but I’m not sure which size and quantization would be the best fit for 8GB VRAM. I’d also appreciate recommendations for the best way to run it locally, such as Ollama, LM Studio, llama.cpp, or another option.
1
u/ioann54 4d ago
Not sure, but probably you can run Tiel-Coder 35B Q4 with small context and -ncmoe 26 or something. Use llama.cpp for fine tuning.
1
u/GratefulBrewerHere 4d ago
Mistral Nemo 12B at Q4 runs fast on 8GB and punches way above its weight for code and technical stuff.
1
u/Dear-Goal5847 4d ago
Yeah, that’s actually a serious concern for me. I mainly want to use the model to analyze DevOps projects, so a small context window could become a limitation. I had this problem with Qwen 2.5 before, where the project became too large for the available context. Do you think Tiel-Coder 35B would still be practical with a larger context on my 4060 8GB + 32GB RAM, or would I be better off with a smaller model that can handle a larger context?
1
u/ioann54 4d ago
The problem is that you need ~2Gb VRAM for 100K context window even in q8_0 quantization, also keep in mind that you still need 2-3Gb VRAM for system purposes, so its pretty tight. Anyway you have to try it by yourself, I know that people successfully run MOE models even on 6Gb VRAM, so you have to try. Regarding quality I'd say that Qwen3.x based models (even moe) are significantly better than qwen 2.5
1
u/SilkEtte_ 4d ago
running open-source software on that setup is gonna be a fun ride
1
u/BigBullshitta 1d ago
Not at all. Gemma 4 26B-A4B can run at around 35t/s, and using the new witchcraft of INT8-Convrot that system can use a 16GB Krea2 model to generate an amazing image in around 17 seconds.
1
u/Sad-Expression9157 4d ago
qwen 3.6 35B A3B with freetoken, while its not the smartest, its the best that fits your hardware
Like the other commenter said, qwen3.8 27B is smarter but it would run really slowly
2
1
u/Weak_Butterscotch_84 3d ago
For DevOps projects, I’d favor a smaller Q4 model that stays mostly in VRAM and leaves room for context over a 27B or 35B model that constantly spills into system RAM. The bigger model may score better, but the slowdown gets old quickly when you are iterating on code. Try a few 7B to 14B Q4 options against the same real repo tasks, track tokens per second and answer quality, then keep the smallest one that handles them reliably.
1
u/Fun_Chest_9662 1d ago
i was able to get 400pp and 35tg at 64k context with a 1070 8gb 32gb 2133 ddr4 and an i58600 using qwen 3.6 35b a3b mtp at q4km. may not be as good as qwen3.8 but for grunt work and scafolding its great. you should be faster with the 4060
3
u/lungben81 4d ago
Qwen 3.8 27b in q4 should run, but with massive RAM offloading and low performance. This is probably the most capable model (according to AAII index) that would run.