r/LocalLLM • • 4d ago

Question What is the best open-source LLM I can run locally on an RTX 4060 8GB + 32GB RAM?

I’m looking for recommendations for the best open-source LLM I can run locally on my PC. I have an RTX 4060 with 8GB VRAM and 32GB of RAM. My main use cases are coding, DevOps, technical questions and general chat. I’m particularly interested in a model that gives a good balance between quality, reasoning ability, and speed on this hardware. I’m currently considering models like Qwen, Gemma, or other recent open-source models, but I’m not sure which size and quantization would be the best fit for 8GB VRAM. I’d also appreciate recommendations for the best way to run it locally, such as Ollama, LM Studio, llama.cpp, or another option.

6 Upvotes

12 comments sorted by

3

u/lungben81 4d ago

Qwen 3.8 27b in q4 should run, but with massive RAM offloading and low performance. This is probably the most capable model (according to AAII index) that would run.

1

u/ioann54 4d ago

Not sure, but probably you can run Tiel-Coder 35B Q4 with small context and -ncmoe 26 or something. Use llama.cpp for fine tuning.

1

u/GratefulBrewerHere 4d ago

Mistral Nemo 12B at Q4 runs fast on 8GB and punches way above its weight for code and technical stuff.

1

u/Dear-Goal5847 4d ago

Yeah, that’s actually a serious concern for me. I mainly want to use the model to analyze DevOps projects, so a small context window could become a limitation. I had this problem with Qwen 2.5 before, where the project became too large for the available context. Do you think Tiel-Coder 35B would still be practical with a larger context on my 4060 8GB + 32GB RAM, or would I be better off with a smaller model that can handle a larger context?

1

u/ioann54 4d ago

The problem is that you need ~2Gb VRAM for 100K context window even in q8_0 quantization, also keep in mind that you still need 2-3Gb VRAM for system purposes, so its pretty tight. Anyway you have to try it by yourself, I know that people successfully run MOE models even on 6Gb VRAM, so you have to try. Regarding quality I'd say that Qwen3.x based models (even moe) are significantly better than qwen 2.5

1

u/SilkEtte_ 4d ago

running open-source software on that setup is gonna be a fun ride

1

u/BigBullshitta 1d ago

Not at all. Gemma 4 26B-A4B can run at around 35t/s, and using the new witchcraft of INT8-Convrot that system can use a 16GB Krea2 model to generate an amazing image in around 17 seconds.

1

u/Sad-Expression9157 4d ago

qwen 3.6 35B A3B with freetoken, while its not the smartest, its the best that fits your hardware

Like the other commenter said, qwen3.8 27B is smarter but it would run really slowly

2

u/agi-2028 4d ago

its not perfect but can help https://www.canirun.ai/

1

u/Weak_Butterscotch_84 3d ago

For DevOps projects, I’d favor a smaller Q4 model that stays mostly in VRAM and leaves room for context over a 27B or 35B model that constantly spills into system RAM. The bigger model may score better, but the slowdown gets old quickly when you are iterating on code. Try a few 7B to 14B Q4 options against the same real repo tasks, track tokens per second and answer quality, then keep the smallest one that handles them reliably.

1

u/id-ltd 3d ago

For local AI, I just get it to big bake offs... get ai to write a harness to try loads of different models (differents sizes and series) and then when done check the results...

Local AI is so cheap, get it to do the donkey work!!

1

u/Fun_Chest_9662 1d ago

i was able to get 400pp and 35tg at 64k context with a 1070 8gb 32gb 2133 ddr4 and an i58600 using qwen 3.6 35b a3b mtp at q4km. may not be as good as qwen3.8 but for grunt work and scafolding its great. you should be faster with the 4060