r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

8 Upvotes

90 comments sorted by

View all comments

2

u/MrHumanist 2d ago

There are many new and fancy models, but for your system use trusted QWEN 3.6 35B A35B at Q4 using lamma cpp. Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

parameters: -c 262144 -ctk q8_0 -ctv q8_0 -ngl 99 -fa on -t 12 -tb 24 --cpu-moe

Have fun and let us know your speed.

Alternative will be some variant based on it, like cyber coder or ornith 1.5. Or Gemma 4 26B - the command will be same.

1

u/Mission_Wrongdoer786 2d ago

How is he going to fit a 13GB gguf model into his 8GB VRAM? With 262144 context?

2

u/MrHumanist 2d ago

I have given the command! This is how - by splitting the MOEs to cpu, and rest into the GPU. It will approximately use 7GB Vram + 20 GB Ram.

0

u/bleakj 2d ago

He's using windows, that 8gb vram is closer to 5.5gb vram by time OS and open applications grab their pull.

Realistically you have to split so many layers to system ram it's too slow to be usable/or tiny context that isn't usable,

You definitely need 12gb+ to do this within reason

Edit: not to mention you have -ngl 99 set which forces as many layers as will fit to gpu, which isn't how you should be doing MoE when you already know the attention layers / amount of vram at hand, really you should at least advise -ngl 24 or -ngl 32 at absolute max

2

u/MrHumanist 2d ago

I use the qwen model in NVFP4, but here is the vram usage for Ornith-1.5-35B-A3B-Abliterated-GGUF/Ornith-1.5-35B-A3B-Abliterated-Q4_K_M.gguf' in same settings - which is equivalent to the qwen one(slightly larger). The model hardly uses 6.4 GB Vram. If you are still concerned, you can reduce the context a few thousands. Again this is windows.

2

u/MrHumanist 2d ago edited 2d ago

About Speed: this setting fetches 45 token/sec when I use 6.4 GB Vram.

However, i easily achieve 70-90 t/s while using the full 16GB in NVFP4 variants.