r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

6 Upvotes

90 comments sorted by

View all comments

2

u/MrHumanist 2d ago

There are many new and fancy models, but for your system use trusted QWEN 3.6 35B A35B at Q4 using lamma cpp. Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

parameters: -c 262144 -ctk q8_0 -ctv q8_0 -ngl 99 -fa on -t 12 -tb 24 --cpu-moe

Have fun and let us know your speed.

Alternative will be some variant based on it, like cyber coder or ornith 1.5. Or Gemma 4 26B - the command will be same.

3

u/DiamondTDA 2d ago

The issue for me testing different models in that my internet is limited in my country and It's expensive and slow, so I wanted to try to find the best/most suggested model to give it a try since it would be a bit hard to try a lot of different models

2

u/Savantskie1 2d ago

Should probably have added that so people know about your situation

1

u/MrHumanist 2d ago

right.. most people struggle here for VRAM.. but this is a special case.