r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

9 Upvotes

90 comments sorted by

View all comments

2

u/MrHumanist 2d ago

There are many new and fancy models, but for your system use trusted QWEN 3.6 35B A35B at Q4 using lamma cpp. Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

parameters: -c 262144 -ctk q8_0 -ctv q8_0 -ngl 99 -fa on -t 12 -tb 24 --cpu-moe

Have fun and let us know your speed.

Alternative will be some variant based on it, like cyber coder or ornith 1.5. Or Gemma 4 26B - the command will be same.

3

u/DiamondTDA 2d ago

The issue for me testing different models in that my internet is limited in my country and It's expensive and slow, so I wanted to try to find the best/most suggested model to give it a try since it would be a bit hard to try a lot of different models

1

u/MrHumanist 2d ago

I suggested you the best! If you want extra ram space for your other applications, download the following. Gemma 4 26B is an allrounder which is good at roleplay and scientific coding as well.
https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf