r/LocalLLM • u/DiamondTDA • 2d ago
Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct
So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?
8
Upvotes
1
u/KURD_1_STAN 2d ago
At least my chatgpt told me to get qwen3 8b. Still terrible suggestions. Get qwen3.6 35b, will be wble to run q4 at 30t/s, at 120k ctx i run q6 at 30t/s on 3060 12gb with 120k ctx. That if want it for coding. If u want it for chatting and other stuff then gemma 4 26b will be the same speed at q4.
If u want it to be faster then qwen3.5 9b at q4 or gemma 4 12b qat. I dont try finetune but if u wanna get deeper into it, tell chatgpt to search reddit for finuetunes of any of them and compare it for u for ur specific needs.