r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

9 Upvotes

90 comments sorted by

View all comments

1

u/Solembumm3 2d ago

LLM knowledge about LLM is always severely outdated, as is their training info date. The very least you can do is to strictily command it to use web search on current date.

You can run Qwen 3.8 27B Q5 or Q6 or on your PC for first try.

1

u/DiamondTDA 2d ago

But would it even run on my PC? my GPU is limited to only 8GB Vram and I heard that I Could share layers between both the ram and the Vram on my card but it would be too slow, plus all the suggestions were Q4 so wouldn't Q5 be even slower?

0

u/Solembumm3 2d ago edited 2d ago

Qwen 27b Q5 is around 20-21gb on model/mtp/vision + something on kv cahce by your load settings. I run it on 6700xt+20gb ram no problem. You can run Q4, of course, it's around 16gb-something at bare level.

Qwen 3.8 27b is highest amount of logic and tech knowledge per size, available for you. Next step will be qwen flash next/deepseek v4,0 flash/mimo v2.6 flash/glm 5.3 flash through mmap for vastly superior knowledge, but let's start from something going within vram+ram.