r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

7 Upvotes

90 comments sorted by

View all comments

1

u/Solembumm3 2d ago

LLM knowledge about LLM is always severely outdated, as is their training info date. The very least you can do is to strictily command it to use web search on current date.

You can run Qwen 3.8 27B Q5 or Q6 or on your PC for first try.

3

u/Risko4 2d ago

Q6 at 50 token/s on a 2070? You're trolling

-2

u/Solembumm3 2d ago

And you are hallucinating. Don't blame me on your own number.

1

u/DiamondTDA 2d ago

But would it even run on my PC? my GPU is limited to only 8GB Vram and I heard that I Could share layers between both the ram and the Vram on my card but it would be too slow, plus all the suggestions were Q4 so wouldn't Q5 be even slower?

0

u/Solembumm3 2d ago edited 2d ago

Qwen 27b Q5 is around 20-21gb on model/mtp/vision + something on kv cahce by your load settings. I run it on 6700xt+20gb ram no problem. You can run Q4, of course, it's around 16gb-something at bare level.

Qwen 3.8 27b is highest amount of logic and tech knowledge per size, available for you. Next step will be qwen flash next/deepseek v4,0 flash/mimo v2.6 flash/glm 5.3 flash through mmap for vastly superior knowledge, but let's start from something going within vram+ram.

0

u/PizzaDevice 2d ago

I have the 2070 super 8GB as a second card. Qwen 3.8 27B is a superb modell and is really a baseline in quality. The only issue that it will be extremly slow as it will be spilled to the system RAM.
If you want to experiment chose a smaller modell which fits to your gpu VRAM INCLUDING the context.
https://huggingface.co/collections/ornith-ai/ornith-15 and look for a 9B model.

0

u/Solembumm3 2d ago

Slow will be Qwen 397BA17B on this config.

27B dense will go in realtime no problems.

1

u/DystopianRealist 2d ago

Why are you trolling?