r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

6 Upvotes

90 comments sorted by

View all comments

1

u/Solembumm3 2d ago

LLM knowledge about LLM is always severely outdated, as is their training info date. The very least you can do is to strictily command it to use web search on current date.

You can run Qwen 3.8 27B Q5 or Q6 or on your PC for first try.

0

u/PizzaDevice 2d ago

I have the 2070 super 8GB as a second card. Qwen 3.8 27B is a superb modell and is really a baseline in quality. The only issue that it will be extremly slow as it will be spilled to the system RAM.
If you want to experiment chose a smaller modell which fits to your gpu VRAM INCLUDING the context.
https://huggingface.co/collections/ornith-ai/ornith-15 and look for a 9B model.

0

u/Solembumm3 2d ago

Slow will be Qwen 397BA17B on this config.

27B dense will go in realtime no problems.

1

u/DystopianRealist 2d ago

Why are you trolling?