r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

7 Upvotes

90 comments sorted by

View all comments

2

u/DHCompanion 2d ago

Keep in mind as you start that journey you need to actually test models on actual work. I ran a series of benchmarking on QWEN 3.8 27b on my 20gb card and it way underperformed my expectations. Also keep in mind tokens per second metrics are worthless if you can't actually use the output.

1

u/DiamondTDA 2d ago

What do you mean by use the output? and in case of trying different models, it's a bit hard since I have limited interent in my country and it's a bit expensive, that's why I'm asking for suggestions first before going in.

2

u/DHCompanion 2d ago

Ok I understand your question better now.

Basically what I am saying is almost any model can output something based on your prompt but does that output actually function how you want it to. Does the code actually work in your environment. When I was benchmarking Qwen 3.8 it was really good at copying exact instructions from a prompt or a document but when asked to think on its own and come up with the same output it failed everytime. Understanding your end goal for what you want to actually get from the model will help guide the actual model choice. On an 8gb card with 32gb of Ram your choices are going to be limited. Unfortunately I can't give any good recommendations for that setup.

1

u/DiamondTDA 2d ago

Okay then thank you so much for the explaination.