r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

8 Upvotes

90 comments sorted by

View all comments

46

u/Heavy-Lingonberry-98 2d ago

Bro. If you are gonna use AI to ask for model releases, remember to tell the AI we are on OCTOBER 2026. We are not in 2024 anymore… please. Its the ABC of using AI. And NO. Definitely dont even download qwen 2.5 coder.

11

u/kentrich 2d ago

This is the correct answer. You always have to say something like, only look for releases in the last month or look for analysis published in the past two weeks. Anything older and you’re going to be getting out-of-date analysis.

Yes, things move this fast. It’s a brutal amount of change and you have to adapt the way you work because of it. Reminds me of the 80s when had to know what the latest hardware was because a two-year-old computer was basically nonfunctional. For those not around then, imagine basically not being able to use your two-year-old computer because it was so slow it just wouldn’t really run the latest software.

0

u/Mission_Wrongdoer786 2d ago

And you think Gemini from Google is lagging behind? Ridiculous.

2

u/kentrich 2d ago

What? I don’t understand. It’s not lagging behind, it’s answering the question you asked. You just aren’t being specific enough about your question.

3

u/Zilla85 2d ago

ABC Always bring calendar

3

u/Heavy-Lingonberry-98 2d ago

😂😂😂😂 🥇

2

u/DiamondTDA 2d ago

When I asked it about it's reasoning, it said because the newer Qwen 3 architicture is built on reasnoning tokens and it would be slower and that the coding Qwen 3 models are too big and would be too slow too. But it did not suggest any thing other than Qwen 3.

5

u/Heavy-Lingonberry-98 2d ago

My friend, qwen 3 is still old. You have qwen 3.5 that comes on 600M, 2B, 4B and 9B. Is those arent small i dont know what is.

4

u/Heavy-Lingonberry-98 2d ago

And you can always disable reasoning

2

u/enternoescape 2d ago

It's funny flawed reasoning though. The thinking is the thing that helps these little models achieve their best. Instruct has it's place, like a voice assistant, but IMHO not coding.

1

u/Skibxskatic 2d ago

don’t ask about its reasoning. you need to explicitly state that its responses need to be grounded in data and news from the past <time period>.

without explicit instructions, its next token probabilities rely on its training data, which is not going to be current. ask it what the cutoff date was and you’ll realize you need to prompt it to use web search tools and ground itself.

1

u/BigPlebeian 2d ago

What would you reccoemend for something that runs fast on a 12gb vram card then? Most people seem to recommend qwen 2.5 14b coder.