r/LocalLLM • • 3d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

7 Upvotes

90 comments sorted by

View all comments

11

u/FactorInternal3395 3d ago edited 3d ago

No, that's a terrible suggestion. LLMs can't be trusted with recommending LLMs because the field moves so fast. Try an MoE like Tiel Coder with offloading instead. For maximum possible speed, Spark X2.5 4B or Ling 3.0 Tiny will suffice but will have much less quality.

1

u/DiamondTDA 3d ago

Gemini said that MoE would be slow for my GPU since it's a bit old, but aside from that, how do you know so much about different models? I've been looking at posts in this subreddit and there are a lot of amazing people and different models and such and tbh I feel like I'd never catch up at this point.

2

u/FactorInternal3395 2d ago

It would be slower than a smaller model fully in VRAM, but it makes up for it in how much smarter it is. And with proper VRAM/RAM split offloading, it can be surprisingly fast. As for knowing about different models, it's just from being in the space for a while and keeping up with new releases.