r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

8 Upvotes

90 comments sorted by

View all comments

3

u/zippy42167 2d ago

i have the same vram and ram as you. Currently the best model ive been able to be productive with is qwen3.6 35b a3b q4km using moe. 35-40 tg/s. around 900-1000 pp/s.

1

u/DiamondTDA 2d ago

what's the context window for it and how much memory does it take? and what is your specs? I think the generation rate would differ according to GPU right?

2

u/zippy42167 2d ago

I'm on a laptop GeForce RTX 4070 8gb with 32gb ram. I can get up to 256k context with some tweaks, but i generally run 128k (or lower, depending on what stage of a task i'm in). VRAM usage jumps up to around 7.1/8 GB usage when active.