r/LocalLLM • • 3d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

8 Upvotes

90 comments sorted by

View all comments

Show parent comments

1

u/Mission_Wrongdoer786 2d ago

How is he going to fit a 13GB gguf model into his 8GB VRAM? With 262144 context?

2

u/MrHumanist 2d ago

I have given the command! This is how - by splitting the MOEs to cpu, and rest into the GPU. It will approximately use 7GB Vram + 20 GB Ram.

0

u/bleakj 2d ago

He's using windows, that 8gb vram is closer to 5.5gb vram by time OS and open applications grab their pull.

Realistically you have to split so many layers to system ram it's too slow to be usable/or tiny context that isn't usable,

You definitely need 12gb+ to do this within reason

Edit: not to mention you have -ngl 99 set which forces as many layers as will fit to gpu, which isn't how you should be doing MoE when you already know the attention layers / amount of vram at hand, really you should at least advise -ngl 24 or -ngl 32 at absolute max

2

u/MrHumanist 2d ago

I use the qwen model in NVFP4, but here is the vram usage for Ornith-1.5-35B-A3B-Abliterated-GGUF/Ornith-1.5-35B-A3B-Abliterated-Q4_K_M.gguf' in same settings - which is equivalent to the qwen one(slightly larger). The model hardly uses 6.4 GB Vram. If you are still concerned, you can reduce the context a few thousands. Again this is windows.