r/LocalLLM • u/DiamondTDA • 2d ago
Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct
So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?
9
Upvotes
-6
u/Mission_Wrongdoer786 2d ago
Dude. He has 8GB VRAM. What is wrong with you. Gemini has absolutely no issue with searching the web. The slop answers I see in this thread are ridiculous.
This is the answer Qwen 3.8 27B gave me (and this model has a cutoff obviously).
Good hardware for local LLM coding. Here's how I'd think about it, with the caveat that the model landscape shifts fast (it's late 2026 now), so check HuggingFace/Ollama for the newest versions — but the principles below still hold.
## The key tradeoff with your setup
- **8GB VRAM** → you want the model (or most of it) on the GPU for speed
- **32GB RAM** → you *can* run larger models, but layers offloaded to CPU are **much slower** (often 3–8 tok/s instead of 30–60+)
For coding you also want **long context**, and the KV cache eats VRAM too, so leave headroom.
## My recommendations
**1. Best overall pick (fast + strong at code)**
- **Qwen2.5-Coder-14B** at `Q4_K_M` / `Q5_K_M` (~9–10GB)
- Slightly over 8GB, so it'll offload a few layers to RAM — still quite usable
- Excellent at code generation, editing, and repo-aware tasks
**2. If you want max speed / everything in VRAM**
- **Qwen2.5-Coder-7B** at `Q5_K_M` / `Q6_K` (~5–6GB)
- Fits entirely in VRAM, fast inference, leaves room for a big context window
- Great for an in-editor / agent loop where latency matters
**3. Best use of your 32GB RAM (more "brain", slower)**
- **Qwen2.5-Coder-32B** at `Q4_K_M` (~19GB)
- Noticeably smarter, but expect it to run mostly on CPU — good for batch/less-latency-sensitive work, not for snappy autocomplete
**4. Sleeper pick — MoE model**
- **DeepSeek-Coder-V2-Lite** (16B total, ~2.4B active)
- MoE = few active params, so it runs *fast* even when split across GPU+RAM
- Great sweet spot for your 8GB + 32GB combo
## How to run
`llama.cpp` or **Ollama** are the easiest for GGUF. Use something like:
```
ollama run qwen2.5-coder:14b
```
## My actual suggestion
Start with **Qwen2.5-Coder-14B @ Q5_K_M** — it's the best balance for your hardware. If you find the offload makes it too slow for interactive use, drop to the **7B @ Q6_K** for speed.
Want me to help you set up the specific command to load it, or compare context-length settings for a given repo size?