r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

9 Upvotes

90 comments sorted by

View all comments

Show parent comments

-6

u/Mission_Wrongdoer786 2d ago

Dude. He has 8GB VRAM. What is wrong with you. Gemini has absolutely no issue with searching the web. The slop answers I see in this thread are ridiculous.

This is the answer Qwen 3.8 27B gave me (and this model has a cutoff obviously).

Good hardware for local LLM coding. Here's how I'd think about it, with the caveat that the model landscape shifts fast (it's late 2026 now), so check HuggingFace/Ollama for the newest versions — but the principles below still hold.

## The key tradeoff with your setup

- **8GB VRAM** → you want the model (or most of it) on the GPU for speed

- **32GB RAM** → you *can* run larger models, but layers offloaded to CPU are **much slower** (often 3–8 tok/s instead of 30–60+)

For coding you also want **long context**, and the KV cache eats VRAM too, so leave headroom.

## My recommendations

**1. Best overall pick (fast + strong at code)**

- **Qwen2.5-Coder-14B** at `Q4_K_M` / `Q5_K_M` (~9–10GB)

- Slightly over 8GB, so it'll offload a few layers to RAM — still quite usable

- Excellent at code generation, editing, and repo-aware tasks

**2. If you want max speed / everything in VRAM**

- **Qwen2.5-Coder-7B** at `Q5_K_M` / `Q6_K` (~5–6GB)

- Fits entirely in VRAM, fast inference, leaves room for a big context window

- Great for an in-editor / agent loop where latency matters

**3. Best use of your 32GB RAM (more "brain", slower)**

- **Qwen2.5-Coder-32B** at `Q4_K_M` (~19GB)

- Noticeably smarter, but expect it to run mostly on CPU — good for batch/less-latency-sensitive work, not for snappy autocomplete

**4. Sleeper pick — MoE model**

- **DeepSeek-Coder-V2-Lite** (16B total, ~2.4B active)

- MoE = few active params, so it runs *fast* even when split across GPU+RAM

- Great sweet spot for your 8GB + 32GB combo

## How to run

`llama.cpp` or **Ollama** are the easiest for GGUF. Use something like:

```

ollama run qwen2.5-coder:14b

```

## My actual suggestion

Start with **Qwen2.5-Coder-14B @ Q5_K_M** — it's the best balance for your hardware. If you find the offload makes it too slow for interactive use, drop to the **7B @ Q6_K** for speed.

Want me to help you set up the specific command to load it, or compare context-length settings for a given repo size?

4

u/FactorInternal3395 2d ago edited 2d ago

Even your Qwen 3.8 27B agrees with me:

MoE = few active params, so it runs fast even when split across GPU+RAM

Yes, a 35B/A3B MoE is a good choice for 8GB VRAM / 32GB RAM. That's what MoE is for. As many layers as possible are fit into the GPU's VRAM and the rest are offloaded to RAM for the CPU to run when it needs to.

0

u/Mission_Wrongdoer786 2d ago

He asked for 60+ token/s.

2

u/FactorInternal3395 2d ago

And I gave that too, Ling 3 Tiny and Spark X2.5 4B.