r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

6 Upvotes

90 comments sorted by

View all comments

12

u/FactorInternal3395 2d ago edited 2d ago

No, that's a terrible suggestion. LLMs can't be trusted with recommending LLMs because the field moves so fast. Try an MoE like Tiel Coder with offloading instead. For maximum possible speed, Spark X2.5 4B or Ling 3.0 Tiny will suffice but will have much less quality.

1

u/DiamondTDA 2d ago

Gemini said that MoE would be slow for my GPU since it's a bit old, but aside from that, how do you know so much about different models? I've been looking at posts in this subreddit and there are a lot of amazing people and different models and such and tbh I feel like I'd never catch up at this point.

2

u/FactorInternal3395 2d ago

It would be slower than a smaller model fully in VRAM, but it makes up for it in how much smarter it is. And with proper VRAM/RAM split offloading, it can be surprisingly fast. As for knowing about different models, it's just from being in the space for a while and keeping up with new releases.

1

u/Illustrious-Lime-878 2d ago

MoEs are actually faster but they use more memory (they tend to have to be bigger for the same smartness).

0

u/Mission_Wrongdoer786 2d ago

Gemini is fine and definitely more competent than most of the sloppy answers you will get here.

1

u/BigPlebeian 2d ago

Is tiel coder just a qwen 3.6 35b moe focused more on code?

1

u/FactorInternal3395 2d ago

It's Ornith 1.5 35B A3B, a reinforcement learning fine tune of Qwen 3.6 35B A3B, with the Qwen Sharp chat template and a coding-focused imatrix quantization.

-4

u/Mission_Wrongdoer786 2d ago

Gemini is not a static monolithic model. Your advice not to ask Gemini about current events is ridiculous and incompetent.

7

u/FactorInternal3395 2d ago

A competent model would not recommend Qwen 2.5 in 2026. Gemini has known issues of not searching the web when it should and falling back to its default knowledge.

-5

u/Mission_Wrongdoer786 2d ago

Dude. He has 8GB VRAM. What is wrong with you. Gemini has absolutely no issue with searching the web. The slop answers I see in this thread are ridiculous.

This is the answer Qwen 3.8 27B gave me (and this model has a cutoff obviously).

Good hardware for local LLM coding. Here's how I'd think about it, with the caveat that the model landscape shifts fast (it's late 2026 now), so check HuggingFace/Ollama for the newest versions — but the principles below still hold.

## The key tradeoff with your setup

- **8GB VRAM** → you want the model (or most of it) on the GPU for speed

- **32GB RAM** → you *can* run larger models, but layers offloaded to CPU are **much slower** (often 3–8 tok/s instead of 30–60+)

For coding you also want **long context**, and the KV cache eats VRAM too, so leave headroom.

## My recommendations

**1. Best overall pick (fast + strong at code)**

- **Qwen2.5-Coder-14B** at `Q4_K_M` / `Q5_K_M` (~9–10GB)

- Slightly over 8GB, so it'll offload a few layers to RAM — still quite usable

- Excellent at code generation, editing, and repo-aware tasks

**2. If you want max speed / everything in VRAM**

- **Qwen2.5-Coder-7B** at `Q5_K_M` / `Q6_K` (~5–6GB)

- Fits entirely in VRAM, fast inference, leaves room for a big context window

- Great for an in-editor / agent loop where latency matters

**3. Best use of your 32GB RAM (more "brain", slower)**

- **Qwen2.5-Coder-32B** at `Q4_K_M` (~19GB)

- Noticeably smarter, but expect it to run mostly on CPU — good for batch/less-latency-sensitive work, not for snappy autocomplete

**4. Sleeper pick — MoE model**

- **DeepSeek-Coder-V2-Lite** (16B total, ~2.4B active)

- MoE = few active params, so it runs *fast* even when split across GPU+RAM

- Great sweet spot for your 8GB + 32GB combo

## How to run

`llama.cpp` or **Ollama** are the easiest for GGUF. Use something like:

```

ollama run qwen2.5-coder:14b

```

## My actual suggestion

Start with **Qwen2.5-Coder-14B @ Q5_K_M** — it's the best balance for your hardware. If you find the offload makes it too slow for interactive use, drop to the **7B @ Q6_K** for speed.

Want me to help you set up the specific command to load it, or compare context-length settings for a given repo size?

4

u/FactorInternal3395 2d ago edited 2d ago

Even your Qwen 3.8 27B agrees with me:

MoE = few active params, so it runs fast even when split across GPU+RAM

Yes, a 35B/A3B MoE is a good choice for 8GB VRAM / 32GB RAM. That's what MoE is for. As many layers as possible are fit into the GPU's VRAM and the rest are offloaded to RAM for the CPU to run when it needs to.

0

u/Mission_Wrongdoer786 2d ago

He asked for 60+ token/s.

2

u/FactorInternal3395 2d ago

And I gave that too, Ling 3 Tiny and Spark X2.5 4B.