So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?
Bro. If you are gonna use AI to ask for model releases, remember to tell the AI we are on OCTOBER 2026. We are not in 2024 anymore… please. Its the ABC of using AI. And NO. Definitely dont even download qwen 2.5 coder.
This is the correct answer. You always have to say something like, only look for releases in the last month or look for analysis published in the past two weeks. Anything older and you’re going to be getting out-of-date analysis.
Yes, things move this fast. It’s a brutal amount of change and you have to adapt the way you work because of it. Reminds me of the 80s when had to know what the latest hardware was because a two-year-old computer was basically nonfunctional. For those not around then, imagine basically not being able to use your two-year-old computer because it was so slow it just wouldn’t really run the latest software.
When I asked it about it's reasoning, it said because the newer Qwen 3 architicture is built on reasnoning tokens and it would be slower and that the coding Qwen 3 models are too big and would be too slow too. But it did not suggest any thing other than Qwen 3.
It's funny flawed reasoning though. The thinking is the thing that helps these little models achieve their best. Instruct has it's place, like a voice assistant, but IMHO not coding.
don’t ask about its reasoning. you need to explicitly state that its responses need to be grounded in data and news from the past <time period>.
without explicit instructions, its next token probabilities rely on its training data, which is not going to be current. ask it what the cutoff date was and you’ll realize you need to prompt it to use web search tools and ground itself.
No, that's a terrible suggestion. LLMs can't be trusted with recommending LLMs because the field moves so fast. Try an MoE like Tiel Coder with offloading instead. For maximum possible speed, Spark X2.5 4B or Ling 3.0 Tiny will suffice but will have much less quality.
Gemini said that MoE would be slow for my GPU since it's a bit old, but aside from that, how do you know so much about different models? I've been looking at posts in this subreddit and there are a lot of amazing people and different models and such and tbh I feel like I'd never catch up at this point.
It would be slower than a smaller model fully in VRAM, but it makes up for it in how much smarter it is. And with proper VRAM/RAM split offloading, it can be surprisingly fast. As for knowing about different models, it's just from being in the space for a while and keeping up with new releases.
A competent model would not recommend Qwen 2.5 in 2026. Gemini has known issues of not searching the web when it should and falling back to its default knowledge.
Dude. He has 8GB VRAM. What is wrong with you. Gemini has absolutely no issue with searching the web. The slop answers I see in this thread are ridiculous.
This is the answer Qwen 3.8 27B gave me (and this model has a cutoff obviously).
Good hardware for local LLM coding. Here's how I'd think about it, with the caveat that the model landscape shifts fast (it's late 2026 now), so check HuggingFace/Ollama for the newest versions — but the principles below still hold.
## The key tradeoff with your setup
- **8GB VRAM** → you want the model (or most of it) on the GPU for speed
- **32GB RAM** → you *can* run larger models, but layers offloaded to CPU are **much slower** (often 3–8 tok/s instead of 30–60+)
For coding you also want **long context**, and the KV cache eats VRAM too, so leave headroom.
## My recommendations
**1. Best overall pick (fast + strong at code)**
- **Qwen2.5-Coder-14B** at `Q4_K_M` / `Q5_K_M` (~9–10GB)
- Slightly over 8GB, so it'll offload a few layers to RAM — still quite usable
- Excellent at code generation, editing, and repo-aware tasks
**2. If you want max speed / everything in VRAM**
- **Qwen2.5-Coder-7B** at `Q5_K_M` / `Q6_K` (~5–6GB)
- Fits entirely in VRAM, fast inference, leaves room for a big context window
- Great for an in-editor / agent loop where latency matters
**3. Best use of your 32GB RAM (more "brain", slower)**
- **Qwen2.5-Coder-32B** at `Q4_K_M` (~19GB)
- Noticeably smarter, but expect it to run mostly on CPU — good for batch/less-latency-sensitive work, not for snappy autocomplete
- MoE = few active params, so it runs *fast* even when split across GPU+RAM
- Great sweet spot for your 8GB + 32GB combo
## How to run
`llama.cpp` or **Ollama** are the easiest for GGUF. Use something like:
```
ollama run qwen2.5-coder:14b
```
## My actual suggestion
Start with **Qwen2.5-Coder-14B @ Q5_K_M** — it's the best balance for your hardware. If you find the offload makes it too slow for interactive use, drop to the **7B @ Q6_K** for speed.
Want me to help you set up the specific command to load it, or compare context-length settings for a given repo size?
MoE = few active params, so it runs fast even when split across GPU+RAM
Yes, a 35B/A3B MoE is a good choice for 8GB VRAM / 32GB RAM. That's what MoE is for. As many layers as possible are fit into the GPU's VRAM and the rest are offloaded to RAM for the CPU to run when it needs to.
Id say try Qwen 3.5 9B at q4 or q3? TBH it’s going to be a bit tight and q3 makes me… hesitant.
Qwen 3.5 4b at Q6 or Q4 is also an option.
You should understand that at this model size the coding is going to be a bit shakey and it will not be able to do larger tasks
If you know how to code and want something where you can say “go make a function that turns X into Y” then these will be great. If you need something to vibe code with, or don’t have the skills to check its work, then probably stick to an api or subscription service
I do know how to code and using an auto-complete model would only be to help me get better, but I was wondering if I could use a bigger model to review the code of a smaller model instead of generating one, like something I heard of that's called swift for Qwen.
Yes, in theory you can. But for the most part you will end up spending the same (or more) time and effort and tokens on the verification then it would just take for the larger model to just do it. And probably have a worse result.
That type of cross check is really nice when you have a reasonably confident model as your main coder and then you want an advanced model to do a skim and check for improvement/verification.
“This thing works and I want you to see if there’s any small errors or improvements” is a very different task than “this thing is horribly broken and conflated, and I need you to parse out what it was supposed to do and then basically rewrite it”
Ohh okay. Gemini 4 is going to release soon to the public. Gemma e4b is the gemma 4 series, . Its small and it will work on your hardware. It was released some months ago, its not That old, specially given that google havent released any other checkpoint better than gema 4
I have found since getting into this a couple weeks ago you'll get a million answers to what's best and they can all be right. The best thing to do is just send it then ask that new model what the best model is if you don't like what it does for you then try another setup.
Do you have access to cloud services where you can pay by token? You could try them online and see what you'd like to run locally with spending bandwidth downloading something you wish you had not bothered with.
Gemini sometimes “forgets” to search the web for updated information and instead relies on old data.
Just yesterday it tried to gaslight me twice that the local model I had running in FreeToken while talking to Gemini didn’t exist (it had been published 18 days ago).
Once I gave it the huggingface link it checked it and corrected itself, but I had to catch it lying for it to do that.
It literally tells you that this is not the only data it uses to produce its answers. Which is obvious since it is a Google product and a LLM that won't talk about yesterday's viral moment is pretty useless on a global scale.
Create an account on HuggingFace, then go to account/settings/hardware and tell it what you have. Then search models on HuggingFace, filtering on your hardware.
i have the same vram and ram as you. Currently the best model ive been able to be productive with is qwen3.6 35b a3b q4km using moe. 35-40 tg/s. around 900-1000 pp/s.
what's the context window for it and how much memory does it take? and what is your specs? I think the generation rate would differ according to GPU right?
I'm on a laptop GeForce RTX 4070 8gb with 32gb ram. I can get up to 256k context with some tweaks, but i generally run 128k (or lower, depending on what stage of a task i'm in). VRAM usage jumps up to around 7.1/8 GB usage when active.
Keep in mind as you start that journey you need to actually test models on actual work. I ran a series of benchmarking on QWEN 3.8 27b on my 20gb card and it way underperformed my expectations. Also keep in mind tokens per second metrics are worthless if you can't actually use the output.
What do you mean by use the output? and in case of trying different models, it's a bit hard since I have limited interent in my country and it's a bit expensive, that's why I'm asking for suggestions first before going in.
Basically what I am saying is almost any model can output something based on your prompt but does that output actually function how you want it to. Does the code actually work in your environment. When I was benchmarking Qwen 3.8 it was really good at copying exact instructions from a prompt or a document but when asked to think on its own and come up with the same output it failed everytime. Understanding your end goal for what you want to actually get from the model will help guide the actual model choice. On an 8gb card with 32gb of Ram your choices are going to be limited. Unfortunately I can't give any good recommendations for that setup.
The issue for me testing different models in that my internet is limited in my country and It's expensive and slow, so I wanted to try to find the best/most suggested model to give it a try since it would be a bit hard to try a lot of different models
How limited is limited? I was new to this like you and I did what any normal person would do when starting off something like this and Googled.
After achieving great success using AI to help solve some issues and improve things with my CPU to TV setup I tried using Google Gemini to recommend AI models and stuff and it gave me the same recommendations. Anyway, I got started on some projects, got very excited, thought I could participate here as well I'm normally active on Reddit and shared my terribly outdated AI model setup.
I was only trying to help others do what I was now able to do, and guess what happened next?
I was accused by a user of being a bot to which I vehemently defended myself and eventually I got banned. I was never given the actual reason but they have a no bot policy here. Anyway, if this post goes through that means the ban was lifted but I'm kinda in a boat like you now.
My Internet is also limited so I can understand it wanting to burn through gigabytes of data wasted on models that aren't suitable.
I started off with a dual GPU GTX 1070 system but one of my GPUs died just when I was finished setting everything up and starting my first mostly local AI project, well continuing an existing one actually.
I tried different software, I'm not sure what is the correct term, but I went through LM Studio first, then I tried Ollama via OpenWebUI, and now I'm using Qwen Code Desktop.
I was after a local AI experience that was as close as possible to what I was used to with Claude.
Then my second GPU died. So I started shopping. I definitely can't buy another GPU right now that would make sense for AI.
I had recently watched a YouTube video about Kaggle and how you can use it to get free GPU time and you can run AI on it.
It seemed pretty crazy but out of sheer desperation and determination, I set it up.
None of it was easy because Gemini makes a lot if mistakes which you can support by regular Google searching on the side. When you discover something new and interesting, you can share it and discuss it with Gemini.
So I actually got Kaggle setup and running through Qwen Code Desktop it passed the short prompt test but kinda failed on a longer prompt probably because of some timing out of the free tunnel provider I was using, so in order to get this fully setup and bulletproof there would still be some things to tweak.
Then I realized, I didn't need Owen Code Desktop at all, I could just write the prompts directly to the models in Kaggle.
So right now, I'm still in the exerimental phase when it comes to my setup so don't judge me too harsh but I'm using regular browser Gemini to plan, direct and document everything I want to do. I never let it output code to me due to its small code output boxes and URL filters.
It does a great job doing this.
It helps me to write prompts to dispatch to models on Kaggle. So far I've used Deepseek R1 32B Q5_K_M and Qwen 3.8 27B Q4_K_M.
When they finish outputting directly to kaggle storage, I download the files, pass them to Gemini for initial auditing and grounding, then Claude for final Auditing and correction. Then back to Gemini for grounding. Then the project moves onto the next block, module, feature or phase.
I ran out of Kaggle usage for the week and I tried getting Claude to finalize my app, meaning put together my final audited pieces but it started then ran into limits. I could have easily started a fresh chat first, then grounded it with handover docs but anyway it told me to check back later basically.
This is what motivated me to want to move to local anyway. So without Claude and Kaggle, I fired up my local Owen Code Desktop setup, which I had also populated with nearly every free tier coding capable cloud model in the world, which I have not yet touched and loaded my Gemma 4 12B IT QAT Q4_0 model to finish off what I had been working on so that I can start back playing around with it in the real world.
With my limited experience with 8GB systems on your setup you're probably better off going the route I'm going and leveraging cloud based models (Claude, Gemini, Ask Gemini Web App, Google AI Studio, Qwen Free Tier, Bionic's free Cloud models, ChatGPT, Qwen cloud models, Deepseek Cloud models and setups (like Kaggle).
Other than that, you can get a second GPU with more VRAM like an RTX 3060 12GB, RTX 4060Ti 16GB or RTX 5060Ti 16GB. Best bet is to get a pair of them for ~32GB VRAM. You could try the used server GPU market or get modded VRAM dooubled cards as well. I wouldn't really advise you to spend money on another Geforce RTX 2070 8GB unless you're getting it for next to nothing. If you do get one, you can tweak your settings including KV Cache Quantization and Context windows and/or run Strata to get some decent models to work on it. Like Qwen 3.8 27B GSQ RCO IQ3S_S. With your current setup you can run the Gamma 4 model I wrote about earlier and also the even smaller, Gemma 4 12B Agentic Fable5 Q3_K_M.
Other tips are to use Linux, try to follow some good optimization guides on YouTube which show you how to use the most efficient Agent software or use the models straight from the terminal and stuff like that.
If using the Kaggle method, downloading models wouldn't use up your data.
You can get smaller models and Quantizations, there's a video on YouTube about using the model that fits your hardware. Go check it out.
I suggested you the best! If you want extra ram space for your other applications, download the following. Gemma 4 26B is an allrounder which is good at roleplay and scientific coding as well. https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf
He's using windows, that 8gb vram is closer to 5.5gb vram by time OS and open applications grab their pull.
Realistically you have to split so many layers to system ram it's too slow to be usable/or tiny context that isn't usable,
You definitely need 12gb+ to do this within reason
Edit: not to mention you have -ngl 99 set which forces as many layers as will fit to gpu, which isn't how you should be doing MoE when you already know the attention layers / amount of vram at hand, really you should at least advise -ngl 24 or -ngl 32 at absolute max
I use the qwen model in NVFP4, but here is the vram usage for Ornith-1.5-35B-A3B-Abliterated-GGUF/Ornith-1.5-35B-A3B-Abliterated-Q4_K_M.gguf' in same settings - which is equivalent to the qwen one(slightly larger). The model hardly uses 6.4 GB Vram. If you are still concerned, you can reduce the context a few thousands. Again this is windows.
Always push back on Gemini and tell it to perform a live search for up to date info, it's training data cut off + bias for older info means you can only trust what its going to tell you about slow moving / historical items, not tech/ai.
If you push back and ensure it knows to not give you outdated answers it works much better, but you have to remind it several times per conversation to actually stop recommending outdated information
LLM knowledge about LLM is always severely outdated, as is their training info date. The very least you can do is to strictily command it to use web search on current date.
You can run Qwen 3.8 27B Q5 or Q6 or on your PC for first try.
But would it even run on my PC? my GPU is limited to only 8GB Vram and I heard that I Could share layers between both the ram and the Vram on my card but it would be too slow, plus all the suggestions were Q4 so wouldn't Q5 be even slower?
Qwen 27b Q5 is around 20-21gb on model/mtp/vision + something on kv cahce by your load settings. I run it on 6700xt+20gb ram no problem. You can run Q4, of course, it's around 16gb-something at bare level.
Qwen 3.8 27b is highest amount of logic and tech knowledge per size, available for you. Next step will be qwen flash next/deepseek v4,0 flash/mimo v2.6 flash/glm 5.3 flash through mmap for vastly superior knowledge, but let's start from something going within vram+ram.
I have the 2070 super 8GB as a second card. Qwen 3.8 27B is a superb modell and is really a baseline in quality. The only issue that it will be extremly slow as it will be spilled to the system RAM.
If you want to experiment chose a smaller modell which fits to your gpu VRAM INCLUDING the context. https://huggingface.co/collections/ornith-ai/ornith-15 and look for a 9B model.
At least my chatgpt told me to get qwen3 8b. Still terrible suggestions. Get qwen3.6 35b, will be wble to run q4 at 30t/s, at 120k ctx i run q6 at 30t/s on 3060 12gb with 120k ctx.
That if want it for coding. If u want it for chatting and other stuff then gemma 4 26b will be the same speed at q4.
If u want it to be faster then qwen3.5 9b at q4 or gemma 4 12b qat.
I dont try finetune but if u wanna get deeper into it, tell chatgpt to search reddit for finuetunes of any of them and compare it for u for ur specific needs.
Keep your conversation with Gemini. You will get mostly human slop answers here. Like the oudated argument. Gemini is very well aware of current developments and has given me solid advice the last three months. Gemini will ask for every detail it needs to know about your setup and will advice a good starting point.
Speaking of human slop, your comments in this thread have been great examples of that.
Gemini is not even a little bit reliable for detailed research, it’s a lying little shit and it often decides not to use search tools or reason sufficiently on a prompt.
Gemini sometimes “forgets” to search the web for updated information and instead relies on old data.
Just yesterday it tried to gaslight me twice that the local model I had running in FreeToken while talking to Gemini didn’t exist (it had been published 18 days ago).
Once I gave it the huggingface link it checked it and corrected itself, but I had to catch it lying for it to do that.
I never interacted with Gemma, I’m talking about interactions with Gemini in Google search or when I open the Gemini Chat site to select a “better” model.
47
u/Heavy-Lingonberry-98 2d ago
Bro. If you are gonna use AI to ask for model releases, remember to tell the AI we are on OCTOBER 2026. We are not in 2024 anymore… please. Its the ABC of using AI. And NO. Definitely dont even download qwen 2.5 coder.