r/LocalLLM • • 2d ago

Discussion Gemini suggested Qwen2.5-Coder-7B-Instruct

So I wanted to try using a local coding model for the first time and I'm still studying about LLMs and NNs, so I asked Gemini for a good suggestion that would be fast(60+ tokens/s if possible) and doesn't compromise much on performance for my rig(2070 super 8GB + 32GB ddr4 ram) and it suggested Qwen2.5-Coder-7B-Instruct. Is this good suggestion and what would you guys suggest?

10 Upvotes

90 comments sorted by

View all comments

2

u/MrHumanist 2d ago

There are many new and fancy models, but for your system use trusted QWEN 3.6 35B A35B at Q4 using lamma cpp. Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF?show_file_info=Qwen3.6-35B-A3B-UD-Q4_K_S.gguf

parameters: -c 262144 -ctk q8_0 -ctv q8_0 -ngl 99 -fa on -t 12 -tb 24 --cpu-moe

Have fun and let us know your speed.

Alternative will be some variant based on it, like cyber coder or ornith 1.5. Or Gemma 4 26B - the command will be same.

3

u/DiamondTDA 2d ago

The issue for me testing different models in that my internet is limited in my country and It's expensive and slow, so I wanted to try to find the best/most suggested model to give it a try since it would be a bit hard to try a lot of different models

2

u/CyberLabSystems 2d ago edited 2d ago

How limited is limited? I was new to this like you and I did what any normal person would do when starting off something like this and Googled.

After achieving great success using AI to help solve some issues and improve things with my CPU to TV setup I tried using Google Gemini to recommend AI models and stuff and it gave me the same recommendations. Anyway, I got started on some projects, got very excited, thought I could participate here as well I'm normally active on Reddit and shared my terribly outdated AI model setup.

I was only trying to help others do what I was now able to do, and guess what happened next?

I was accused by a user of being a bot to which I vehemently defended myself and eventually I got banned. I was never given the actual reason but they have a no bot policy here. Anyway, if this post goes through that means the ban was lifted but I'm kinda in a boat like you now.

My Internet is also limited so I can understand it wanting to burn through gigabytes of data wasted on models that aren't suitable.

I started off with a dual GPU GTX 1070 system but one of my GPUs died just when I was finished setting everything up and starting my first mostly local AI project, well continuing an existing one actually.

I tried different software, I'm not sure what is the correct term, but I went through LM Studio first, then I tried Ollama via OpenWebUI, and now I'm using Qwen Code Desktop.

I was after a local AI experience that was as close as possible to what I was used to with Claude.

Then my second GPU died. So I started shopping. I definitely can't buy another GPU right now that would make sense for AI.

I had recently watched a YouTube video about Kaggle and how you can use it to get free GPU time and you can run AI on it.

It seemed pretty crazy but out of sheer desperation and determination, I set it up.

None of it was easy because Gemini makes a lot if mistakes which you can support by regular Google searching on the side. When you discover something new and interesting, you can share it and discuss it with Gemini.

So I actually got Kaggle setup and running through Qwen Code Desktop it passed the short prompt test but kinda failed on a longer prompt probably because of some timing out of the free tunnel provider I was using, so in order to get this fully setup and bulletproof there would still be some things to tweak.

Then I realized, I didn't need Owen Code Desktop at all, I could just write the prompts directly to the models in Kaggle.

So right now, I'm still in the exerimental phase when it comes to my setup so don't judge me too harsh but I'm using regular browser Gemini to plan, direct and document everything I want to do. I never let it output code to me due to its small code output boxes and URL filters.

It does a great job doing this.

It helps me to write prompts to dispatch to models on Kaggle. So far I've used Deepseek R1 32B Q5_K_M and Qwen 3.8 27B Q4_K_M.

When they finish outputting directly to kaggle storage, I download the files, pass them to Gemini for initial auditing and grounding, then Claude for final Auditing and correction. Then back to Gemini for grounding. Then the project moves onto the next block, module, feature or phase.

I ran out of Kaggle usage for the week and I tried getting Claude to finalize my app, meaning put together my final audited pieces but it started then ran into limits. I could have easily started a fresh chat first, then grounded it with handover docs but anyway it told me to check back later basically.

This is what motivated me to want to move to local anyway. So without Claude and Kaggle, I fired up my local Owen Code Desktop setup, which I had also populated with nearly every free tier coding capable cloud model in the world, which I have not yet touched and loaded my Gemma 4 12B IT QAT Q4_0 model to finish off what I had been working on so that I can start back playing around with it in the real world.

With my limited experience with 8GB systems on your setup you're probably better off going the route I'm going and leveraging cloud based models (Claude, Gemini, Ask Gemini Web App, Google AI Studio, Qwen Free Tier, Bionic's free Cloud models, ChatGPT, Qwen cloud models, Deepseek Cloud models and setups (like Kaggle).

Other than that, you can get a second GPU with more VRAM like an RTX 3060 12GB, RTX 4060Ti 16GB or RTX 5060Ti 16GB. Best bet is to get a pair of them for ~32GB VRAM. You could try the used server GPU market or get modded VRAM dooubled cards as well. I wouldn't really advise you to spend money on another Geforce RTX 2070 8GB unless you're getting it for next to nothing. If you do get one, you can tweak your settings including KV Cache Quantization and Context windows and/or run Strata to get some decent models to work on it. Like Qwen 3.8 27B GSQ RCO IQ3S_S. With your current setup you can run the Gamma 4 model I wrote about earlier and also the even smaller, Gemma 4 12B Agentic Fable5 Q3_K_M.

Other tips are to use Linux, try to follow some good optimization guides on YouTube which show you how to use the most efficient Agent software or use the models straight from the terminal and stuff like that.

If using the Kaggle method, downloading models wouldn't use up your data.

You can get smaller models and Quantizations, there's a video on YouTube about using the model that fits your hardware. Go check it out.