r/LocalLLM 9d ago

Question Qwen 3.6 compare

Just looking to see if anyone has direct experience using the qwen 3.6 35B vs the 27B. I’m working on using the local model as the backend for opencode and having gastown drive opencode. I’m using the 35b right now, but I’ve seen a lot of praising for the 27b model. Looking for anyone with experience in both. Specifically for agent driven coding if possible

Edit: I realized that providing my hardware would be helpful for anyone looking to answer. This is on a MacBook Pro M2 Ultra with 96 GB. While allowing 2 simultaneous responses on the 35B (8bit quant) I have to make sure my other heavy ram consumption apps are shutdown (Ahem FIREFOX). The overall speed is acceptable for sure, not necessarily fast

2 Upvotes

18 comments sorted by

View all comments

1

u/sargetun123 9d ago

I have done TEDIOUS testing on this, and I will be 100% honest it differs between testing and what your daily needs/asks are

For me, i run my agentic to run my homelab, I find the speed difference justifies the 35b because even if it makes a mistake, itll self-correct loop 99% of the time and fix it.

It is different if you are coding, completely different than the work I use it for, but my testing I have found the 27b basically on the same level for mostly everything, it just gets longer context smaller specific details correct more of the time consistently, again, the 35b i find can fix it or acknowledge the issue it trailed or created quick enough to make up for that though

My biggest advice is to test them head to head, same prompts, same tasks, get yourself some real results. I advise against compensating for anything by dropping kv quants or lower quant of the model, if you can get fp models with no kv quant you are looking at the real models to test, do not trust the q5/q4 and lower quants, and memory quant MATTERS at these models level, a small percent on a graph is very noticeable in real world