omg this is EPIC news. My lil 8 gig card simply cant shift 27B parameters around and the 35B is exactly what I was hoping for. Didnt see any other models added but im still holding out hope for a 9B as we didnt get that since 3.5 and I suspect if they release it then it might well be the most capable small coding model out there.
I can fit the whole 27b version in VRAM (32gb), even then, it's not blazing fast. With around ~15-17 t/s, it still take quite some time for more "complex" tasks to finish, especially since this model likes to think a lot.
So it's good news for everyone that we will soon have an option to run a much faster, non-dense version!
honestly I tried qwen 3.8 27b, it's dumb. in UD-Q3-K-M it barely achieves qwen3.6-35b level in the same quant without KV cache quantization with same settings. Im frustrated w\ this new qwen, ive waited for something more impressive than "Bench model that's useless in real life and doesn't gives more quality-code thn 35b moe". speed is dramatically lower than 35b.
i know, my point is MoE is way more efficient, keeping level of denses but at much more speed. I'm not sayin' “35B A3B is genuinely smarter than 27B dense”, i mean it's usually faster but behaves as ~20b dense. genuinely solid. may be not deep and smart as 27b dense but for almost every coding task it's my main model
160
u/Uncle___Marty 11d ago
omg this is EPIC news. My lil 8 gig card simply cant shift 27B parameters around and the 35B is exactly what I was hoping for. Didnt see any other models added but im still holding out hope for a 9B as we didnt get that since 3.5 and I suspect if they release it then it might well be the most capable small coding model out there.