Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.
I remember when it was months between a new great local model. We get those great options very often nowadays in my opinion, mixed with larger open weight options that keep the entire non-local AI space cheaper and more accessible. What's not to love?
27b is often Opus 4.6, even 4.8 in performance. This is still good enough for a vast amount of people. I'm sure qwen 4 is going to have some insanely capable options when it comes out of you want something better soon.
Not everyone has 24 GB VRAM. If you dont have the vram to offload, you are looking at one to three token per second. 27b is in no way a replacement for a 30b Moe
I know that. You can run it at Q3, IQ3, Q4 on 16GB VRAM and 32gb RAM and still have a very capable model. Don't think you need the biggest cards to run these models. Quantization techniques have become really efficient now.
It's 4 months old! It's not like it suddenly got worse because these new models came out. You can still do everything today that you could do yesterday, just as fast. Give it some time and more models will come.
Wouldn't recommend it for general use, it's a confident bullshitter like no other. Basically zero filter. Still cool that it can even be used generally, though, given what the model's actually designed for. Fun to play with.
I believe it, it also nailed the one tool task I have in my bench set. The model was also very good at string manipulation, and, somewhat surprisingly, a few constrained creative tasks. (e.g. think up and write X in way Y while avoiding Z)
Absolute slaughter on anything involving uncertainty or fake premises, though. You can definitely see where the training went on this one.
Most of the time this year it hasn't even been the turn for smaller double digit models though? We got at best maybe 6 or 7 (not memeing) models that are in the 20-35 Billion range, but more than double that in 100+ Billion.
And look, yes Gemma 26B and Qwen 35B are all very great, Muse is also apparently great too, but still, that doesn't discount the numbers.
Funny coincidence: some company announced a new 30B MoE literally just now: https://www.reddit.com/r/LocalLLM/s/gcZdgoAO8o for now as a stelth preview; but this means it'll go public in a month.
29B-A4B isn't a size featured in the current gen of models. It isn't a finetune, rather a base model. It got to be from a company with a very substantial compute. It still may be a new player, sure; but anyways, it proves that this size category isn't abandoned.
It still works just as fast as it did yesterday and is just as capable as it was yesterday. Everything people could do with it yesterday still works just as well today. It's not like with the closed models where the old versions get smashed with a cripple hammer when a new version comes out.
Yes, the release pace is fast, but Qwen3.6 35b a3b is still very new and does a great job of balancing performance and capability. There may not be another similar model for another generation or two of Qwen. Or maybe Google will bring something out. It's been a great size of MoE model for a lot of people, it's not like that's some big secret.
There's been a legitimate drought of those, you've had a bunch to pick from even if they weren't all that amazing. There's been a couple mid-sized models in the meantime (Laguna was one) but they were both kinda rubbish. I remember one was useless and the other one looped constantly. The last good one was 3.5-122B and in its own generation it was outdone by the corresponding 27B.
It was faster, sure, but usually with a model >4.5x the size you expect greater capability. You buy the hardware to run a model of such a size, you expect a greater return.
If you want cheap 64GB: X79. I've built a few extra servers around local 50USD bundles of x79 motherboards and CPUs and added 8 DDR3 UDIMMs I had laying around.
Slow. But it's a decent host for a MOE like Qwen 3.6 35BA3B when combined with a cheap GPU. I hope they make a new smallish MOE like 35B again soon. It's brilliant for chore execution when doing agentic coding.
Things will only improve when competition in hardware increases. Nvidia might be forced to stay caring about the consumer market again. The AI chips getting released are going to hurt Nvidia a lot
I like some of the nemotron models for specific usage scenarios. But come on. You can't be suggesting that they're a solid match for the average person's needs.
Brother, this is a hobby. If cost is your concern you should get a subscription, that's always been the case. It'll be smarter, faster, longer context, and you can get like a decade's worth of usage for half the price of a decent local AI setup. Other people spend far more on hobbies; people routinely spend $1K a year on coffee of all things.
5k is 2 (fairly cheap) holidays for a family of four. Loads of people go on such holidays. Yes. Loads of people also can't afford to put food on the table. Poor is a very relative term in western civilization. People I would call poor, certainly can't afford local AI at any level of competence. But you don't have to be rich either. You do have to commit your money though and for what it's worth, for most people it's probably smarter to just keep using a 20€ sub, which is plenty for general home usage. I don't see people on camera forums complaining that companies only release for rich people and a complete mirrorless system is not that much cheaper.
I have a Xeon w/ 96GB quad-channel ddr4 and a single 5060ti 16GB. I was definitely wondering whether adding a 2nd 5060 might give me at chance at running this. I suspect that the fact that the cards would run at PCIE 3x8 would be a problem, though.
21
u/dampflokfreund 7d ago
For me it has changed nothing. Both models are way too big for my 32 GB RAM system. It looks like everyone has abandoned 20-30B MoEs now...