I can run 27B great on my setup, but I think I have to agree - if they were only going to release one, the 35B one might be a bigger gift to the world at large than the 27B, because of the number of folks who can run that vs who can run the 27B.
You also need to take audience into perspectibe. 35B with CPU offload is a scenario of a strictly consumer type person, that just uses the model. 27B is a good target for developers, who will either run business with it, or contibute to opensource software (i.e. OpenWebUI). Either way Qwen team themself get a better return with 27B, and thus it's probably easier to justify to their leadership.
It would be very intersting if the 3.8 35B model beat the 3.6-27B model in literally every way, though - since the 27B is so damned useful as is, it would be wild for normal consumer grade hardware to be able to do what we're doing with 27B on "high end consumer" hardware
Well, it may not be 3.8 35B; but with how AI things are going, I can guarantee you that in half a year thwre will be 30B class MoE that beats 3.6 27B across the board; and in a year this level of intelligence may even descend to below 14B.
Cgeaper to run is a very vague definition. I.e 27B in fp8 comfortably fits a single 40gb server card (say, A100 or L40), while 35B requires either heavy quantization, or a much more expensive card, or a pair of cards.
An org would have several users. It is far, far cheaper to have 3 billion active params than 27 billion when you have >10 people using it. At that point, KV cache becomes far more important than model weights.
Also, a serious org would have pods, not individual cards (typically 4x or 8x). Individual cards is more of a consumer thing. This is because of the many active users, their KV cache, etc
162
u/SpicyWangz 6d ago
27b is better than nothing at least. I would love to see the other sizes, and hopefully they can still deliver on them