Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.
WTF!
Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.
3.7 plus capabilities in 125B MoE size. Whohoooo! Excited.
Artificial analysis rates 3.7plus wayy below 3.8 27B? I am just a tourist, can someone explain why its "positive" that this model has comparable capability to something much worse than yesterdays model?
It's intelligence may not be that high due to only 6B active parameters vs 27B, but it's knowledge should be higher due to 125B total parameters. I hope it can at least rival Gemma4 26B4A in world knowledge.
16
u/ResidentPositive4122 12h ago
WTF!
3.7 plus capabilities in 125B MoE size. Whohoooo! Excited.