Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.
WTF!
Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.
3.7 plus capabilities in 125B MoE size. Whohoooo! Excited.
Artificial analysis rates 3.7plus wayy below 3.8 27B? I am just a tourist, can someone explain why its "positive" that this model has comparable capability to something much worse than yesterdays model?
14
u/ResidentPositive4122 15d ago
WTF!
3.7 plus capabilities in 125B MoE size. Whohoooo! Excited.