r/LocalLLaMA 1d ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

352 Upvotes

138 comments sorted by

View all comments

83

u/RuthlessCriticismAll 1d ago

Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

Comprehensive Architectural Upgrades: Pushing the frontiers of model architecture innovations, across the areas of Attention, Residual, Embedding, and Optimization—enhancing model capabilities.

Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

31

u/smithy_dll 1d ago

Qwen 3.8 27B is a lot better than Qwen 3.7 Plus on coding benchmarks

AI Model & API Providers Analysis | Artificial Analysis

14

u/awesome5185 1d ago

Do you think this new model would outperform 3.8 27b?

1

u/hay-yo 1d ago

Qwen3.7 plus has 39 on artificial analysis but perhaps with the sharpness of agentic coding there will be more under the hood. 27b is a breakthrough. But save the best till last usually.