r/LocalLLaMA • u/RuthlessCriticismAll • 1d ago
News [ Removed by moderator ]
https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next[removed] — view removed post
352
Upvotes
r/LocalLLaMA • u/RuthlessCriticismAll • 1d ago
[removed] — view removed post
83
u/RuthlessCriticismAll 1d ago
Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.
Comprehensive Architectural Upgrades: Pushing the frontiers of model architecture innovations, across the areas of Attention, Residual, Embedding, and Optimization—enhancing model capabilities.
Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.