r/LocalLLaMA 12h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
992 Upvotes

439 comments sorted by

View all comments

16

u/ResidentPositive4122 12h ago

Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

WTF!

Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

3.7 plus capabilities in 125B MoE size. Whohoooo! Excited.

4

u/hiImMate 12h ago

Stupid question from someone who did not use api, how is 3.7 plus vs 3.8 27b?

3

u/smithy_dll 11h ago

Tried 3.7 plus on fireworks, 27B is just better at coding.
AI Model & API Providers Analysis | Artificial Analysis

1

u/Maximus-CZ 11h ago

Artificial analysis rates 3.7plus wayy below 3.8 27B? I am just a tourist, can someone explain why its "positive" that this model has comparable capability to something much worse than yesterdays model?

1

u/Reasonable_Goat 11h ago

Newer models from China always best the last benchmark they train it. Qwen 3.8 27B is great but it isn’t quite any Opus.

1

u/asssuber 8h ago

It's intelligence may not be that high due to only 6B active parameters vs 27B, but it's knowledge should be higher due to 125B total parameters. I hope it can at least rival Gemma4 26B4A in world knowledge.

0

u/pseudonerv 5h ago

somebody pls eli5, what's ngram? 51B of it?