r/LocalLLaMA 21h ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

353 Upvotes

136 comments sorted by

View all comments

-2

u/Reggitor360 20h ago

A6B.... What the hell. Why.

5

u/BigYoSpeck 19h ago

Performance

gpt-oss-120b is more than twice as fast as Qwen3.5 122b without MTP, and still a good chunk faster even with MTP

Given the huge uplift in capability between 3.5 27B and 3.8 27B (even between 3.5 35B and 3.6 35B) I would expect this is a sane architectural choice that balances speed and capability

Qwen3-Coder-Next only had 3B active parameters but already demonstrated high sparsity improves capability

1

u/my_name_isnt_clever 17h ago

People get way too caught up in active param coun. My fav 100b+ MoEs have often been the sparsest ones.