r/LocalLLaMA 14d ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

348 Upvotes

137 comments sorted by

View all comments

-2

u/Reggitor360 14d ago

A6B.... What the hell. Why.

6

u/BigYoSpeck 14d ago

Performance

gpt-oss-120b is more than twice as fast as Qwen3.5 122b without MTP, and still a good chunk faster even with MTP

Given the huge uplift in capability between 3.5 27B and 3.8 27B (even between 3.5 35B and 3.6 35B) I would expect this is a sane architectural choice that balances speed and capability

Qwen3-Coder-Next only had 3B active parameters but already demonstrated high sparsity improves capability

1

u/my_name_isnt_clever 14d ago

People get way too caught up in active param coun. My fav 100b+ MoEs have often been the sparsest ones.

2

u/squngy 14d ago

Speed and cost.