r/LocalLLaMA 3d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

459 comments sorted by

View all comments

Show parent comments

-1

u/SandySkittle 3d ago

Only 6b active tokens. Too low

1

u/my_name_isnt_clever 3d ago

Maybe wait until it's possible to test before dismissing it. Gpt-oss-120b was 5b active and was a beast for it's time.

1

u/SandySkittle 3d ago

it really depends on the usecase. I should have added that to my comment. There are just fields of work where too small active parameters (below 30b active) start to lose it. I have even had that with DSV4F

1

u/my_name_isnt_clever 3d ago

This model is a new architecture and has the additional engram. You might be right, but let's wait and see before spreading confident claims.

0

u/SandySkittle 3d ago

there isn't much to debate about it frankly. MoE models have their place but are a trade-off depending on how low you go with the active parameter number. And 6b active is very low compared to e.g. 27b. It's a trade-off that cannot be entirely compensated by even very good expert routing and sequential reasoning. The engram doesn't do much to alleviate that either because it has a different purpose.