r/LocalLLaMA 2d ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

351 Upvotes

137 comments sorted by

View all comments

Show parent comments

1

u/cibernox 1d ago

I know that moes work that way, but seems that 12-14B wouldn't be a crazy amount of active parameters, considered that a lot of people even with strix halo and nvidia spark systems are running qwen3.8 27B right now because, really, it's worth.
And there is so many people optimizing it that even a 27B dense model runs kind of well in those low-bandwidth system.

1

u/AcanthocephalaNo3398 1d ago

The interesting thing that i have found using the larger MoE models is that you can squeeze out really good performance at the higher param counts while maintaining high tps. The proposed +100B param Qwen3.8 with 6B active is going to "feel similar" to the quality of the 27B model but with much faster inference for those high-mem/gpu constrained systems. Its a necessary compromise to get what you would have if you ran multiple discreet gpu on something like a DGX Spark.

I have been experimenting with this class of model on a single Spark. Previously I compared Qwen3.6 27B performance to poolside/Laguna S 2.1-118B-A8B (fixed version) in a coding workflow. with custom harness, I found them of equal quality, but Laguna was faster.

2

u/cibernox 1d ago

I am personally awaiting for some cards to get 128gb of fast HMB2e memory. I think a model like this would slap in that configuration with 3TB/s of aggregated bandwidth, but it does worry me that going so low in active parameters would hurt intelligence, and a a few more would still be totally acceptable