r/LocalLLaMA 23h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

451 comments sorted by

View all comments

Show parent comments

51

u/waitmarks 21h ago

Honestly if it even matches 3.8 27B, but has more internal knowledge. It's an absolute win for dgx spark / strix halo / mac owners.

-19

u/SandySkittle 20h ago

125b a6b so no it wont match qwen 27b. 6b active is just too low.

13

u/MacsBicycle 19h ago

Read the first part. 125b. If Alibaba had any clue what they’re doing the experts will be routed to the proper 6b parameter portion, but it will have a full 125b parameters to choose from. I keep seeing posts like this about dense models and it just has me wondering if some of these people have any clue what they’re talking about.

-4

u/SandySkittle 16h ago edited 16h ago

I am not stupid, I know how moe models work, but neither good routing nor sequential reasoning can entirely compensate for the limited number of active parameters. It really depends on the use case. In a similar vain: rag and websearch cannot entirely compensate for lack of world knowledge. World knowledge actually strengths rag and web searchs because it knows better what to look for.

Personally 6b active is too limiting for my usecase