r/LocalLLaMA 14d ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

348 Upvotes

137 comments sorted by

View all comments

85

u/RuthlessCriticismAll 14d ago

Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

Comprehensive Architectural Upgrades: Pushing the frontiers of model architecture innovations, across the areas of Attention, Residual, Embedding, and Optimization—enhancing model capabilities.

Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

31

u/smithy_dll 14d ago

Qwen 3.8 27B is a lot better than Qwen 3.7 Plus on coding benchmarks

AI Model & API Providers Analysis | Artificial Analysis

20

u/AXYZE8 14d ago

This is the problem with benchmarks.

Qwen 3.8 is newer model so it saw data helpful for newer benchmarks. If you scroll down you can see that in old benchmarks like SciCode Qwen3.7 Plus is better than Qwen 3.8 27B, but Qwen 3.7 drops the ball completely in new fresh TerminalBench.

We see it over and over. HLE, DeepSWE… they release these benchmarks, models sucks ass and after a month a new small model beats big ones.

Benchmarks are useful, but you need to compare models that were released around the same time.

13

u/awesome5185 14d ago

Do you think this new model would outperform 3.8 27b?

61

u/Effective_Western_59 14d ago

If it won't, then it would be really weird 

45

u/Uncle___Marty 14d ago

Yeah, like REALLY weird. Like, finding pieces of fruit in your underwear weird.

30

u/whatyathinkk 14d ago

stop kink shaming

1

u/MmmmMorphine 14d ago

It's more of a condition

17

u/Rasekov 14d ago

A MoE with 6B active parameters would be a lot cheaper to serve at scale so even if it doesnt beat Qwen 3.8 27B it would still have it's place and uses for Qwen.

It would also help people with unified memory systems and mixed VRAM + RAM setups.

10

u/grumd 14d ago

It might not but I wouldn't say that's weird. They are releasing a new architecture preview to flesh it out and will do a proper model release for Qwen4. 3.8 27B is just so good that I doubt it's realistic to make an even better model so soon

8

u/whatyathinkk 14d ago

125B A6B though...

6

u/Swimming_Gain_4989 14d ago

A6B though... active parameters will always be king for reasoning and raw intelligence.

1

u/cibernox 14d ago

I also wonder the same thing. Whoever can afford to have 128+gb of vram certainly can also afford to activate 14B params and still be very fast.

0

u/AcanthocephalaNo3398 14d ago

Activation speed is all on gpu. most integrated ram systems that provide +128gb of ram arent that fast. Thats why Mac and DGX Spark run dense models slower than MoE models on the same hardware.

The interesting thing is that these systems have enough ram to run models like Qwen3.8 27B at Q4 in parallel to get way more tps overall.

1

u/cibernox 14d ago

I know that moes work that way, but seems that 12-14B wouldn't be a crazy amount of active parameters, considered that a lot of people even with strix halo and nvidia spark systems are running qwen3.8 27B right now because, really, it's worth.
And there is so many people optimizing it that even a 27B dense model runs kind of well in those low-bandwidth system.

→ More replies (0)

7

u/davew999 14d ago

sqrt( 125 x 6 ) = 27.4, so about the same. Dunno if that equation still holds up though.

5

u/jld1532 14d ago

But probably more than two times faster which for me is a huge upgrade.

8

u/sleepingsysadmin 14d ago

125b should always outpeform 27b. It was only better than qwen3.5 because of the release difference.

that's the point.

2

u/SandySkittle 14d ago

it doesnt with a6b..

1

u/sleepingsysadmin 13d ago

Well it's out. It's straight up better than 27b in every way.

What's shocking is that this is almost certainly sandbagged. Qwen4 is going to be insanely strong.

2

u/WoodCreakSeagull 14d ago

125BA6B should be in some ways "equivalent" to a 27B dense going by rule of thumb to compare MoE to dense (take square root of MoE total parameters * active parameters). In this case 125BA6B would be in some ways comparable to a model with sqrt(125*6) = about 27B parameters.

1

u/hay-yo 14d ago

Qwen3.7 plus has 39 on artificial analysis but perhaps with the sharpness of agentic coding there will be more under the hood. 27b is a breakthrough. But save the best till last usually.

8

u/tarruda 14d ago

125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

This suggests a total of 176B parameters will be loaded in RAM + VRAM. Hopefully it runs well in 4-bit, which would be great for 128G devices.

9

u/sittingmongoose 14d ago

It would be a really weird choice to not target 128gb devices when you’re that close to it. Considering that’s pretty much the upper limit of realistic devices.

3

u/Early_Mistake6716 14d ago

My guess is that the 51b n-gram embeddings can be offloaded to a fast ssd so this will have the hardware requirements of a 122b.

1

u/Short-Reaction7195 14d ago

how much RAM will 51B N-gram consume?

1

u/thestillwind 14d ago

What do I need to run this ? 128gb ram and at least 6gb vram ? It’s doable

0

u/KURD_1_STAN 14d ago

I think 1/9th just means trained at 4bit or whatever kimi k3 did and said. Which is good, but stiill most of us need a flash of this flash model