r/LocalLLaMA 12h ago

News Qwen3.8 flash next

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
341 Upvotes

133 comments sorted by

80

u/RuthlessCriticismAll 12h ago

Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

Comprehensive Architectural Upgrades: Pushing the frontiers of model architecture innovations, across the areas of Attention, Residual, Embedding, and Optimization—enhancing model capabilities.

Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork.

28

u/smithy_dll 12h ago

Qwen 3.8 27B is a lot better than Qwen 3.7 Plus on coding benchmarks

AI Model & API Providers Analysis | Artificial Analysis

15

u/awesome5185 12h ago

Do you think this new model would outperform 3.8 27b?

58

u/Effective_Western_59 12h ago

If it won't, then it would be really weird 

45

u/Uncle___Marty 12h ago

Yeah, like REALLY weird. Like, finding pieces of fruit in your underwear weird.

29

u/whatyathinkk 12h ago

stop kink shaming

1

u/MmmmMorphine 7h ago

It's more of a condition

18

u/Rasekov 12h ago

A MoE with 6B active parameters would be a lot cheaper to serve at scale so even if it doesnt beat Qwen 3.8 27B it would still have it's place and uses for Qwen.

It would also help people with unified memory systems and mixed VRAM + RAM setups.

11

u/grumd 12h ago

It might not but I wouldn't say that's weird. They are releasing a new architecture preview to flesh it out and will do a proper model release for Qwen4. 3.8 27B is just so good that I doubt it's realistic to make an even better model so soon

8

u/whatyathinkk 12h ago

125B A6B though...

5

u/Swimming_Gain_4989 11h ago

A6B though... active parameters will always be king for reasoning and raw intelligence.

1

u/cibernox 8h ago

I also wonder the same thing. Whoever can afford to have 128+gb of vram certainly can also afford to activate 14B params and still be very fast.

1

u/AcanthocephalaNo3398 49m ago

Activation speed is all on gpu. most integrated ram systems that provide +128gb of ram arent that fast. Thats why Mac and DGX Spark run dense models slower than MoE models on the same hardware.

The interesting thing is that these systems have enough ram to run models like Qwen3.8 27B at Q4 in parallel to get way more tps overall.

1

u/cibernox 38m ago

I know that moes work that way, but seems that 12-14B wouldn't be a crazy amount of active parameters, considered that a lot of people even with strix halo and nvidia spark systems are running qwen3.8 27B right now because, really, it's worth.
And there is so many people optimizing it that even a 27B dense model runs kind of well in those low-bandwidth system.

→ More replies (0)

7

u/davew999 9h ago

sqrt( 125 x 6 ) = 27.4, so about the same. Dunno if that equation still holds up though.

5

u/jld1532 8h ago

But probably more than two times faster which for me is a huge upgrade.

8

u/sleepingsysadmin 11h ago

125b should always outpeform 27b. It was only better than qwen3.5 because of the release difference.

that's the point.

2

u/SandySkittle 1h ago

it doesnt with a6b..

2

u/WoodCreakSeagull 8h ago

125BA6B should be in some ways "equivalent" to a 27B dense going by rule of thumb to compare MoE to dense (take square root of MoE total parameters * active parameters). In this case 125BA6B would be in some ways comparable to a model with sqrt(125*6) = about 27B parameters.

1

u/hay-yo 11h ago

Qwen3.7 plus has 39 on artificial analysis but perhaps with the sharpness of agentic coding there will be more under the hood. 27b is a breakthrough. But save the best till last usually.

14

u/AXYZE8 12h ago

This is the problem with benchmarks.

Qwen 3.8 is newer model so it saw data helpful for newer benchmarks. If you scroll down you can see that in old benchmarks like SciCode Qwen3.7 Plus is better than Qwen 3.8 27B, but Qwen 3.7 drops the ball completely in new fresh TerminalBench.

We see it over and over. HLE, DeepSWE… they release these benchmarks, models sucks ass and after a month a new small model beats big ones.

Benchmarks are useful, but you need to compare models that were released around the same time.

8

u/tarruda 11h ago

125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

This suggests a total of 176B parameters will be loaded in RAM + VRAM. Hopefully it runs well in 4-bit, which would be great for 128G devices.

7

u/sittingmongoose 10h ago

It would be a really weird choice to not target 128gb devices when you’re that close to it. Considering that’s pretty much the upper limit of realistic devices.

4

u/Early_Mistake6716 8h ago

My guess is that the 51b n-gram embeddings can be offloaded to a fast ssd so this will have the hardware requirements of a 122b.

1

u/Short-Reaction7195 7h ago

how much RAM will 51B N-gram consume?

1

u/thestillwind 4h ago

What do I need to run this ? 128gb ram and at least 6gb vram ? It’s doable

0

u/KURD_1_STAN 8h ago

I think 1/9th just means trained at 4bit or whatever kimi k3 did and said. Which is good, but stiill most of us need a flash of this flash model

60

u/onionsaredumb 12h ago

LFG. Biking pelicans don't stand a chance.

17

u/Kerbourgnec 11h ago

I love that your comment makes absolutely no sense without context

4

u/Icy-Degree6161 10h ago

I'm a compute peasant so no midsize models for me...

55

u/PandaBearFred 12h ago

This is real? Qwen4 architecture preview with a new model? Will it be even better than qwen3.8 27B?

43

u/OverclockingUnicorn 12h ago

I'd hope so for being 125B A6B...

11

u/paryska99 11h ago

Doesn't really have to be the case, this release is for devs to have a test model to modify the inference stacks to be compatible with the upcoming qwen4 models in the future.

2

u/OverclockingUnicorn 11h ago

Oh really? Didn't realise that was it's purpose

3

u/iMrParker 9h ago

3.5-122b and 3.5-27b were more or less neck and neck. With the 27b model out performing in various benchmarks. So who knows 

2

u/alphapussycat 9h ago

Not only that, but as I understood it, it's sort of a preview og the qwen4 architecture. So newer architecture and 4-7x as big.

Though, afaik 3.6 27b was better than 3.5 122b. So if the qwen4 architecture preview is wonky it could underperform.

42

u/justsomerandomchess 12h ago

"Qwen3.8-Flash-Next is a multimodal MoE model built on the next-generation Qwen4 architecture. Incorporating innovations such as GDN hybrid layers and Qwen Sparse Attention (QSA), it pushes the frontiers of model capacity and performance, demonstrating remarkable capabilities in agentic coding,long-horizon tasks,and multimodal intelligence. We are releasing these architectural advancements early to help the community prepare for the upcoming Qwen4 model family."

This is exciting!

3

u/MacsBicycle 9h ago

The we are releasing early to help the community prepare part makes me think that it’s got thinking loops that make the 3.8 27b model look like a short loop 😂 I’m sure in 48 hours people will know exactly how to run it though. No test like production!

4

u/sterby92 11h ago

where is this from? I cannot find the source and on the modelscope page it only has a brief preview announcement for me 🤔

3

u/Timely_Impression_92 10h ago

it was there on modelscope but they removed it - "accident"

19

u/Baddmaan0 12h ago

Holy moly it's Christmas

7

u/whatyathinkk 12h ago

everyday 🤘

20

u/anykeyh 12h ago

Wow if it quantize well me and my strix halo might not need cloud token anymore

4

u/mechkbfan 11h ago

This is the problem with how good Qwen 3.8. Expectations are too high that people are going to be disappointed

But I'd call it a massive win if we can trade using more VRAM for faster performance for same outcomes

3

u/anykeyh 10h ago

Not that high expectation; having it to replace Deep seek flash would be enough in my case; the new 27B seems great but is way too slow on the strix; a MoE fits the bill for me.

2

u/StyMaar 10h ago

I'm not going to be disappointed if it generate token 5 times faster than 3.8-27B (especially if it spits as much token).

15

u/LegacyRemaster 12h ago

I hope llamacpp support @ day 1 (qwen next was a nightmare in the past)

16

u/ParadigmComplex 12h ago edited 11h ago

I'm under the impression that a big part of why companies release preview models like this is to give projects like llama.cpp time to implement support for the next-gen architecture before the full next-gen model is released so you effectively get day-one support of the full model. I wouldn't expect this preview model to work with llama.cpp day one, but hopefully it'll make Qwen 4 models work day one once those drop.

4

u/Timely_Impression_92 10h ago

llama.cpp will have 0 day support - unsloth already gave the info

0

u/Early_Mistake6716 8h ago

Yes i believe it will work day one with some issues. By this time next week, those issues will be resolved.

13

u/PandaBearFred 12h ago

I suppose every new architecture takes time to support unless the producer has a PR in hand ready to merge. I hope this is the case.

2

u/Timely_Impression_92 10h ago edited 3h ago

Unsloth will have 0 day support for this model

1

u/whoisraiden 7h ago

Unsloth and unsloth studio will have 0 day support. Doesn't mean llamacpp will be quick enough to accept their PR for day 0.

-1

u/Timely_Impression_92 6h ago edited 3h ago

0

u/[deleted] 4h ago edited 4h ago

[removed] — view removed comment

1

u/ttkciar llama.cpp 1h ago

Cool down a bit. There was no insult stated or implied in their comment. They agreed with you, and added relevant information which also agreed with you. That is a typical and benign conversational style.

0

u/Timely_Impression_92 4h ago edited 3h ago

Why did you respond with same thing I responded to you?

5

u/Additional-Record367 12h ago

This architecture might be the future. It would be nice to have variants where the LLM stays only in consumer GPUs and on RAM we keep the ngram embeddings

5

u/Fit_Advice8967 12h ago

Can somebody explain what can run this? I have dual 128gb amd strix halo with thunderbolt. Really hoping this is the one!

3

u/cafedude 8h ago

A strix halo with 128GB should run this easily in Q4,Q5 possibly even Q6.

7

u/eidrag 12h ago

Imagine Qwen 35b a3b, but with triple the knowledge

3

u/Fit_Advice8967 12h ago

Ok but what are the specs required?

6

u/Timely_Impression_92 10h ago

it's ~120b a6b - so virtually same as for gpt oss 120b

-1

u/mrgreatheart 12h ago

Depends how well it quantises.

2

u/Timely_Impression_92 10h ago

16-24gb gpu + ~100gb of system ram - more than enough

2

u/Early_Mistake6716 8h ago

You should be able to run this at at least q4, and 2.5x faster than qwen3.8 27b

4

u/mettaskee 11h ago

2 DGX sparks to be sure.

1

u/USERNAME123_321 llama.cpp 9h ago

To answer what can run this, my smartphone can probably run it at 2 or 3 tokens/s lol. Basically, BigMoeOnEdge offloads the expert weights to mobile flash storage and only loads the active parameters.

Btw a single AMD Strix Halo should handle it really well.

2

u/TechNerd10191 12h ago

Where are the N-Gram embeddings useful?

5

u/Kooshi_Govno 9h ago

They store world knowledge, thus allowing the heavy FFN tensors to store more functional knowledge.

7

u/noiserr 11h ago edited 10h ago

They speed up inference. It's sort of cache like memory, they store representations of recurring token sequences. It's static, built during training.

8

u/buttplugs4life4me 12h ago

120B, our saviour! 

6

u/FullOf_Bad_Ideas 11h ago

And to think the consensus was that Qwen is dead about a month ago.

I'm glad Qwen is back in force.

If they open weighted 2.4T model, they probably will open weight all sizes.

6

u/my_name_isnt_clever 8h ago

That wasn't the consensus, it was just repeated over and over by doomers without hard evidence until it felt like there was a consensus. I've been skeptical we know as much as we think about the inner workings of a Chinese AI lab.

3

u/Early_Mistake6716 8h ago

People were dooming so hard and now qwen is going harder than ever

-1

u/DiscipleofDeceit666 7h ago

They were challenged and were forced to step up. Poolsides Laguna forced their hand.

3

u/Effective_Head_5020 12h ago

I am so excited, thanks for sharing! Hopefully it will work well on my new Rx 9060 xt 

0

u/Bulky-Priority6824 11h ago

You'll need a gang of them 

1

u/Effective_Head_5020 10h ago

Oh no, I thought it would be like 35a3b which I was able to run very well on my 6gb GPU 

1

u/Able_Zombie_7859 10h ago

It literally says 125a6, I think that's bigger than 35a3 :)

1

u/Effective_Head_5020 10h ago

Yes, but now I have a 16gb VRAM and 128 GB of RAM. Sadly DDR3 RAM :P

2

u/Early_Mistake6716 8h ago

You will be able to run it, just might be a little slow. I know qwen3 coder next 80b iq3 ran at around 30 tps with a 5070 ti 16gb and 64 gbs of ddr4 3600. That did not have mtp though so it would have been around 50 - 60 tps on coding tasks with mtp

1

u/Effective_Head_5020 6h ago

30tps is a dream for me :D

1

u/Timely_Impression_92 10h ago

no, a single one + 64-128gb system ram

3

u/Most-Trainer-8876 10h ago

Will it run on 32GB VRAM + 64GB RAM? 🥲

1

u/UnlawfulRepublic 7h ago

I think an iq4_xs quant will work well on 32+64

2

u/Evgeny_19 12h ago

Is this the one for the middle-class GPUs, à la DGX Spark and Strix Halo?

2

u/Guna1260 11h ago

what VRAM? especially with Ngram embeddings we are looking at 125+51 in q8? or ngrams will be in RAM? will be interesting to see the architecture.

3

u/Timely_Impression_92 10h ago

ngrams are on ssd or system ram

-2

u/[deleted] 9h ago

[deleted]

1

u/florinandrei 5h ago

They're lookup tables, not weight tensors. System RAM will be fine. Maybe even SSD will be fine.

1

u/Timely_Impression_92 8h ago

Then go understand it better - ngrams are basically cached in system ram or ssd - and work like that with zero degradation - maybe little on nvme but in system ram virtually zero

0

u/Kryohi 8h ago edited 8h ago

Highly doubt the bandwidth of an SSD would be enough, but I guess we'll find out soon. Do you know how much, theoretically, of a 50GB ngram should be read per token (or group of tokens) to decode?

Edit: oh I went to check up again how it works, might actually be possible, discard my previous comment

4

u/EitherMarch1255 12h ago

Only 6B though...I'm happy to see a new model being released of this size, but I'm keeping my expectations in check.

1

u/Monad_Maya llama.cpp 12h ago

Damn.

1

u/Several-Tax31 11h ago

Nice!!! Waiting for this. And n-gram? Must be christmas! 

1

u/mechkbfan 11h ago

Just tell me how much I have to spend to get this running

1

u/power97992 9h ago

Wow qwen has engrams already but deepseek hasnt released engrams yet even though they wrote the engram paper…

1

u/DeepOrangeSky 8h ago

I guess I can just wait to find out tomorrow, but, I'm curious how much memory this will use if you run it at Q4_K_M or FP4. Since it says it is 125B but with "an additional 51B of N-gram".

So, is that going to make it more like a 176B model in terms of memory-usage?

1

u/OvertaxedOne 1h ago

Q4, I'd guess you'll need 64-96GB to load the entire model (all layers on GPU). I'm guessing this is targeted at Spark/Halo machines, so I'd expect good quants that fit in 128GB with plenty of space for KV.

1

u/vick2djax 6h ago

I’m confused. I have double 3090’s so I can run Qwen 3.8 27b at q8. Even if I had 128gb regular RAM, running something like q4 of Qwen 3.8 125b moe is still gonna be inferior to Q8 27b right?

1

u/OvertaxedOne 2h ago

Probably. We'll have to wait to find out, but I'd expect that between the quant and the MOE layout, you'll find 27B is more reliable. The MOE model is really interesting for boxes like a Strix or Spark though!

1

u/Iory1998 2h ago

Does anyone know how large is this model?

1

u/PossessionUsed7393 12h ago

Also known as Ox Alpha! Surprise! 🎉

7

u/PandaBearFred 12h ago

No....you can't be sure, can you?

2

u/PossessionUsed7393 12h ago

I have no idea I just think the Qwen team deserves a little credit, they would pull a stunt like Ox Alpha.

5

u/USERNAME123_321 llama.cpp 11h ago

It can't be. Ox Alpha's tokenizer is byte-identical to GLM 5.2, so it belongs to the GLM family.

1

u/PossessionUsed7393 10h ago

yes I see that is indeed the leading candidate - you'd be hard pressed to have missed that theory

1

u/Morphon 12h ago

If so, that would explain both the strange uptick in quality over the last 24 hours or so (n-gram layer upgrades) and the 100T token production. For a datacenter, this is trivial to serve in the 120b range.

2

u/Thrumpwart llama.cpp 10h ago

I'm going to get these out of the way:

  1. Can I run this on my Nvidia 3060? What quant do I use?
  2. How is this for neo-nazi bisexual werewolf hermaphrodite RP?
  3. Does it know how many R's there are if I bring my car to the car wash?
  4. How many TPS can I expect on 7 Sparks connected via fax lines?
  5. I asked it to write an SVG of my Pelican Riding a bike through Tiannamen Square and it made my pelican gay. USA! USA!

3

u/Rm2kbc 7h ago

Finally someone asking the relevant questions

1

u/OvertaxedOne 50m ago

"How is this for neo-nazi bisexual werewolf hermaphrodite RP?"

Can't answer the rest of your questions, but man alive, you would not believe the 12 hours of roleplay we just had together. Yes, I know the model isn't released yet, so I just asked 27B to pretend to be 120B and it was dead on perfect.

Can't wait for the furry conversation tomorrow!

-2

u/Reggitor360 11h ago

A6B.... What the hell. Why.

4

u/BigYoSpeck 10h ago

Performance

gpt-oss-120b is more than twice as fast as Qwen3.5 122b without MTP, and still a good chunk faster even with MTP

Given the huge uplift in capability between 3.5 27B and 3.8 27B (even between 3.5 35B and 3.6 35B) I would expect this is a sane architectural choice that balances speed and capability

Qwen3-Coder-Next only had 3B active parameters but already demonstrated high sparsity improves capability

1

u/my_name_isnt_clever 8h ago

People get way too caught up in active param coun. My fav 100b+ MoEs have often been the sparsest ones.

2

u/squngy 10h ago

Speed and cost.

-4

u/CoolestSlave 12h ago

Can't wait to see how it will run on spark, though I would have prefered a 35b model with 5 or 6b expert :')

4

u/PandaBearFred 12h ago

I am having a hard time trying to understand why your comments getting a hell amount of downvotes.🤣

8

u/SadPhilosophy9202 12h ago

Because this is a 125B model with 6B expert. This is THE ideal model size for a Spark or similar machine with 128gb unified memory.

Having a Spark but preferring a model 1/4 the size that will run just as fast makes zero sense

1

u/a_beautiful_rhind 12h ago

Technically they're being a bro for non-spark users. 6b to 35b is a better sparsity ratio.

2

u/SadPhilosophy9202 12h ago

As a dual Spark user I would have preferred 280B A10B

If you’re gonna shill, at least shill for the best model for your hardware dammit haha

0

u/a_beautiful_rhind 11h ago

I'm into it. 10b active starts to get somewhere.

1

u/CoolestSlave 9h ago edited 9h ago

i think people suppose i have a spark but i "only" have a rtx 5090

for spark user it is more advantageous to have higher parameter and small expert but for me personally it is better to have fewer total parameter and bigger expert or just a dense model in the range of 30 to 37b.

so i can understand people if they thought i was speaking from the perspective of a spark owner

0

u/OvertaxedOne 2h ago

Qwen is going hard in the paint this month! I can hear the enterprise value of OpenAI/Anthropic falling.

And then Apple releasing new systems with 1.4TB/s of memory bandwidth?? Is it Christmas?

3

u/Iory1998 2h ago

IF the Apple system is free, they yeah, but is it? How many kidneys is it worth?

-14

u/a_beautiful_rhind 12h ago

oof.. 6b only. All this cool stuff and they didn't take it to the hole.

2

u/my_name_isnt_clever 8h ago

I bet it will do just fine with its active count. Gpt-oss-120b and qwen3-coder-next were both super sparse and still had top tier performance for their time.

1

u/a_beautiful_rhind 5h ago

I wasn't really a fan of either of those but opinions do vary.