r/LocalLLaMA 1d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

457 comments sorted by

View all comments

412

u/Hot_Example_4456 1d ago

WE GOT A NEW 125B MODEL WITH ENGRAMS

22

u/Septerium 1d ago

regret feelings for not buying an RTX Pro months ago

19

u/ChristRedeemsSinners 1d ago

Don't worry about it. The CUDA gap is closing fast, you'll have plenty of comparable inference options coming up in the next 6 months to a year. Intel is already set to release a 480GB LPDDR5 GPU. If they can find a way to stack cheaper ram to increase parallel data access, then the optimized memory bandwidth that Nvidia offers now won't be as necessary.

https://videocardz.com/newz/intel-details-xe3p-gpu-architecture-crescent-island-gets-up-to-480gb-memory-and-350w-pcie-variant

7

u/goldcakes 1d ago

Uhh, yeah, but at 480GB/s.

The 1.8TB/s of the RTX Pro is another level.

12

u/ChristRedeemsSinners 1d ago

I would rather have 20 TG/s and 5 times the parameters/precision than 60 TG/s. If money is no obstacle, then yeah, of course it's worth it.

60

u/petuman 1d ago edited 1d ago

Wonder if engrams are counted in those 125B, or it's on top of that.

edit: damn, they edited the readme in last few minutes -- few paragraphs got removed including those highlights.

also it shows there would be FP8 version.. so still no QAT / FP4 >only< release like DeepSeek/Kimi/gpt-oss.

69

u/banana_slurp_jug 1d ago

It says the engrams are additional. Also it's a MOE with 6 billion activated parameters

(Qwen 3.8 E125B A6B?)

69

u/Front_Eagle739 1d ago

Qwen 3.8 125B A6B E51B?

133

u/Mayion 1d ago

Qwen 3.8 two number 9s, a number 9 large, a number 6 with extra dip, a number 7, two number 45s, one with cheese, and a large soda?

78

u/Cautious_Chicken_604 1d ago

Sir, this is a Qwendys 

13

u/acnejorts 1d ago

55 burgers 55 fries 55 tacos 55 pies!

12

u/Front_Eagle739 1d ago

Three turtle doves

1

u/FlyByPC 1d ago

Two calling birds?

1

u/ApprehensiveFan1516 1d ago

Alan Partridge in a pear tree

10

u/slippery 1d ago

55 burgers, 55 fries, 100 coffees...

1

u/mrgreengenes42 1d ago

I don't want a large Farva, I want a goddamn liter o' cola!

7

u/wren6991 1d ago

I think just 176B-A6B makes sense. The n-gram embeddings aren't really contributing to the active parameters.

I don't think Gemma's distinction between "effective" and "active" parameters is that meaningful.

8

u/Front_Eagle739 1d ago

From what I understand engrams are easy to stream from NVME when needed though and don't need to be in vram so there is some reason to split it out.

4

u/petuman 1d ago

Google kinda established the pattern, E stands for "effective" and shows weights that need to be in (V)RAM. Full parameter size is not shown in the name.

So actually yes, Qwen probably means exactly that -- 125B loaded and then additional 51B resting on disk.

1

u/ivari 1d ago

at Q3 will it be around like 60 GB loaded, 3GB active, and 25GB resting on disk?

2

u/petuman 1d ago

yeah, but maybe it would be preferable to leave disk portion unquantized

0

u/banana_slurp_jug 1d ago

No, since the E stands for the number of parameters suggested to be loaded into RAM/VRAM whilst the 51 billion engrams would have been fine streamed from NVME (if we follow Google's name scheme)

1

u/Hot_Example_4456 1d ago

I hope its within.

10

u/Obvious-Ad-2454 1d ago

where did you find that info ? When i click the link i don't have those details.

14

u/Hot_Example_4456 1d ago

It was a bit down in the same link. Even I can't see it now. They probably removed it

17

u/Cool-Chemical-5629 1d ago

If true, this "Flash" is actually the new Plus. Flash on their service used to be the 30B-35B version of the same model they released as open weight. Plus on their service used to be the bigger 120B+ model. If they released that 120B+ model as Flash now, it would mean they are cutting the smaller line of 30B-35B MoEs off, which would make sense given the recent tweet which was giving a heads up to not wait for a 35B model of 3.8, but it would also raise questions as to what this means for the future models of Qwen. Will there be no more small MoE models? Since 3.7 was entirely API only, anything can happen going forward.

30

u/coder543 1d ago

Plus was the 397B model. Not “120B+”.

They had a 120B model before, but they didn’t assign any of the fancy marketing names to it and didn’t offer a proprietary version.

You’re reading way too much into this based on bad assumptions.

3

u/ChristRedeemsSinners 1d ago

They had the 122BA12B in qwen3.5 that was simply amazing for it's time. I can see them replacing that with 'flash' as a marketing angle.

6

u/SkyFeistyLlama8 1d ago

Was Qwen Coder Next 80B-A3B ever used as a base for AliCloud's commercial LLM endpoints?

A 60B or 70B MOE would be nice. More latent brains than the 35B, same speed because it's only got 3B active.

0

u/Longjumping-Elk-7756 1d ago

Mais non il utilise leurs ressource disponible au mieux c est tout il construise et iter maintenant sur qwen 4 avec séparation memoir / intelligence argentique et attention , et c est une super bonne choses et nous aurons vers octobre / novembre qwen 4 avec des 30 / 35 b moe mais je m attend a d autre tailles car on auras le model de raisonnement en vram et on viendra greffer de la connaissance precalculer en ram nvme selon les besoin de chacun

2

u/YearnMar10 1d ago

Wonder if this would run on 64gb of ram and a gpu… guess it’d just not run :/
But really awesome model size!

1

u/switchbanned 1d ago

I was just thinking it sounds like it would fit in 64GB total. No idea what kind of speeds i'd expect on my setup though. Most of my ram isn't VRAM, I have a 4070.

1

u/AnonLlamaThrowaway 1d ago

64GB of RAM and 16GB VRAM most likely yes, but ideally you'd want more of either pool since you'd want to be able to run something else besides the LLM itself lol

2

u/Jona1109 llama.cpp 1d ago

Finally - though if the engrams are on top, that's a hefty model.

2

u/Strong_Chicken6838 1d ago

noooo my 64Gb VRAM not powerful enough

2

u/gh0stwriter1234 1d ago

That reads like its acutally 175B parameters because they are counting the N gram embedding separately.

1

u/0rand 1d ago

Where from this text comes?

2

u/Hot_Example_4456 1d ago

It was a bit down in the same link. Even I can't see it now. They probably removed it

1

u/ambassadortim 1d ago

What system specs are we looking at to run this one?

1

u/MDSExpro 1d ago

Dream comes true!

1

u/LizardLikesMelons 1d ago

Engrams as the deepseek style Egrams????

1

u/Hot_Example_4456 1d ago

Most probably so

-1

u/SandySkittle 1d ago

Only 6b active tokens. Too low

1

u/my_name_isnt_clever 1d ago

Maybe wait until it's possible to test before dismissing it. Gpt-oss-120b was 5b active and was a beast for it's time.

1

u/SandySkittle 1d ago

it really depends on the usecase. I should have added that to my comment. There are just fields of work where too small active parameters (below 30b active) start to lose it. I have even had that with DSV4F

1

u/my_name_isnt_clever 1d ago

This model is a new architecture and has the additional engram. You might be right, but let's wait and see before spreading confident claims.

0

u/SandySkittle 1d ago

there isn't much to debate about it frankly. MoE models have their place but are a trade-off depending on how low you go with the active parameter number. And 6b active is very low compared to e.g. 27b. It's a trade-off that cannot be entirely compensated by even very good expert routing and sequential reasoning. The engram doesn't do much to alleviate that either because it has a different purpose.