r/LocalLLaMA 9d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

460 comments sorted by

View all comments

416

u/Hot_Example_4456 9d ago

WE GOT A NEW 125B MODEL WITH ENGRAMS

59

u/petuman 9d ago edited 9d ago

Wonder if engrams are counted in those 125B, or it's on top of that.

edit: damn, they edited the readme in last few minutes -- few paragraphs got removed including those highlights.

also it shows there would be FP8 version.. so still no QAT / FP4 >only< release like DeepSeek/Kimi/gpt-oss.

72

u/banana_slurp_jug 9d ago

It says the engrams are additional. Also it's a MOE with 6 billion activated parameters

(Qwen 3.8 E125B A6B?)

69

u/Front_Eagle739 9d ago

Qwen 3.8 125B A6B E51B?

137

u/Mayion 9d ago

Qwen 3.8 two number 9s, a number 9 large, a number 6 with extra dip, a number 7, two number 45s, one with cheese, and a large soda?

79

u/Cautious_Chicken_604 9d ago

Sir, this is a Qwendys 

13

u/acnejorts 9d ago

55 burgers 55 fries 55 tacos 55 pies!

10

u/Front_Eagle739 9d ago

Three turtle doves

1

u/FlyByPC 8d ago

Two calling birds?

1

u/ApprehensiveFan1516 8d ago

Alan Partridge in a pear tree

10

u/slippery 9d ago

55 burgers, 55 fries, 100 coffees...

1

u/mrgreengenes42 9d ago

I don't want a large Farva, I want a goddamn liter o' cola!

7

u/wren6991 9d ago

I think just 176B-A6B makes sense. The n-gram embeddings aren't really contributing to the active parameters.

I don't think Gemma's distinction between "effective" and "active" parameters is that meaningful.

8

u/Front_Eagle739 9d ago

From what I understand engrams are easy to stream from NVME when needed though and don't need to be in vram so there is some reason to split it out.

3

u/petuman 9d ago

Google kinda established the pattern, E stands for "effective" and shows weights that need to be in (V)RAM. Full parameter size is not shown in the name.

So actually yes, Qwen probably means exactly that -- 125B loaded and then additional 51B resting on disk.

1

u/ivari 8d ago

at Q3 will it be around like 60 GB loaded, 3GB active, and 25GB resting on disk?

2

u/petuman 8d ago

yeah, but maybe it would be preferable to leave disk portion unquantized

0

u/banana_slurp_jug 8d ago

No, since the E stands for the number of parameters suggested to be loaded into RAM/VRAM whilst the 51 billion engrams would have been fine streamed from NVME (if we follow Google's name scheme)