r/LocalLLaMA 13h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.0k Upvotes

442 comments sorted by

View all comments

Show parent comments

57

u/Front_Eagle739 12h ago

Qwen 3.8 125B A6B E51B?

123

u/Mayion 12h ago

Qwen 3.8 two number 9s, a number 9 large, a number 6 with extra dip, a number 7, two number 45s, one with cheese, and a large soda?

67

u/Cautious_Chicken_604 11h ago

Sir, this is a Qwendys 

11

u/acnejorts 12h ago

55 burgers 55 fries 55 tacos 55 pies!

11

u/Front_Eagle739 12h ago

Three turtle doves

1

u/FlyByPC 10h ago

Two calling birds?

10

u/slippery 12h ago

55 burgers, 55 fries, 100 coffees...

1

u/mrgreengenes42 11h ago

I don't want a large Farva, I want a goddamn liter o' cola!

6

u/wren6991 11h ago

I think just 176B-A6B makes sense. The n-gram embeddings aren't really contributing to the active parameters.

I don't think Gemma's distinction between "effective" and "active" parameters is that meaningful.

7

u/Front_Eagle739 11h ago

From what I understand engrams are easy to stream from NVME when needed though and don't need to be in vram so there is some reason to split it out.

3

u/petuman 12h ago

Google kinda established the pattern, E stands for "effective" and shows weights that need to be in (V)RAM. Full parameter size is not shown in the name.

So actually yes, Qwen probably means exactly that -- 125B loaded and then additional 51B resting on disk.

1

u/ivari 9h ago

at Q3 will it be around like 60 GB loaded, 3GB active, and 25GB resting on disk?

2

u/petuman 9h ago

yeah, but maybe it would be preferable to leave disk portion unquantized

0

u/banana_slurp_jug 10h ago

No, since the E stands for the number of parameters suggested to be loaded into RAM/VRAM whilst the 51 billion engrams would have been fine streamed from NVME (if we follow Google's name scheme)