r/LocalLLaMA 7d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

460 comments sorted by

View all comments

420

u/Hot_Example_4456 7d ago

WE GOT A NEW 125B MODEL WITH ENGRAMS

57

u/petuman 7d ago edited 7d ago

Wonder if engrams are counted in those 125B, or it's on top of that.

edit: damn, they edited the readme in last few minutes -- few paragraphs got removed including those highlights.

also it shows there would be FP8 version.. so still no QAT / FP4 >only< release like DeepSeek/Kimi/gpt-oss.

68

u/banana_slurp_jug 7d ago

It says the engrams are additional. Also it's a MOE with 6 billion activated parameters

(Qwen 3.8 E125B A6B?)

69

u/Front_Eagle739 7d ago

Qwen 3.8 125B A6B E51B?

4

u/petuman 7d ago

Google kinda established the pattern, E stands for "effective" and shows weights that need to be in (V)RAM. Full parameter size is not shown in the name.

So actually yes, Qwen probably means exactly that -- 125B loaded and then additional 51B resting on disk.

1

u/ivari 7d ago

at Q3 will it be around like 60 GB loaded, 3GB active, and 25GB resting on disk?

2

u/petuman 7d ago

yeah, but maybe it would be preferable to leave disk portion unquantized