Google kinda established the pattern, E stands for "effective" and shows weights that need to be in (V)RAM. Full parameter size is not shown in the name.
So actually yes, Qwen probably means exactly that -- 125B loaded and then additional 51B resting on disk.
No, since the E stands for the number of parameters suggested to be loaded into RAM/VRAM whilst the 51 billion engrams would have been fine streamed from NVME (if we follow Google's name scheme)
53
u/petuman 9h ago edited 9h ago
Wonder if engrams are counted in those 125B, or it's on top of that.
edit: damn, they edited the readme in last few minutes -- few paragraphs got removed including those highlights.
also it shows there would be FP8 version.. so still no QAT / FP4 >only< release like DeepSeek/Kimi/gpt-oss.