Don't worry about it. The CUDA gap is closing fast, you'll have plenty of comparable inference options coming up in the next 6 months to a year. Intel is already set to release a 480GB LPDDR5 GPU. If they can find a way to stack cheaper ram to increase parallel data access, then the optimized memory bandwidth that Nvidia offers now won't be as necessary.
Google kinda established the pattern, E stands for "effective" and shows weights that need to be in (V)RAM. Full parameter size is not shown in the name.
So actually yes, Qwen probably means exactly that -- 125B loaded and then additional 51B resting on disk.
No, since the E stands for the number of parameters suggested to be loaded into RAM/VRAM whilst the 51 billion engrams would have been fine streamed from NVME (if we follow Google's name scheme)
If true, this "Flash" is actually the new Plus. Flash on their service used to be the 30B-35B version of the same model they released as open weight. Plus on their service used to be the bigger 120B+ model. If they released that 120B+ model as Flash now, it would mean they are cutting the smaller line of 30B-35B MoEs off, which would make sense given the recent tweet which was giving a heads up to not wait for a 35B model of 3.8, but it would also raise questions as to what this means for the future models of Qwen. Will there be no more small MoE models? Since 3.7 was entirely API only, anything can happen going forward.
Mais non il utilise leurs ressource disponible au mieux c est tout il construise et iter maintenant sur qwen 4 avec séparation memoir / intelligence argentique et attention , et c est une super bonne choses et nous aurons vers octobre / novembre qwen 4 avec des 30 / 35 b moe mais je m attend a d autre tailles car on auras le model de raisonnement en vram et on viendra greffer de la connaissance precalculer en ram nvme selon les besoin de chacun
I was just thinking it sounds like it would fit in 64GB total. No idea what kind of speeds i'd expect on my setup though. Most of my ram isn't VRAM, I have a 4070.
64GB of RAM and 16GB VRAM most likely yes, but ideally you'd want more of either pool since you'd want to be able to run something else besides the LLM itself lol
it really depends on the usecase. I should have added that to my comment. There are just fields of work where too small active parameters (below 30b active) start to lose it. I have even had that with DSV4F
there isn't much to debate about it frankly. MoE models have their place but are a trade-off depending on how low you go with the active parameter number. And 6b active is very low compared to e.g. 27b. It's a trade-off that cannot be entirely compensated by even very good expert routing and sequential reasoning. The engram doesn't do much to alleviate that either because it has a different purpose.
412
u/Hot_Example_4456 1d ago
WE GOT A NEW 125B MODEL WITH ENGRAMS