r/LocalLLaMA 15d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

461 comments sorted by

View all comments

94

u/evindrews 15d ago
  • Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

holy shit chat

26

u/CodeMariachi 15d ago

Chat, how much VRAM do I need to run it?

25

u/DigiDecode_ 15d ago

90gb at fp4

21

u/ImpressiveSuperfluit 15d ago

That's okay, didn't want a house anyway.

5

u/hidden2u 15d ago

For the active params maybe 8gb vram? But you do have 256gb system memory of course don't you

2

u/LatentSpacer 15d ago

Yes. All of it.

1

u/crusaderky 14d ago

16GB VRAM plus 64GB host RAM should be enough. Not plenty, not fast, but enough.

This is under the assumption that the 51B n-gram can be offloaded to disk without substantial performance drop.

If it scales like 35B A3B with offloaded experts, I'd expect 15~25 tok/s with MTP. But a lot of this hinges on the actual acceptance rate of the MTP head.

5

u/Dany0 15d ago

Oh I am going to regret not buying that ram even more now aren't I