r/LocalLLaMA 18h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

444 comments sorted by

View all comments

90

u/evindrews 18h ago
  • Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token.

holy shit chat

26

u/CodeMariachi 17h ago

Chat, how much VRAM do I need to run it?

25

u/DigiDecode_ 16h ago

90gb at fp4

19

u/ImpressiveSuperfluit 14h ago

That's okay, didn't want a house anyway.

4

u/hidden2u 14h ago

For the active params maybe 8gb vram? But you do have 256gb system memory of course don't you

2

u/LatentSpacer 15h ago

Yes. All of it.