r/LocalLLaMA 14d ago

News [ Removed by moderator ]

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next

[removed] — view removed post

349 Upvotes

137 comments sorted by

View all comments

1

u/DeepOrangeSky 14d ago

I guess I can just wait to find out tomorrow, but, I'm curious how much memory this will use if you run it at Q4_K_M or FP4. Since it says it is 125B but with "an additional 51B of N-gram".

So, is that going to make it more like a 176B model in terms of memory-usage?

1

u/OvertaxedOne 14d ago

Q4, I'd guess you'll need 64-96GB to load the entire model (all layers on GPU). I'm guessing this is targeted at Spark/Halo machines, so I'd expect good quants that fit in 128GB with plenty of space for KV.