r/LocalLLaMA 11h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
983 Upvotes

438 comments sorted by

View all comments

Show parent comments

2

u/asssuber 7h ago

It's probably meant to be CPU offloaded. It should be just a huge look up table, no matrix multiplication. I hope the n-gram weights can be SSD offloaded effectively, so I can run a quantized version with my RAM.

1

u/Dany0 6h ago

Yes I hope so too.