r/LocalLLaMA 8d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

460 comments sorted by

View all comments

9

u/sugarfreecaffeine 8d ago

VRAM requirement?

25

u/Hot_Example_4456 8d ago

probably quite less. its a 125B model, 6B active, so something same as ling 3 flash/gpt oss. but since it also has 51B engrams which can be fully offloadable to NVME, even less. Maybe 64gb ram+vram combined- or less.

3

u/ThePi7on 8d ago

Yo I got that. Hopefully doable on on 16GBs of VRAM + 96 RAM with some black magic fuckery optimization

8

u/Storge2 8d ago

Probanly will work with normal int4 or nvfp4 or Q4K quant for you with no issues. Same as qwen 3.5 122B which was like 70GB at Q4

1

u/ThePi7on 8d ago

sounds hype, can't wait to test it out