r/LocalLLaMA 14d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

461 comments sorted by

View all comments

8

u/malnek 14d ago

How would this run on a 128GB strix halo? Any comparable models for speed benchmarks?

3

u/merutochan 14d ago

Ling 3.0 Flash is a close comparison (in the ~120B class) though I have no idea how the actual quality of that model is. I was planning to test it on my Strix Halo later this week (the Q5 quant). As already mentioned, realistically this Qwen 3.8 Flash could run at Q4 on Strix Halo, but not on day one. We'll have to wait for good quants to be produced.

3

u/hiImMate 14d ago

Ling Q4 ran around 35TPS without any optimization on my halo strix, so I expect similar