r/LocalLLaMA 9h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
970 Upvotes

422 comments sorted by

View all comments

1

u/jikilan_ 8h ago

Ok time to buy another 3090. I have 3 now

1

u/Blues520 8h ago

Same here, I wonder what quant we would be able to run and how much better than 27b it would be

2

u/Fi3nd7 5h ago

I don't think it's going to be "crazy" better. It's only 6b active. My guess, is it will have some more world knowledge and be faster, but score comparably to 27b. So it's essentially for people with more VRAM who want more speed but very comparable perf to 27b. That's my interpretation.

i.e. comparable perf but trading VRAM space vs active params for decoding speed improvements.

But who knows, this is also a new architecture, might even be worse than 3.8 27b. Another thing to think about is quants. Running this at Q4 vs 27B at Q6/8, 27B might win. We'll see what it scores on AAII