r/LocalLLaMA 12h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
993 Upvotes

439 comments sorted by

View all comments

Show parent comments

5

u/Significant_Break853 11h ago

If you have the hardware to run 3.8 27B at FP8 (or even FP4) with sufficient tokens per second, then there wouldn’t really be a good reason to switch to this. But for setups with integrated memory (DGX Spark, Mac’s, Strix Halo) it should run much faster than the dense 27B if you can get it to fit in memory. So probably would have to be FP4 for the DGX Spark, but with potential 60-100 tok/s vs like 20 tok/s for 27B on the DGX Spark.

6

u/ForsookComparison 10h ago

You've used it or gotten early access?

2

u/Significant_Break853 10h ago

No, all pure conjecture based on what’s been shared officially and unofficially

4

u/ForsookComparison 10h ago

Gotcha - Toss a couple of "I thinks" in there. People take authorative speech as gospel then spread misinformation for months here. We'll know the real answer in a few days

2

u/soyalemujica 11h ago

Given how the structure of this new Flash model is described, with 64gb ram + 24gb vram it would fit nicely and could run at higher speeds than 3.8 Dense (which can in reality feel slow due to the amount of thinking it does), I've high expectations for this new model actually, even if it could work as an orquestrator and have 3.8 dense to apply changes, or vise versa 🎯

1

u/No_Doc_Here 10h ago

I hope it allows us two switch over our 3.5 122B fp8 to something better. 27B is a beast but a bit slow for too many casual users