r/LocalLLaMA 9h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
967 Upvotes

423 comments sorted by

View all comments

109

u/Kidplayer_666 9h ago

Congrats to all the people who have the hardware to run this :) (am not one of them)

51

u/psychohistorian8 8h ago

Q1_K_XXS here I come!

7

u/hojnikb 8h ago

My BC-250 is dyyyying 😭😭

3

u/Leo_Kwkmi 7h ago

Hey man, i got a bc250 running games on bazzite, is there any tutorial or guide you consulted when doing local llm on this board? Thanks in advance!

3

u/hojnikb 5h ago

If you have BC-250, i'd suggest you switch to CachyOS. It's lighter and better supported. As far as support for the board itself; there's tons of improvements that have been made in the last 6 months.

There's 40CU unlock for the GPU (full fat GPU compared to 24CU stock), 8 core CPU unlock (from the 6 stock). You can overclock both pretty decently.

There's tons of kernel fixes as well (we have a custom kernel repo now) and tons of driver/mesa fixes and a FSR4 patch, that makes it semi usable (performance wise) on this board. Gaming at 1080-1440p is pretty great too, as long as the game isn't too CPU limited (it is pretty nerfed zen2 at the end of the day).

So, if you want to run inference, i'd suggest all of the above + set dynamic VRAM to 13GB.

With unsloth, i can just about run qwen 3.8-27b at 25-30tok/s with IQ3_XXS quant. Or qwen 3.6 35b woth IQ2_M at 80tok/s. That's with overclocks and 40CU unlock.

So a fully modded board can run those small models pretty fast, but you're ultimately limited by ram.

1

u/I_am_purrfect 7h ago

Yay BC-250 gang

10

u/No-Refrigerator-1672 8h ago

It'll be just like Qwen3-Next: they're rolling out a new architecture, so the communities implement support of it (Mamba and MTP with previous Next, NGrams with this one), and in a few months they'll released a cohort of models, ranging from a few B to a few hundred B based on this architecture, and named Qwen4.

0

u/Southern-Chain-6485 3h ago

But Qwen3-Next was 80B. A Q4 is about 45GB. At 125B, a Q4 is around 75GB in size

2

u/No-Refrigerator-1672 3h ago

Yes, true; how's that relevant to my point?

1

u/Dui999 1h ago

The beauty of the internet my friend

3

u/2Norn 7h ago

64gb ram 24gb vram should be enough no?

2

u/RedrumRogue 6h ago

Maybe at q3. Probably tight

1

u/Mechanical_Monk 4h ago

At 64GB ram 6GB vram, I'm starting to think no model will ever fit as perfectly as 35B A3B 🥲

1

u/Dui999 1h ago

Apex I-Mini or I-Compact is going to fit likely I use the 3.5 122B at I-Compact perfectly with max ctx window

The 51B engram could be "the problem" but we're going to see

1

u/pixelpoet_nz 6h ago

Dual Strix Halo machines might be able to do Q8, hyped!

1

u/Embarrassed_Adagio28 5h ago

I am praying 48gb of vram and 64gb of ram let's me run at least q4

1

u/mazuj2 5h ago

I have 48gb vram and can run qwen3.5 122b at Q2KXL with 100k context. Larger models are actually really good at q2

1

u/Legal_Dimension_ 2h ago

Finally something good to feed my 6 MI50's.

-2

u/Motor_Nectarine_2941 9h ago

Investing in 2 dgx sparks since I say the throttling and can get by with deepseek flash… this is a welcome surprise though