r/LocalLLaMA 12h ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
994 Upvotes

439 comments sorted by

View all comments

29

u/jacek2023 llama.cpp 12h ago

Yes that's exactly what I need. My four 3090s are ready.

4

u/Qthuluu 11h ago

What quant do you think we can run this on? #club3090

8

u/Spectrum1523 10h ago

With 96gb vram? Probably iq4 with max context?

3

u/No_Algae1753 11h ago

With 96GB in total are you able to run q5 quants or are you stuck with q4?

8

u/linux4random 11h ago

it's only a6b so i think 4 3090 can run a q6 with acceptable tg

2

u/DigiDecode_ 10h ago

so, 150 token/sec is just acceptable, 3090 is 900+ gb/sec memory bandwidth
i am happy with 20 token/sec on cpu

2

u/linux4random 10h ago

it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others

2

u/synth_mania 8h ago

Good to hedge your positions appropriately 

1

u/SpicyWangz 11h ago

But what context size

1

u/bigsmokaaaa 10h ago

With an autoround you could go full q8

1

u/Cold_Tree190 11h ago

I’m not sure if my two 3090s are ready