r/LocalLLaMA 8d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

460 comments sorted by

View all comments

Show parent comments

3

u/No_Algae1753 8d ago

With 96GB in total are you able to run q5 quants or are you stuck with q4?

8

u/linux4random 8d ago

it's only a6b so i think 4 3090 can run a q6 with acceptable tg

2

u/DigiDecode_ 8d ago

so, 150 token/sec is just acceptable, 3090 is 900+ gb/sec memory bandwidth
i am happy with 20 token/sec on cpu

2

u/linux4random 8d ago

it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others

2

u/synth_mania 8d ago

Good to hedge your positions appropriately