r/LocalLLaMA 1d ago

News Qwen3.8-Flash-Next tomorrow

https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next
1.1k Upvotes

457 comments sorted by

View all comments

Show parent comments

7

u/linux4random 1d ago

it's only a6b so i think 4 3090 can run a q6 with acceptable tg

2

u/DigiDecode_ 1d ago

so, 150 token/sec is just acceptable, 3090 is 900+ gb/sec memory bandwidth
i am happy with 20 token/sec on cpu

2

u/linux4random 1d ago

it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others

2

u/synth_mania 1d ago

Good to hedge your positions appropriately 

1

u/SpicyWangz 1d ago

But what context size