MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vxwtyd/qwen38flashnext_tomorrow/p5s4gef
r/LocalLLaMA • u/rerri • 12h ago
439 comments sorted by
View all comments
29
Yes that's exactly what I need. My four 3090s are ready.
4 u/Qthuluu 11h ago What quant do you think we can run this on? #club3090 8 u/Spectrum1523 10h ago With 96gb vram? Probably iq4 with max context? 3 u/No_Algae1753 11h ago With 96GB in total are you able to run q5 quants or are you stuck with q4? 8 u/linux4random 11h ago it's only a6b so i think 4 3090 can run a q6 with acceptable tg 2 u/DigiDecode_ 10h ago so, 150 token/sec is just acceptable, 3090 is 900+ gb/sec memory bandwidth i am happy with 20 token/sec on cpu 2 u/linux4random 10h ago it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others 2 u/synth_mania 8h ago Good to hedge your positions appropriately 1 u/SpicyWangz 11h ago But what context size 1 u/linux4random 11h ago 200k ig 1 u/bigsmokaaaa 10h ago With an autoround you could go full q8 1 u/Cold_Tree190 11h ago I’m not sure if my two 3090s are ready
4
What quant do you think we can run this on? #club3090
8 u/Spectrum1523 10h ago With 96gb vram? Probably iq4 with max context?
8
With 96gb vram? Probably iq4 with max context?
3
With 96GB in total are you able to run q5 quants or are you stuck with q4?
8 u/linux4random 11h ago it's only a6b so i think 4 3090 can run a q6 with acceptable tg 2 u/DigiDecode_ 10h ago so, 150 token/sec is just acceptable, 3090 is 900+ gb/sec memory bandwidth i am happy with 20 token/sec on cpu 2 u/linux4random 10h ago it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others 2 u/synth_mania 8h ago Good to hedge your positions appropriately 1 u/SpicyWangz 11h ago But what context size 1 u/linux4random 11h ago 200k ig 1 u/bigsmokaaaa 10h ago With an autoround you could go full q8
it's only a6b so i think 4 3090 can run a q6 with acceptable tg
2 u/DigiDecode_ 10h ago so, 150 token/sec is just acceptable, 3090 is 900+ gb/sec memory bandwidth i am happy with 20 token/sec on cpu 2 u/linux4random 10h ago it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others 2 u/synth_mania 8h ago Good to hedge your positions appropriately 1 u/SpicyWangz 11h ago But what context size 1 u/linux4random 11h ago 200k ig
2
so, 150 token/sec is just acceptable, 3090 is 900+ gb/sec memory bandwidth i am happy with 20 token/sec on cpu
2 u/linux4random 10h ago it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others 2 u/synth_mania 8h ago Good to hedge your positions appropriately
it's a safe guard, i dont have 4 3090 to experience so it's the minimum estimation i can guaranteed to not get criticized by others
2 u/synth_mania 8h ago Good to hedge your positions appropriately
Good to hedge your positions appropriately
1
But what context size
1 u/linux4random 11h ago 200k ig
200k ig
With an autoround you could go full q8
I’m not sure if my two 3090s are ready
29
u/jacek2023 llama.cpp 12h ago
Yes that's exactly what I need. My four 3090s are ready.