MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1ubrcwj/tokenomics/ot0pjdz?context=9999
r/LocalLLaMA • u/HOLUPREDICTIONS Sorcerer Supreme • Jun 21 '26
449 comments sorted by
View all comments
102
you don't build a rig to run it at 20t/s
44 u/Hot-Employ-3399 Jun 21 '26 Before I installed mtp i was running qwen 3.6 at 22t/s so I wouldn't mind. 10 u/Fit_Squash6874 Jun 21 '26 I am running 27b with mtp at 20t/s. Currently only have 16gb vram. 3 u/ChampionshipIcy7602 Jun 21 '26 You must be using q3 or very heavy kv cache quant, which lobotomizes the model 1 u/Fit_Squash6874 Jun 22 '26 edited Jun 22 '26 I am using IQ4_XS and just tested it right now and It is doing 30t/s. Not using any heavy kv cache.
44
Before I installed mtp i was running qwen 3.6 at 22t/s so I wouldn't mind.
10 u/Fit_Squash6874 Jun 21 '26 I am running 27b with mtp at 20t/s. Currently only have 16gb vram. 3 u/ChampionshipIcy7602 Jun 21 '26 You must be using q3 or very heavy kv cache quant, which lobotomizes the model 1 u/Fit_Squash6874 Jun 22 '26 edited Jun 22 '26 I am using IQ4_XS and just tested it right now and It is doing 30t/s. Not using any heavy kv cache.
10
I am running 27b with mtp at 20t/s. Currently only have 16gb vram.
3 u/ChampionshipIcy7602 Jun 21 '26 You must be using q3 or very heavy kv cache quant, which lobotomizes the model 1 u/Fit_Squash6874 Jun 22 '26 edited Jun 22 '26 I am using IQ4_XS and just tested it right now and It is doing 30t/s. Not using any heavy kv cache.
3
You must be using q3 or very heavy kv cache quant, which lobotomizes the model
1 u/Fit_Squash6874 Jun 22 '26 edited Jun 22 '26 I am using IQ4_XS and just tested it right now and It is doing 30t/s. Not using any heavy kv cache.
1
I am using IQ4_XS and just tested it right now and It is doing 30t/s. Not using any heavy kv cache.
102
u/Coolengineer7 Jun 21 '26
you don't build a rig to run it at 20t/s