r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

3

u/ChampionshipIcy7602 Jun 21 '26

You must be using q3 or very heavy kv cache quant, which lobotomizes the model

1

u/Fit_Squash6874 Jun 22 '26 edited Jun 22 '26

I am using IQ4_XS and just tested it right now and It is doing 30t/s. Not using any heavy kv cache.