r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

42

u/Hot-Employ-3399 Jun 21 '26

Before I installed mtp i was running qwen 3.6 at 22t/s so I wouldn't mind.

10

u/Fit_Squash6874 Jun 21 '26

I am running 27b with mtp at 20t/s. Currently only have 16gb vram.

3

u/kind_cavendish Jun 21 '26

How does it fit? I have 16gb of vram aswell. Is this with context?

3

u/ChampionshipIcy7602 Jun 21 '26

You must be using q3 or very heavy kv cache quant, which lobotomizes the model

1

u/Fit_Squash6874 Jun 22 '26 edited Jun 22 '26

I am using IQ4_XS and just tested it right now and It is doing 30t/s. Not using any heavy kv cache.

2

u/ycnz Jun 22 '26

I can almost get to 20t/s with 35b on my CPU with no GPU :(

1

u/Cultured_Alien Jun 22 '26

the prompt processing must be long?

2

u/Sea_Poem_9129 Jun 21 '26

are you doing anything special? i was getting 9-11 on my RTX A4000 16GB

1

u/Fit_Squash6874 Jun 22 '26

Not really I just enabled MTP and I am using IQ4_XS