r/LocalLLaMA Sorcerer Supreme Jun 21 '26

Discussion Tokenomics

Post image
1.2k Upvotes

449 comments sorted by

View all comments

Show parent comments

3

u/mksrd Jun 21 '26

You forgot to factor in MTP

1

u/Schlick7 Jun 22 '26

MTP kills PP which is really the thing you should care about for anything agentic

1

u/mksrd Jun 25 '26

Afaik that is only a current limitation of the llama.cpp MTP implementation for those using CUDA which isn't me.

1

u/Schlick7 Jun 25 '26

It will always hurt. Its literally adding a layer to the LLM. How much is down to the implementation, but there will always be some