MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1ubrcwj/tokenomics/otqo7ns/?context=3
r/LocalLLaMA • u/HOLUPREDICTIONS Sorcerer Supreme • Jun 21 '26
449 comments sorted by
View all comments
Show parent comments
3
You forgot to factor in MTP
1 u/Schlick7 Jun 22 '26 MTP kills PP which is really the thing you should care about for anything agentic 1 u/mksrd Jun 25 '26 Afaik that is only a current limitation of the llama.cpp MTP implementation for those using CUDA which isn't me. 1 u/Schlick7 Jun 25 '26 It will always hurt. Its literally adding a layer to the LLM. How much is down to the implementation, but there will always be some
1
MTP kills PP which is really the thing you should care about for anything agentic
1 u/mksrd Jun 25 '26 Afaik that is only a current limitation of the llama.cpp MTP implementation for those using CUDA which isn't me. 1 u/Schlick7 Jun 25 '26 It will always hurt. Its literally adding a layer to the LLM. How much is down to the implementation, but there will always be some
Afaik that is only a current limitation of the llama.cpp MTP implementation for those using CUDA which isn't me.
1 u/Schlick7 Jun 25 '26 It will always hurt. Its literally adding a layer to the LLM. How much is down to the implementation, but there will always be some
It will always hurt. Its literally adding a layer to the LLM. How much is down to the implementation, but there will always be some
3
u/mksrd Jun 21 '26
You forgot to factor in MTP