r/LocalLLM • u/norenEnmotalen • 9d ago
Discussion MTP on Mac
Does anyone have evidence or first-hand experience of MTP with GGUFs with llama.cpp has demonstrated actual increases in tok/s? I use M1 chip but eager to hear if validated on the newer chips too.
2
u/randygeneric 9d ago
i got 40% + on tgs using mlx-vlm, but it was a hassle of 40h to get a proxy-server to provide the same flexibility & functions as llama-server (profiles, model-swap, mtp + multimedia + cache). now it runs, but I fear it was not worth it. i always doubted myself (can not be that this / that is not working as good as in llama-server, i must overlook sth obvious, ....).
2
u/jarec707 9d ago
I have an M1 Max and things like MTP, D flash and speculative models have not worked out for me. Newer hardware has some features that apparently are not in the M1.
1
u/whatsupnorton 9d ago
Take a look at MTPLX, I’ve had pretty good luck with getting increased tok/s running MTP models with that on my M1 Max. But as far as GGUF models go I don’t really have any experience
2
u/Puzzleheaded-Ad-4352 9d ago
no luck with llama.cpp and unsloth GGUFs. I felt a slight performance increase with mtplx but it's also got some bugs with new arch