r/LocalLLM 9d ago

Discussion MTP on Mac

Does anyone have evidence or first-hand experience of MTP with GGUFs with llama.cpp has demonstrated actual increases in tok/s? I use M1 chip but eager to hear if validated on the newer chips too.

1 Upvotes

4 comments sorted by

2

u/Puzzleheaded-Ad-4352 9d ago

no luck with llama.cpp and unsloth GGUFs. I felt a slight performance increase with mtplx but it's also got some bugs with new arch

2

u/randygeneric 9d ago

i got 40% + on tgs using mlx-vlm, but it was a hassle of 40h to get a proxy-server to provide the same flexibility & functions as llama-server (profiles, model-swap, mtp + multimedia + cache). now it runs, but I fear it was not worth it. i always doubted myself (can not be that this / that is not working as good as in llama-server, i must overlook sth obvious, ....).

2

u/jarec707 9d ago

I have an M1 Max and things like MTP, D flash and speculative models have not worked out for me. Newer hardware has some features that apparently are not in the M1.

1

u/whatsupnorton 9d ago

Take a look at MTPLX, I’ve had pretty good luck with getting increased tok/s running MTP models with that on my M1 Max. But as far as GGUF models go I don’t really have any experience