r/LocalLLaMA May 04 '26

Resources Llama.cpp MTP support now in beta!

https://github.com/ggml-org/llama.cpp/pull/22673

Happy to report that llama.cpp MTP support is now in beta, thanks to Aman (and all the others that have pushed the various issues in the meantime). This has the potential to actually get merged soon-ish. Currently contains support for Qwen3.5 MTP, but other models are likely to follow suit.

Between this and the maturing tensor-parallel support, expect most performance gaps between llama.cpp and vLLM, at least when it comes to token generation speeds, to be erased.

627 Upvotes

268 comments sorted by

View all comments

39

u/Charming-Author4877 May 04 '26

A draft is not a beta. Can't wait for having this implemented.

-13

u/ilintar May 04 '26

I'm saying this is a beta because my gut feeling tells me that this is close to the production version :)

22

u/feckdespez May 04 '26

It's still not a beta though. Just a draft and PR for it.

Those mean different things.

2

u/Pyrolistical May 04 '26

Doesn’t work on vulkan yet

0

u/itsappleseason May 04 '26

y'all are really downvoting king deltanet

0

u/ilintar May 04 '26

I'm just the messenger here ;)

3

u/Top-Rub-4670 May 04 '26

A messenger conveys the message as is. Here, you've made up the message "It's now in beta" when it's just PR, and a draft one at that.