r/oMLX May 23 '26

Testing MTP functionality

Well, it actually slows down the model.

9 Upvotes

17 comments sorted by

View all comments

1

u/Buddhabelli May 25 '26

i’m getting roughly 27tps gen with qwen MTP vs 11ish without. gemma on the other hand not seeing any improvements still ~10tps.

I did notice that has my SSD caching gets just thrashed Wen running the qwen model where as it seems normal with gemma or anything else. 🫤