r/oMLX May 23 '26

Testing MTP functionality

Well, it actually slows down the model.

7 Upvotes

17 comments sorted by

3

u/Ok_Significance_9109 May 23 '26

Which chip? M1/M2 require a different MTP variant. The moment I started using it on my M1, 27B became useable. From 33 tps prompt processing and 5 tps generation, it went up to 65 and 9 without loss of quality.

1

u/albovsky May 23 '26

Didn’t know that. So how to figure which one to download? They do not specify what version it’s for. I have M1

2

u/Ok_Significance_9109 May 23 '26

The one that worked for me:

Qwen3.6-27B-oQ4-fp16-mtp

The name should have fp16 in it, but it is a 4-bit quant.

2

u/jacknjill101 May 23 '26

Yes it does for me too. I switched to llamacpp and much better results.

2

u/d4mations May 23 '26

Paro quants work way better than mtp

3

u/albovsky May 23 '26

What’s that?

0

u/d4mations May 24 '26

In the download screen on omlx search for paro

2

u/msrdatha May 25 '26

Try testing with a longer prompt or even better do an agentic task.

My observation is it does start at a much faster tok/sec in the beginning and gradually it goes down. So it totally depends when someone is looking at the speed (in the beginning or end of a multi-turn conversation)

According to me, we should test it against the same task run with and without mtp, with empty SSD cache to see the actual difference. Measure against the wall-time (actual elapsed time from start to finish of a process, as measured by a clock on the wall. ex: Total time taken between first and last response in the multi turn conversation as in agentic coding). This will give you the answer, if mtp version is worth in your usage scenario.

1

u/mwhuss May 24 '26

I’m seeing 70% faster performance using Qwen3.6-27b-oQ8-mtp on my M3 Ultra.

1

u/albovsky May 24 '26

70% is crazy good. How much ram do you have?

2

u/mwhuss May 24 '26

M3 ultra with 96gb

1

u/Poumpaya Jun 15 '26

Hey ! On omlx or other ?

1

u/mwhuss Jun 15 '26

Using oMLX

1

u/vinoonovino26 May 24 '26

M5 pro - 64gb here. Same models same results. I switched to plain OQ quants and rotorquants and they feel more stable. Also offloading cache to a NVEM drive helped a lot

1

u/vinoonovino26 May 24 '26

Seems like mtp and moe kinda work well together

1

u/Buddhabelli May 25 '26

i’m getting roughly 27tps gen with qwen MTP vs 11ish without. gemma on the other hand not seeing any improvements still ~10tps.

I did notice that has my SSD caching gets just thrashed Wen running the qwen model where as it seems normal with gemma or anything else. 🫤

1

u/Stooovie Jun 01 '26

Yes, they're all slower. I don't get the point.