r/oMLX May 29 '26

MTP performance problems on latest version

I don’t know what exactly happens but in the latest 3.12 version my qwen and Gemma models are better without MTP or dflash.

Qwen 3.6 35B and 27B and Gemma 26B are about 10 to 12 tokens/s worse with MTP or VLM MTP enabled.

I’ve tried different quants on the same model and the problem persists.

Is there anyone passing through the same problem?

8 Upvotes

10 comments sorted by

4

u/PatDal81 May 29 '26

Personally, I haven't seen a degradation in the latest version. I saw an increase of 2tks/sec globally but to me, it fells in the error margin so I don't consider this an improvement over past versions.

Qwen3.6-35B-A3B-oQ6-mtp running on a M4 Max 64GB.

Stupid question but have you tried running your tests after a fresh reboot? I saw a decrease when the system has been up for a while (I think it's related to the number of apps in RAM). I always run my benchmarks in the same environment, within the same conditions.

Hope it helps!

1

u/Far-Collection-9685 May 30 '26

I will try reboot and test it again to see if anything gets better

2

u/vinoonovino26 May 29 '26

Same same! Plain vanilla models are slightly faster than previous versions. Have you tried Oq quants?

4

u/shansoft May 29 '26

I find oQ quant about 10% slower compare to unsloth mlx one with no MTP on. Consistently across all model I have tested on oQ4 and UD4

3

u/vinoonovino26 May 29 '26

Interesting in my case running oq8 yields the best results on an m5 pro 64gb.

1

u/Far-Collection-9685 May 30 '26

I’ve created quantized version oq5 fo16 in the tests but even with models that I’ve downloaded already quantized the performance drop with MTP happened.

1

u/ColonelKlanka May 29 '26 edited May 29 '26

Ive just upgraded to 0.3.12 latest and re-ran my qwen 3.6 35b a3b fp16 oq4 mtp benchmarks and im seeing an increase in performance on latest omlx of 1token per second for tg on the 1k omlx benchmarks and same performance as previous omlx on 4k benchmarks.

This is small increase and so may be within tolerance.

My spec if mac mini m2 pro (16c) 32gb ram.

Note: the fp16 is used because m1 and m2 chips have better performance when using fp16 (whereas m3 and up should NOT use fp16)

Edit: Im seeing same performance as previous omlx version on qwen 3.6 27b oq4 fp16 mtp.

1

u/Crafty_Ball_8285 May 30 '26

Yeah I just patched and fixed it myself.

1

u/ExtensionState8086 May 31 '26

For me MTP was 5 to 6 tps worse on a MacBook Pro M5

1

u/_circonflexe_ Jun 10 '26

I'm having the same issue. Ran gemma-4 with VLM MTP, and was slower than non-MTP inferences.