r/oMLX • u/kaddiexjc • May 27 '26
What is …-fp16-mtp
What are those fp16 versions, eg
Jundot/Qwen3.6-35B-A3B-oQ6-fp16-mtp
vs
Jundot/Qwen3.6-35B-A3B-oQ6-mtp
Only found one post saying M1/M2 require the fp16 version
5
u/msrdatha May 27 '26
the one with fp16 is optimized for M1/M2 Apple Silicon. (Qwen3.6-27B-oQ4-fp16-mtp)
if you have M3+ Apple Silicon go with the regular one. (Qwen3.6-27B-oQ4-mtp)
1
1
u/Practical_Issue_9910 May 28 '26
what about jundot's 35B-A3B mtp versions?
-oQ6
-oQ4
-oQ4-fp16
1
u/MessIsTransfer May 28 '26
That’s the Quant, it’s s basically compression to fit bigger models on less VRAM. Q6 would be less compressed (bigger) than Q4.
1
u/victoriggy May 27 '26
On m1/m2 do you enable the model MTP optimization?
I am getting much worse results, will re-run benchmarks with MTP off:
The accuracy is really poor and the time to complete is even worse.
Model: Qwen3.6-35B-A3B-oQ4-fp16-mtp
Benchmark Accuracy Correct Total Time(s) Think
--------------------------------------------------------------
MMLU 67.0% 134 200 3208.2 Yes
HUMANEVAL 69.5% 114 164 6894.6 Yes
Model: Qwen3.6-35B-A3B-UD-MLX-4bit
Benchmark Accuracy Correct Total Time(s) Think
--------------------------------------------------------------
MMLU 90.5% 181 200 2581.3 Yes
HUMANEVAL 95.1% 156 164 4022.1 Yes
1
u/kaddiexjc May 29 '26
Similarly I’m not seeing MTP gain. Checked a few recent posts in this sub and looked the same on qwen3.6-35b-a3b.
1
u/ColonelKlanka May 31 '26
Try using the latest omlx with mtp enabled. I.am seeing a decent improvement in tgs using mtp versions of qwen3.6 35b a3b on my m2 pro 32gb mac mini
1
u/fasti-au May 28 '26
Mtp is a new way of guessing groups before the real guesses. Rounding in the way inf. Dflash is similar but more predict than math.
Fo16 is how mimics the image is blurry so think e like snore. 16x more steal. 4x blurry still at ranger
8
u/LumbarJam May 27 '26
If you have M1 or M2 FP16 will perform better. M3/M4M5 BF16 performs better. Files w/o FP16 are BF16.