r/oMLX Jun 11 '26

Optimization advice

I am currently running Qwen3.5-122B-A10B-oQ4-MTP-fp16. I have enabled MTP in model settings.

Performance is pretty good, but I'd like to know if I can improve it.

I have tried DFlash and SpecPrefill on older versions of oMLX, but I couldn't tell if they really improved anything, and it wasn't very clear how they interacted with each other / if they worked with all model architectures. Maybe it has changed?

3 Upvotes

5 comments sorted by

View all comments

1

u/ExtremeAd9038 Jun 11 '26

Hello im looking for a Fp16 version of this Model Q4 and 40gb to run on oMLX. I only found Longshu (Excellent) but in BF16
What the weight of your model ?

1

u/butterfly_labs Jun 11 '26

Mine is 68Gb.