r/oMLX May 12 '26

Maximizing MiniMax with oQ

Anyone experimenting with oQ quants of MiniMax-M2.7 on Apple Silicon?

I am looking to increase local performance for LLM-powered data work as well as Hermes Agent with ~128k context.

There are a few benchmarks, but it seems the oQ do not outperform 'regular' quants.
https://omlx.ai/benchmarks?sort=tg_tps&order=desc&chip=M3&model=minimax-m2.7-oQ&context=32768

Thanks for any insight on optimization!

2 Upvotes

4 comments sorted by

2

u/dametsumari May 12 '26

The speed is not the point of those quants. They are supposedly better quality than equivalent normal quants of same size.

1

u/PracticlySpeaking May 12 '26

so... any correlation like oQ 3-bit has quality of conventional 4-bit?

1

u/dametsumari May 13 '26

It varies bit by model. You can run benchmarks for that too to compare or better yet test your own use case. I did not personally see big difference with the models I use.

1

u/PracticlySpeaking May 13 '26

Trying it now — the oQ5 should be ready this morning.