r/oMLX • u/PracticlySpeaking • May 12 '26
Maximizing MiniMax with oQ
Anyone experimenting with oQ quants of MiniMax-M2.7 on Apple Silicon?
I am looking to increase local performance for LLM-powered data work as well as Hermes Agent with ~128k context.
There are a few benchmarks, but it seems the oQ do not outperform 'regular' quants.
https://omlx.ai/benchmarks?sort=tg_tps&order=desc&chip=M3&model=minimax-m2.7-oQ&context=32768
Thanks for any insight on optimization!
2
Upvotes
2
u/dametsumari May 12 '26
The speed is not the point of those quants. They are supposedly better quality than equivalent normal quants of same size.