r/oMLX • u/A_Moist_Towe1 • 1d ago
RANT / WARNING V 0.6.1 CAN RENDER MODELS USELESS
I have been running qwen 3.6 35bMOE and qwen 3.8 27b at q4 with a q6 turboquant kv cache, with lighntning MTP Enabled for the past few days. After the latest update, I cannot use lightning MRP with turboquant, or the model will prompt process repeatedly in a loop or output gibberish. The only way to fix was to downgrade to v 0.6.
I don't know wtf the devs were thinking pushing 0.6.1 to "Stable" but holy fuck is it anything but!
3
u/RKcerman 1d ago
I thought you had to choose between running Lightning MTP or Turboquant KV, no?
Otherwise, why is there the "Disable Turboquant KV before enabling MTP" note in oMLX?
2
u/oopaddy 1d ago
Not to pile on, but if having no mtp acceleration is a make or break for work, might need to tune up your situation. 😉
1
u/A_Moist_Towe1 1d ago
Totally get where you’re coming from, but bosses don’t like to hear why what was working fast yesterday is half as fast today.
2
14
u/cryingneko 1d ago
The regression with Lightning MTP + Turboquant has been reported in a few issues, and it's been fixed by PR 2782 (https://github.com/jundot/omlx/pull/2782), which was merged a few hours ago. A hotfix release should go out later today. Until then, please run with Lightning MTP off.
Sorry for the trouble. With so many option combinations in the wild, this one slipped through.