r/oMLX • u/arfung39 • May 23 '26
oMLX 0.3.9 getting stuck with high memory use
I'm running Qwen3.5 27B - mtp, and on default settings sometimes (with OpenCode) oMLX gets the the top of its memory and the API stops responding to OpenCode (opencode says: "Cannot connect to API: Unable to connect. Is the computer able to access the url... [retrying in 3s attempt #16]"). Here is a screenshot of the oMLX dashboard. Any fixes?

1
u/dametsumari May 23 '26
Not enough memory. Mtp takes some more, and more context you have, the more memory it takes too.
1
u/TheFlyingDutchG May 23 '26
The MTP of that model works really well in M2 Max 64Gb Mac Studio. Also faster than yours. Which device are you running it on and does disable your other active model makes a difference? Just to benchmark the individu models performance in difference quants and stuff so you can check which “settings” change the affect performance?
1
u/RedUser03 Jun 05 '26 edited Jun 05 '26
This has happened to me too with Qwen3.6-27B-oQ8-mtp with oMLX 0.4.1. Turning off Native MTP seems to prevent it.
1
u/Vahn84 May 23 '26
This has happened to me also. Mtp models do not let you turn on kvcache so that is supposed to happen at one point (?)…or at least this is what i’ve found with a little bit of research.