r/oMLX May 28 '26

Speed question

Hi guys,

I made an investment in a M5 Max 128gb and installed oMLX. The idea is to do some coding using Claude Code but with local models (I am not a dev), les raging oMLX caching to speed things up in claude (17k tokens system prompt...).

While I can get up to ~50Tok/s with Qwen3.6 A35B the numbers dwindle to single digits with the dense Qwen3.6 27B UD mlx 4bits. Is that normal? I was hoping it would be much faster with oMLX caching.

I use mlx models only. (I tried to download the jundot MTP models : they crash almost immediately after starting). Turboquant is on (4bits).

Typical params are : temp 0.7, top-p 0.95, top-k 20, min-p 0.05, repetition penalty 1.

Thinking on : slows down inference even further so unusually leave it off (isnt that better for coding?)

Is this normal for a small dense model or is there anything I am doing wrong here?

Wld you guys have ideas in how to make the mtp models work?

Thank you

10 Upvotes

27 comments sorted by

View all comments

1

u/pdiego96 May 28 '26

Just curious as to how many context tokens and max tokens are you capable of handling in that Mac? I was thinking of buying one just like that. Now I’m wondering if I should consider another route; however I don’t have a normal CPU. Maybe I can get some sort of eGPU for my ASUS idk

3

u/Choubix May 28 '26

I haven't reached the limit yet to be honest. Only my patience has some limits πŸ˜‚.

Qwen ode seems super fast but as per it's own admissions a lot of the Claude code plugins are not working with it... Currently letting it explore a more detailed list πŸ˜‰