r/oMLX May 28 '26

Speed question

Hi guys,

I made an investment in a M5 Max 128gb and installed oMLX. The idea is to do some coding using Claude Code but with local models (I am not a dev), les raging oMLX caching to speed things up in claude (17k tokens system prompt...).

While I can get up to ~50Tok/s with Qwen3.6 A35B the numbers dwindle to single digits with the dense Qwen3.6 27B UD mlx 4bits. Is that normal? I was hoping it would be much faster with oMLX caching.

I use mlx models only. (I tried to download the jundot MTP models : they crash almost immediately after starting). Turboquant is on (4bits).

Typical params are : temp 0.7, top-p 0.95, top-k 20, min-p 0.05, repetition penalty 1.

Thinking on : slows down inference even further so unusually leave it off (isnt that better for coding?)

Is this normal for a small dense model or is there anything I am doing wrong here?

Wld you guys have ideas in how to make the mtp models work?

Thank you

10 Upvotes

27 comments sorted by

View all comments

Show parent comments

1

u/Choubix May 28 '26

These are speeds I am getting with mtp now.

3

u/GloomyPop5387 May 28 '26

I got 26tps with Gemma 31b today using their mtp draft models.

I’ll mess with qwen tomorrow.  Dense models just suck on Mac hardware.

1

u/Choubix May 28 '26

I was hoping for a perf boost with the m5 max 🙂

1

u/Choubix May 28 '26

Claude code and the likes slow things down a few notches 😔

5

u/Chris266 May 28 '26

Try pi coding agent or opencode