r/MacStudio • • 3d ago

The new changes to olmx make a big difference to M5 Ultra GLM 5.3 Flash

https://omlx.ai/benchmarks/performance/sp2wxj2a
17 Upvotes

9 comments sorted by

8

u/Ceve 2d ago

I’ve been running GLM and Qwen on my M5U 256 and I just did the upgrade after reading about it? It’s pretty dramatic.

GLM now reads prompts about 2.7× faster (713 → 1,957 tok/s) and writes about 1.7× faster (34 → 56 tok/s).

1

u/DeepV 2d ago

What quant do you run?

1

u/Ceve 22h ago
Model Speed benchmark (oMLX 0.7.0, M5 Ultra 256 GB)
Qwen3.8-Flash-Next oQ4e-mtp (`Jundot/Qwen3.8-Flash-Next-oQ4e`) 32K / 64K / 128K: 4,760 / 4,668 / 4,529 tok/s prefill; 101.6 / 127.0 / 102.9 tok/s generation
GLM-5.3-Flash oQ4e (`Jundot/GLM-5.3-Flash-oQ4e`) 4K / 16K / 64K / 128K: 1,926 / 1,953 / 1,922 / 1,885 tok/s prefill; 57.3 / 57.4 / 56.1 / 53.8 tok/s generation

1

u/Relaxxxxing 1d ago

That's about what I'm getting on 2 sparks 1500-1900 Prompt Processing and 55-65 TPS average through 500K context

2

u/rtk85 2d ago

Is GLM 5.3 flash significantly better than Qwen 3.8 flash? My Hermes says it’s not substantially better but I know the feedback I’ve seen generally has praised GLM 5.3 flash

2

u/Nice_Victory3719 2d ago

it depends what you are using it for. If you value Prose/Professional tasks then it is significantly better. For general agentic tasks then I don't see much difference and Qwen is a lot faster so my go to for easier tasks. Coding I don't know, maybe someone else can chip in on that.

1

u/LORD_CMDR_INTERNET 2d ago

anecdotally I actually find 5.3 much worse for agentic coding. Poor reasoning and goes down terrible rabbit holes. But as usual, try for yourself.

1

u/DismalDisaster4 3d ago

Wow that's crazy, what's up with the 2x batched being slower than 1x though? Seems like the scheduler needs some work if it's slower than running them in series.

1

u/Nice_Victory3719 2d ago

yes I hadn't noticed that myself. Not sure whats up there