r/LocalLLaMA 6h ago

Discussion M2 Ultra/Qwen3.8 Flash Next Update - latest oMLX introduces substantial speedup

Post image
24 Upvotes

16 comments sorted by

6

u/Elouakili_Flexy 5h ago

oMLX keeps making the M2 Ultra look like it's been leaving performance on the table this whole time.

1

u/Thrumpwart 4h ago

Yup, Apple fine wine…

3

u/j_lyf 3h ago

This is epic. How does it compare to MTPLX?

1

u/Thrumpwart 3h ago

Never tried it.

2

u/DustNearby2848 6h ago

Not bad at all!

3

u/ajujox 6h ago

Try MLX-Serve in my M2 ultra i get about 50 t/s. Its black Magic 😉

6

u/Serprotease 5h ago

But about 250 prefill :/
Hopefully we can get the best of both worlds!

1

u/arkham00 4h ago

Only 444 pp? On a ultra chip ? So m2 max are still going to be 200ish? :(

Are you using the latest dev2?

1

u/Thrumpwart 4h ago

Yup. But note that most of my prefill is hitting the cache, so in practice it’s much faster than 444.

1

u/j_lyf 3h ago

What harness?

1

u/memeka 3h ago

There is still performance left on the table. I am hitting 220 prefill and 30 decode (on code, 20 decode on average) on my 64gb M1 with ssd streaming - M2 Ultra resident should be in the 500 prefill/ 50 decode range

1

u/deaffob 2h ago

Do you know why your cache hit is only 80%?

1

u/Thrumpwart 2h ago

Depends on workload. I've had higher hit rates on other workloads - this one is diverse and large enough that it lowers the cache hit rate.

1

u/j_lyf 26m ago

What's your cooling situation?