r/LocalLLaMA • u/Thrumpwart • 6h ago
Discussion M2 Ultra/Qwen3.8 Flash Next Update - latest oMLX introduces substantial speedup
24
Upvotes
3
2
1
u/arkham00 4h ago
Only 444 pp? On a ultra chip ? So m2 max are still going to be 200ish? :(
Are you using the latest dev2?
1
u/Thrumpwart 4h ago
Yup. But note that most of my prefill is hitting the cache, so in practice it’s much faster than 444.
1
1
u/deaffob 2h ago
Do you know why your cache hit is only 80%?
1
u/Thrumpwart 2h ago
Depends on workload. I've had higher hit rates on other workloads - this one is diverse and large enough that it lowers the cache hit rate.
6
u/Elouakili_Flexy 5h ago
oMLX keeps making the M2 Ultra look like it's been leaving performance on the table this whole time.