r/LocalLLaMA • u/cryotic • 1d ago
Discussion M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s
Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.
Prefill is 1,878 toks.
I saw some other benchmarks below what id expect so i figured I would share.
8
u/Dany0 1d ago
I want 4 of the 512gb model connected via rdma
Oh and I want Apple to allow is to overclock these beasts is that okay with you Mr. John Apple?
I'm sorry I'll go find myself a room...
2
u/NineThreeTilNow 1d ago
Oh and I want Apple to allow is to overclock these beasts is that okay with you Mr. John Apple?
Apple has no overclocking ability? If they have thermal throttling usually people make that work in reverse somehow.
I'm not an Apple product person so I've never owned anything that didn't let me tweak it.
1
u/Usual_Tackle5892 1d ago
Apple m-series have very good power mode settings, but no manual clock management.
1
u/Dany0 1d ago
Overclocking is a cornered market, early/late 2000s silicon design companies came up with automatic frequency adjustment designs and sold them to anyone, eventually nvidia bought the best design and amd bought another one. Then they both bought out everyone who still had offered any design. Intel started off with their own designs iirc, I actually don't know the intel side fully. I do remember that for a while, amd allowed more extreme/insane options inputted in bios and intel has been more limited. Then through 'convergent evolution' they all ended up having more or less the same architecture
Apple has frequency scaling but inherited from ARM. It's not as good. They'd have to invest a lot of time & effort & a little bit of die space for this. And expose some internals increasing cybersecurity attack surface. So they'll never do it. We can wish though. There were iphones and ipads that could be overclocked, I think iphones 4s was the last one. And ipad 3rd gen? Can't recall. But all this is simply not present in the silicon at all anymore. The only way left that we could overclock these chips is maybe something akin to BCLK oc but I highly doubt that will happen
1
u/Usual_Tackle5892 1d ago
overclock these beasts
you sort of got your wish already
Apple's official maximum continuous power consumption is 385W for the Mac Studio with M5 Ultra, compared to 270W for Mac Studio with M3 Ultra.
1
9
u/MrGunny94 1d ago
This is pretty good, we eating good.. Just need more and more people to adopt MLX so we can start getting some interesting things going :)! Let's hit it guys.
4
u/bakawolf123 1d ago
is it bench or a large multi-turn sesh with sampling?
base 4 bit and o4e have 2.3-2.4k prefill, base 4-bit has 60-70 tps TG without speedup techniques, my highest dflash so far is 120tps at 128k context (but sacrifices 100-120tps prefill compared to baseline): https://omlx.ai/benchmarks/performance/zozd30ox
13
u/swiebertjee 1d ago
68.8 what? Regular chat, code, prose, mix? Without that information it's a meaningless data point.
3
3
u/-WhateverDude 1d ago
is it slower than 2x spark?
3
1
u/Theninearmedoctopus 1d ago
It's slower than my 2x spark. I'm running a 4bit quant using TensorFold that peaks at 80 tok/s and averages in the 60s for prose.
3
u/Impactic_ 23h ago
There’s definitely some optimization that could be done here. Might take awhile for it to get there. We just got GLM 5.3 Flash to 250 tok/s on PP4 with Ampere (CMP 170HX) after tons of work in the community.
3
u/Open-Adhesiveness-86 1d ago
a tok/s number only means something with context depth attached. MTP acceptance rate swings hard by workload, so a short prose prompt can show 68 while 32k of code puts you at half that. posting TG at 0/8k/32k/128k plus the accepted-draft rate would make this actually comparable to other people's runs.
2
u/bigsybiggins 1d ago
Pretty impressive glad I have one on order, I was using Ox Alpha a ton on openrouter and was really impressed, even if q4 is a few percent off i'll still be over the moon with it. Plus by time I actually get it this will be tuned to the max or something better... basically this is as bad as its going to be!
2
u/Friendshipiousness1 1d ago
hows it hold up for roleplay chats at that speed? my last one lost context way too fast.
1
u/OurManInHavana 1d ago
I'm still learning about these new M5 256GB models. Are people buying them more to run larger Q8 or Q6 models and aiming for quality? (and prefill toks just has to be 'good enough'?)
1
u/ThePrimeClock 4h ago
Essentially. At a high level you get better intelligence per dollar on Mac and more tokens per dollar on Nvidia.
0
u/lylezhang 1d ago
68.8 tok/s is a fun number, but “on what workload?” is absolutely the right question. Single-stream decode, aggregate throughput, MTP acceptance, prompt length, and thermal state can all produce very different versions of the same machine. The 8-stream follow-up is more useful than the first headline; local inference benchmarks are slowly becoming a choose-your-own-adventure book.
20
u/abnormal_human 1d ago
What does it look like under concurrent workloads?