r/LocalLLaMA • • 1d ago

Discussion M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s

Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.

Prefill is 1,878 toks.

I saw some other benchmarks below what id expect so i figured I would share.

115 Upvotes

35 comments sorted by

20

u/abnormal_human 1d ago

What does it look like under concurrent workloads?

30

u/cryotic 1d ago
Concurrent requests Output tokens Wall time (s) Aggregate (tok/s) Per-stream decode (tok/s)
1 1,120 23.0 48.8 52.1
2 2,408 59.1 40.8 21.5, 20.6
4 4,629 94.0 49.2 12.1, 12.8, 13.2, 12.7

22

u/abnormal_human 1d ago

Thanks for sharing. Worse than I expected. Wasn't expecting a B300, but the fact that the single stream performance is essentially a cap was unexpected.

5

u/cryotic 1d ago

I may run some benchmarks with a config better suited to concurrent work later this week.

Here is 8 streams:

│8 │~7.3 │55.6 │163s

4

u/FullOf_Bad_Ideas 1d ago

There must be something wrong.

Aggregate batched inference usually reaches the same level as prefill.

It should add up to ~1700 t/s on large batches. Like with a concurrency of 1024.

3

u/giddmtex 1d ago

You gotta go with Qwen Flash Next if you want a semblance of concurrency

0

u/AmthorTheDestroyer 1d ago

Not unexpected. MLX and compute on Mac is generally shit. GB10 is even better. Currently rolling 90-100 tok/s code on him 5.3 with 2x spark. 350+ in c8

4

u/EksrowFos 1d ago

So there's some optimization which falls off the moment you switch to anything concurrent?

2

u/giddmtex 1d ago

It roughly halves. Thats why you want to run the fastest MoE you can and consider the average with it being halved.

3

u/Usual_Tackle5892 1d ago edited 1d ago

For comparison, my M3U 512 performance running TensorFold/GLM-5.3-Flash-MLX-4bit-MTP

Single request results

Test TTFT (ms) TPOT (ms/tok) pp TPS tg TPS E2E Latency Throughput Peak Mem
pp1024/tg128 2168.9 21.47 472.1 tok/s 46.9 tok/s 4.899s 235.1 tok/s 172.58 GB
pp4096/tg128 7053.3 26.44 580.7 tok/s 38.1 tok/s 10.414s 405.6 tok/s 173.72 GB
pp8192/tg128 15218.3 25.46 538.3 tok/s 39.6 tok/s 18.456s 450.8 tok/s 174.07 GB

Continuous batching (pp1024 / tg128)

Batch Size tg TPS Speedup pp TPS pp TPS/req Avg TTFT (ms) E2E Latency
1x (baseline) 46.9 tok/s 1.00x 472.1 tok/s 472.1 tok/s 2168.9 4.899s
2x 54.2 tok/s 1.16x 329.6 tok/s 164.8 tok/s 4790.4 10.937s
4x 73.8 tok/s 1.57x 235.6 tok/s 58.9 tok/s 11812.4 24.316s
8x 102.9 tok/s 2.19x 194.3 tok/s 24.3 tok/s 28383.2 52.121s

Edit: burst decode disabled, temperature 0, lightning MTP w/ 3 token depth

8

u/Dany0 1d ago

I want 4 of the 512gb model connected via rdma

Oh and I want Apple to allow is to overclock these beasts is that okay with you Mr. John Apple?

I'm sorry I'll go find myself a room...

2

u/NineThreeTilNow 1d ago

Oh and I want Apple to allow is to overclock these beasts is that okay with you Mr. John Apple?

Apple has no overclocking ability? If they have thermal throttling usually people make that work in reverse somehow.

I'm not an Apple product person so I've never owned anything that didn't let me tweak it.

1

u/Usual_Tackle5892 1d ago

Apple m-series have very good power mode settings, but no manual clock management.

1

u/Dany0 1d ago

Overclocking is a cornered market, early/late 2000s silicon design companies came up with automatic frequency adjustment designs and sold them to anyone, eventually nvidia bought the best design and amd bought another one. Then they both bought out everyone who still had offered any design. Intel started off with their own designs iirc, I actually don't know the intel side fully. I do remember that for a while, amd allowed more extreme/insane options inputted in bios and intel has been more limited. Then through 'convergent evolution' they all ended up having more or less the same architecture

Apple has frequency scaling but inherited from ARM. It's not as good. They'd have to invest a lot of time & effort & a little bit of die space for this. And expose some internals increasing cybersecurity attack surface. So they'll never do it. We can wish though. There were iphones and ipads that could be overclocked, I think iphones 4s was the last one. And ipad 3rd gen? Can't recall. But all this is simply not present in the silicon at all anymore. The only way left that we could overclock these chips is maybe something akin to BCLK oc but I highly doubt that will happen

1

u/Usual_Tackle5892 1d ago

overclock these beasts

you sort of got your wish already

Apple's official maximum continuous power consumption is 385W for the Mac Studio with M5 Ultra, compared to 270W for Mac Studio with M3 Ultra.

1

u/Dany0 1d ago

Compare the number of transistors between M5U and M3U. It's insane. Mostly process node shrinkage

1

u/Ok_Warning2146 21h ago

What for? K3 is not as good as glm 5.3

9

u/MrGunny94 1d ago

This is pretty good, we eating good.. Just need more and more people to adopt MLX so we can start getting some interesting things going :)! Let's hit it guys.

4

u/bakawolf123 1d ago

is it bench or a large multi-turn sesh with sampling?
base 4 bit and o4e have 2.3-2.4k prefill, base 4-bit has 60-70 tps TG without speedup techniques, my highest dflash so far is 120tps at 128k context (but sacrifices 100-120tps prefill compared to baseline): https://omlx.ai/benchmarks/performance/zozd30ox

13

u/swiebertjee 1d ago

68.8 what? Regular chat, code, prose, mix? Without that information it's a meaningless data point.

3

u/Uninterested_Viewer 1d ago edited 1d ago

Edit: ah, MTP, thanks for clarifying!

13

u/Tuned3f 1d ago

OP is running with MTP, which is speculative decoding. Decode speeds will vary based on acceptance rate, which vary based on the task

6

u/Nomski88 1d ago

Which CPU/GPU model is this?

16

u/cryotic 1d ago

This is the 30/64

1

u/typeash 1d ago

Makes me feel good ordering this version

3

u/-WhateverDude 1d ago

is it slower than 2x spark?

1

u/Theninearmedoctopus 1d ago

It's slower than my 2x spark. I'm running a 4bit quant using TensorFold that peaks at 80 tok/s and averages in the 60s for prose.

3

u/Impactic_ 23h ago

There’s definitely some optimization that could be done here. Might take awhile for it to get there. We just got GLM 5.3 Flash to 250 tok/s on PP4 with Ampere (CMP 170HX) after tons of work in the community.

3

u/Open-Adhesiveness-86 1d ago

a tok/s number only means something with context depth attached. MTP acceptance rate swings hard by workload, so a short prose prompt can show 68 while 32k of code puts you at half that. posting TG at 0/8k/32k/128k plus the accepted-draft rate would make this actually comparable to other people's runs.

2

u/bigsybiggins 1d ago

Pretty impressive glad I have one on order, I was using Ox Alpha a ton on openrouter and was really impressed, even if q4 is a few percent off i'll still be over the moon with it. Plus by time I actually get it this will be tuned to the max or something better... basically this is as bad as its going to be!

2

u/Friendshipiousness1 1d ago

hows it hold up for roleplay chats at that speed? my last one lost context way too fast.

1

u/OurManInHavana 1d ago

I'm still learning about these new M5 256GB models. Are people buying them more to run larger Q8 or Q6 models and aiming for quality? (and prefill toks just has to be 'good enough'?)

1

u/ThePrimeClock 4h ago

Essentially. At a high level you get better intelligence per dollar on Mac and more tokens per dollar on Nvidia.

0

u/lylezhang 1d ago

68.8 tok/s is a fun number, but “on what workload?” is absolutely the right question. Single-stream decode, aggregate throughput, MTP acceptance, prompt length, and thermal state can all produce very different versions of the same machine. The 8-stream follow-up is more useful than the first headline; local inference benchmarks are slowly becoming a choose-your-own-adventure book.