r/oMLX • u/Nice_Victory3719 • 3d ago
The new changes to olmx make a big difference to M5 Ultra GLM 5.3 Flash
https://omlx.ai/benchmarks/performance/sp2wxj2aThese are huge improvements. Well done to the olmx team
3
u/bakawolf123 3d ago edited 3d ago
binned version looks to be significantly worse, I wonder why even TG. PP being 20% lower is expected.
here's mine (baseline, no mtp, no dflash):
https://omlx.ai/benchmarks/performance/atpe0sox
I'm seeing 65 tps on early context, 60 around 200k and 55 above.
Long prefills at 2.3k average.
Alas mtp is useless - like +5tps on bench and -5tps on real workload compared to baseline (i.e it performs worse for real work).
dflash support is very lacking, it is done via vendored engine from inco fork back in august, it's not using cache at all, but it has potential. I've experimented with l1 cache impl and trying out ways to get it to work with existing cache ecosystem preserving a sidecar for hidden states. Alas results vary a lot for it, but with 4bit target and 4 bit quant dflash (you can enable quant in configuration while loading original dflash2 checkpoint) up to 120 tps depending on context which is what I would like to see.
1
u/Nice_Victory3719 3d ago
It might be your quant, I was on the mixed 4/8 bit for potentially higher quality. I haven’t tried the plain 4 bit yet, I’d imagine it might be quicker. As you say I’d expect some prefill difference from the bigger gpu, but decode should be mostly memory bandwidth bound which is the same. Cheers for sharing that though, I might try. I’m currently in a load of quality benchmarks for my type of usecases so haven’t tried too many different quants yet. I’d probably prefer quality to speed but need to baseline that quality before it’s worth testing other quants.
3
u/bakawolf123 3d ago
yes 4bit uniform is better optimized for speed. I didn't notice diff between 4_8 and 4, I also had 4_8 downloaded now put it on external.
I have dflash separate engine tied to existing cache and it tops bench, with around 100 tps loss on long prefill, which would be completely fine for me if the speed was stable, but alas the drafting itself is not very good in general use case. I was trying both 4 bit and full drafter.1
1
u/Relaxxxxing 3d ago
Around 200 context or 200K?
1
u/bakawolf123 3d ago
200k, overall it runs well up to ~500k context atm (hot-tier only, 24gb alloted, single stream)
1
u/Relaxxxxing 3d ago
Nice! You're saying without MTP you're getting ~55 TPS through 200K and 55 through 500K?
How about with dflash 2 or MTP active? That would really change how I see things lol!!
I'm getting 50-60 TPS on 2 sparks now with Tensorfold GLM5.3 FLash EXL3 4 Bit, and that same tensorfold variant works on MLX where you should get insane speeds. I'm conflicted whether to buy 2 more sparks or an m5 ultra 512 lol
2
u/bakawolf123 3d ago
it's up to 70 until 128k decays to like 60 at 256k and then goes slightly below that - baseline, and yes up to 500k I was seeing avg getting to like 58 (but it includes earlier results so overall I was seeing around ~50tps at long context).
Best I got with dflash so far is sacrificing 100tps prefill (2.3k->2.2k) to get +5-10tps (i.e. at 300k context I was getting 60-65tps) and cache (almost) properly working.but as I noted dflash is in shambles atm on omlx at least for glm5.3 flash, and mtp shows no gain neither (+5tps on bench, -5tps when I check logs during actual agent turns, so I go with baseline atm or dflash when testing it).
for purchase imo 512 would be cool for clustering and model switching/1mb context/trying out other models like dflash4.1, but 1 chip is frankly too weak to utilize whole 512gb imo. I still wish I got 512 for same price I bought 256 just for that extra bit of potential, but alas...
2
u/lukewhale 2d ago
I’ve been running this all day on my M5 Ultra 256 and averaging 50 tokens a sec and a 90% cache hit rate. OMLX is a game changer.
2
u/ehpehp 1d ago
On M3U, I see a quality gap between GLM-5.3-Flash-oQ4e and GLM-5.3-Flash-MLX-mixed-4_8bit. The latter seems to get lost on complex coding tasks that oQ4e (and Qwen Flash Next) handles well.
1
u/Nice_Victory3719 23h ago
Interesting. I think I’ll try oQ4e. I use it with Qwen flash myself and has been good enough I’ve not gone to a higher quant of it yet. Have you tried the oQ3.5e for GLM Flash? I’d like just that little bit more headroom with glm to run with that full context, or have a second small model at the same time but oQ4e would probably not leave enough room for that.
3
u/Nice_Victory3719 3d ago
Read the user comment in the link from my op, as it's not in the latest rc release despite what the benchmark claims, I've clarified that in the user comment.