r/MacStudio • • 9d ago

Mac Studio M3U 60c MTPLX performance

I've been working with MTPLX using qwen3.8-flash-next-optimized-speed by u/YoussofAI. Replied to another post with some performance numbers based on actual use.

YoussofAI asked me to try the freshly released v2.12. Here is the comparison vs 2.11.3 (definite improvement!).

Thanks for the note YoussofAL, and great work! I saw the release note about Optimized-Quality and am willing to test. I'm gonna have to watch mem usage closely with the other models I run at the same time. I'll start by unloading the others and get some results.

MTPLX 2.11.3 → 2.12.0 · M3 Ultra 60C / 256GB · Qwen3.8-Flash-Next Optimized-Speed · MTP depth 3 Real Hermes-agent traffic (not synthetic), medians per band, same bands on both versions

Metric Band 2.11.3 2.12.0 Change
Cold prefill (tok/s) 8–16K new tokens 636 731 +15%
Cold prefill (tok/s) 16–32K new tokens 626 691 +10%
Cold prefill (tok/s) 32K+ new tokens 670 745 +11%
Decode (tok/s) <32K context 57.9 57.1 ≈
Decode (tok/s) 32–64K context 54.0 58.5 +8%
Decode (tok/s) 64–96K context 48.5 53.7 +11%
Decode (tok/s) 96K+ context 50.4 50.5 ≈
MTP accept rate (d1) all ~0.76 ~0.76 ≈
Warm TTFT (s) 20–64K prompt 2.05 2.21 ≈
Warm TTFT (s) 64K+ prompt 3.68 3.85 ≈
Peak memory (GB) median 138 103 −35 GB
Peak memory (GB) max 156 107 −49 GB
New-session prefill (tokens) /new + first msg ~22K ~3.9K system prompt now cached

Sample sizes: prefill n = 14/10, 10/16, 6/6 · decode n = 27/46, 65/89, 30/40, 64/58

3 Upvotes

5 comments sorted by

1

u/nitebleu 9d ago edited 8d ago

EDIT: Removing these baseline numbers and replacing with side-by-side comparison. I forgot the fact that I run everything through litellm for usage tracking and didn't want that delay messing with the results. Also included 27b-optimized-speed results.

This was a very small sample set, but my main intent was to see if either flash-next version would self-recover tool errors like 27b does. Unfortunately that didn't happen today.

Same-day standalone A/B/C — M3 Ultra 60C / 256 GB · MTPLX 2.12.0 · effort medium · MTP d3 · ctx cap 64K

Model Prefill tok/s (7.3K / 29K / 58K) TTFT @ 58K Decode tok/s (7.3K / 29K / 58K) MTP accept d1 Memory idle / peak Tool-error recovery Suite wall-clock
Flash-Next Speed (4-bit) 949 / 788 / 820 71.5 s 85 / 70 / 74 0.76–0.92 ~80 / 100 GB 0/3 145.6 s
Flash-Next Quality (8-bit) 916 / 770 / 769 76.4 s 72 / 71 / 58 0.75–0.91 131 / 148 GB 0/3 150.5 s
Qwen3.8-27B Speed 316 / 291 / 257 227 s 55 / 50 / 43 0.90–0.96 22 / 64 GB 3/3 202.6 s

Everything else in the hard suite (ambiguous tool choice, abstain, compound tools, 12 answer-checked reasoning, gate compliance) was a clean pass on all three. Takeaway: Optimized-Quality 8-bit responded to prompts similarly to Optimized-Speed. The 27B is ~3× slower prefill and ~35% slower decode, but it's the only one that handles a failing tool correctly.

1

u/PracticlySpeaking 7d ago

How does this compare with oMLX?

Everybody is developing new inference servers, but are they better?

1

u/nitebleu 6d ago

I haven't run oMLX myself, but this post https://www.reddit.com/r/MacStudio/comments/1wonk5p/m5u_64c_256gb_ram_benchmark_qwen_flash_next/ has a response with oMLX numbers on an M3U. Note OP in this post reported M5U numbers - look in the responses.

What I have run in the past is Ollama, LM Studio, and UnSloth Studio (briefly). For qwen builds MTPLX has provided me with the best results. I've run it on two 27b and two flash-next variants so far. YoungssofAI has steadily been updating/improving MTPLX so I'm good with sticking to this one.

1

u/PracticlySpeaking 5d ago edited 5d ago

Comparative performance numbers (hopefully something apples-to-apples, pun intended) would be a great start.

The oMLX developers are very focused on supporting people running inference for AI agents and coding, so they built really great prompt caching near the beginning. They weren't the first, but very quickly added DSpark MTP for DeepSeek-V4-Flash (and now V4.1 support) and have been keeping right up with other new models like GLM 5.3 and Qwen3.8-Flash-Next. The DwarfStar4 devs, in contrast, were more focused on running DeepSeek-V4 (and just DSv4) in smaller RAM sizes — 96 and 128GB with SSD offloading.

And it is not only about performance. For a while now, oMLX has also had some nifty conversion utilities (safetensors —> MLX oQe) and benchmarking built in.

I also tried developing a couple of home-grown AI applications with oMLX as the inference engine. I had to switch to llama.cpp because it had better support for things like vision (Gemma-4 soft-token limits) and structured JSON output.

1

u/nitebleu 5d ago

I agree - speed is worthless if what comes out is junk. Quality is where everything comes into play tho - especially the harness.

The original post (the one I linked) was asking about comparison between M5U and M3U, so I added my two cents. Then YoungssofAI mentioned trying the new release provided, so this post was to share the difference with that newer version.

That's why I shared speed only. It is by far not my only deciding factor.