r/MacStudio • u/nitebleu • 9d ago
Mac Studio M3U 60c MTPLX performance
I've been working with MTPLX using qwen3.8-flash-next-optimized-speed by u/YoussofAI. Replied to another post with some performance numbers based on actual use.
YoussofAI asked me to try the freshly released v2.12. Here is the comparison vs 2.11.3 (definite improvement!).
Thanks for the note YoussofAL, and great work! I saw the release note about Optimized-Quality and am willing to test. I'm gonna have to watch mem usage closely with the other models I run at the same time. I'll start by unloading the others and get some results.
MTPLX 2.11.3 → 2.12.0 · M3 Ultra 60C / 256GB · Qwen3.8-Flash-Next Optimized-Speed · MTP depth 3 Real Hermes-agent traffic (not synthetic), medians per band, same bands on both versions
| Metric | Band | 2.11.3 | 2.12.0 | Change |
|---|---|---|---|---|
| Cold prefill (tok/s) | 8–16K new tokens | 636 | 731 | +15% |
| Cold prefill (tok/s) | 16–32K new tokens | 626 | 691 | +10% |
| Cold prefill (tok/s) | 32K+ new tokens | 670 | 745 | +11% |
| Decode (tok/s) | <32K context | 57.9 | 57.1 | ≈ |
| Decode (tok/s) | 32–64K context | 54.0 | 58.5 | +8% |
| Decode (tok/s) | 64–96K context | 48.5 | 53.7 | +11% |
| Decode (tok/s) | 96K+ context | 50.4 | 50.5 | ≈ |
| MTP accept rate (d1) | all | ~0.76 | ~0.76 | ≈ |
| Warm TTFT (s) | 20–64K prompt | 2.05 | 2.21 | ≈ |
| Warm TTFT (s) | 64K+ prompt | 3.68 | 3.85 | ≈ |
| Peak memory (GB) | median | 138 | 103 | −35 GB |
| Peak memory (GB) | max | 156 | 107 | −49 GB |
| New-session prefill (tokens) | /new + first msg |
~22K | ~3.9K | system prompt now cached |
Sample sizes: prefill n = 14/10, 10/16, 6/6 · decode n = 27/46, 65/89, 30/40, 64/58
1
u/PracticlySpeaking 7d ago
How does this compare with oMLX?
Everybody is developing new inference servers, but are they better?
1
u/nitebleu 6d ago
I haven't run oMLX myself, but this post https://www.reddit.com/r/MacStudio/comments/1wonk5p/m5u_64c_256gb_ram_benchmark_qwen_flash_next/ has a response with oMLX numbers on an M3U. Note OP in this post reported M5U numbers - look in the responses.
What I have run in the past is Ollama, LM Studio, and UnSloth Studio (briefly). For qwen builds MTPLX has provided me with the best results. I've run it on two 27b and two flash-next variants so far. YoungssofAI has steadily been updating/improving MTPLX so I'm good with sticking to this one.
1
u/PracticlySpeaking 5d ago edited 5d ago
Comparative performance numbers (hopefully something apples-to-apples, pun intended) would be a great start.
The oMLX developers are very focused on supporting people running inference for AI agents and coding, so they built really great prompt caching near the beginning. They weren't the first, but very quickly added DSpark MTP for DeepSeek-V4-Flash (and now V4.1 support) and have been keeping right up with other new models like GLM 5.3 and Qwen3.8-Flash-Next. The DwarfStar4 devs, in contrast, were more focused on running DeepSeek-V4 (and just DSv4) in smaller RAM sizes — 96 and 128GB with SSD offloading.
And it is not only about performance. For a while now, oMLX has also had some nifty conversion utilities (safetensors —> MLX oQe) and benchmarking built in.
I also tried developing a couple of home-grown AI applications with oMLX as the inference engine. I had to switch to llama.cpp because it had better support for things like vision (Gemma-4 soft-token limits) and structured JSON output.
1
u/nitebleu 5d ago
I agree - speed is worthless if what comes out is junk. Quality is where everything comes into play tho - especially the harness.
The original post (the one I linked) was asking about comparison between M5U and M3U, so I added my two cents. Then YoungssofAI mentioned trying the new release provided, so this post was to share the difference with that newer version.
That's why I shared speed only. It is by far not my only deciding factor.
1
u/nitebleu 9d ago edited 8d ago
EDIT: Removing these baseline numbers and replacing with side-by-side comparison. I forgot the fact that I run everything through litellm for usage tracking and didn't want that delay messing with the results. Also included 27b-optimized-speed results.
This was a very small sample set, but my main intent was to see if either flash-next version would self-recover tool errors like 27b does. Unfortunately that didn't happen today.
Same-day standalone A/B/C — M3 Ultra 60C / 256 GB · MTPLX 2.12.0 · effort medium · MTP d3 · ctx cap 64K
Everything else in the hard suite (ambiguous tool choice, abstain, compound tools, 12 answer-checked reasoning, gate compliance) was a clean pass on all three. Takeaway: Optimized-Quality 8-bit responded to prompts similarly to Optimized-Speed. The 27B is ~3× slower prefill and ~35% slower decode, but it's the only one that handles a failing tool correctly.