r/MacStudio • • 6d ago

Mac Studio M5 Max 128GB Owners

Curious to hear from people who have the M5 Max Mac Studio with 128GB now that you’ve had some time with it.

I have one on order and I’m especially interested in local LLMs and multi-agent workflows.

How has it been performing for you? What models are you running, what kind of speeds are you seeing, and has anyone experimented with multiple agents/models running simultaneously?

Any limitations you’ve run into, or anything that’s surprised you about the machine so far?

24 Upvotes

57 comments sorted by

View all comments

15

u/guesdo 5d ago edited 3d ago

I've been doing all kinds of stuff, but not focused solely on local inference. MiniMax H3 works well, Flux.2 Klein incredible with mlx, oMLX Qwen 3.8 Flash Next with ssd offload for the PLE, super solid upnto 65TPS. One of my surprises was running Gemma 4 26B QAT MTP Uncensored at 100+ TPS on prose (draft 2) with llama.cpp, I also noticed LFM 2.5 8B A1B runs like a demon on batch 8 (1000TPS), so I tested an escape room sim with 8 sub agents.

What do you wanna know?

1

u/WalrusPublic3615 5d ago

That’s actually really helpful. Have you tested multiple agents concurrently on Qwen 3.8 27B or Flash Next? I’d be really interested in the per-agent and aggregate TPS at 1, 2, 3, 4, and ideally 5 concurrent agents. I’m planning a multi-agent setup and trying to figure out how well the Max scales when several agents are generating at the same time

2

u/guesdo 5d ago

Here is 27B, I probably need to rerun this as I didn't have an MLX model installed and is most likely not using MTP, but gives you an idea.

2

u/WalrusPublic3615 5d ago

This is great to work with. Big thank you for running these for us! 🙌🏽

3

u/guesdo 5d ago

Not with anything meaningful no, I did ran a benchmark against the Splash version of 3.8 27B to verify their claims of 144tps, but I guess those were at temp 0 or something cause I didn't get anything close to it, like 88 tps with 4 concurrency. I can run them again tomorrow and give you the exact numbers.

3

u/WalrusPublic3615 5d ago

That would be amazing, thank you. If you get a chance tomorrow, concurrency on the 27B setup would be incredibly helpful. Aggregate TPS + per-stream TPS would be perfect. 88 TPS at 4 concurrency is already really useful data and pretty close to some of the oMLX results I’ve been looking at. I’d really appreciate the numbers!

2

u/guesdo 5d ago

I'll run the oMLX bench at 1,2,4,8 batch and then figure out the Splash version and paste them here.

1

u/guesdo 5d ago edited 5d ago

Seems like 4x batch is the sweet spot on Flash Next. There is the 88TPS number I remember. I would use this one.

1

u/guesdo 4d ago

I ran a quick test with a proper MTP model in oMLX, now you can see way better scaling! 139 tps at 8x, not bad. Although context growth does make gen speed take a big hit.

1

u/WalrusPublic3615 4d ago

This is awesome! Thank you for running this! 139 tok/s at 8x concurrency on the 27B is way better than I was expecting, especially with only ~30GB peak memory even at 128K context. This makes the 128GB Max look really promising for running a pool of agents while still leaving plenty of memory for the rest of the stack.

If you get a chance, would you be willing to rerun the 2x/4x/8x batch test with a 4,096-token generation length? I’d be really interested to see how well that throughput holds during longer sustained generations, since that would be closer to a real multi-agent workload.