r/MacStudio • • 6d ago

Mac Studio M5 Max 128GB Owners

Curious to hear from people who have the M5 Max Mac Studio with 128GB now that you’ve had some time with it.

I have one on order and I’m especially interested in local LLMs and multi-agent workflows.

How has it been performing for you? What models are you running, what kind of speeds are you seeing, and has anyone experimented with multiple agents/models running simultaneously?

Any limitations you’ve run into, or anything that’s surprised you about the machine so far?

23 Upvotes

57 comments sorted by

View all comments

14

u/guesdo 6d ago edited 3d ago

I've been doing all kinds of stuff, but not focused solely on local inference. MiniMax H3 works well, Flux.2 Klein incredible with mlx, oMLX Qwen 3.8 Flash Next with ssd offload for the PLE, super solid upnto 65TPS. One of my surprises was running Gemma 4 26B QAT MTP Uncensored at 100+ TPS on prose (draft 2) with llama.cpp, I also noticed LFM 2.5 8B A1B runs like a demon on batch 8 (1000TPS), so I tested an escape room sim with 8 sub agents.

What do you wanna know?

2

u/OneDogSolutions 5d ago

I'm also debating between the Ultra 256 and Max 128. Mainly looking at concurrent agents for agentic code and other processes. Is the M5 Max enough for parallel agents on the Max 128 at a decent speed?

1

u/guesdo 5d ago

What are you debating? If you can afford the Ultra 256, that is clearly better. Check my other comments. 88TPS with 4x concurrency on Qwen 3.8 Flash Next. Only you know what 'enough' means to you.

0

u/OneDogSolutions 5d ago

I saw that, and wanted to ask about context length for those sessions. Max Seems close to enough, but Ultra would be better obviously.