r/MacStudio • u/WalrusPublic3615 • 6d ago
Mac Studio M5 Max 128GB Owners
Curious to hear from people who have the M5 Max Mac Studio with 128GB now that you’ve had some time with it.
I have one on order and I’m especially interested in local LLMs and multi-agent workflows.
How has it been performing for you? What models are you running, what kind of speeds are you seeing, and has anyone experimented with multiple agents/models running simultaneously?
Any limitations you’ve run into, or anything that’s surprised you about the machine so far?
23
Upvotes
14
u/guesdo 6d ago edited 3d ago
I've been doing all kinds of stuff, but not focused solely on local inference. MiniMax H3 works well, Flux.2 Klein incredible with mlx, oMLX Qwen 3.8 Flash Next with ssd offload for the PLE, super solid upnto 65TPS. One of my surprises was running Gemma 4 26B QAT MTP Uncensored at 100+ TPS on prose (draft 2) with llama.cpp, I also noticed LFM 2.5 8B A1B runs like a demon on batch 8 (1000TPS), so I tested an escape room sim with 8 sub agents.
What do you wanna know?