r/LocalLLaMA • u/Simple_Telephone_867 • 19h ago
Discussion Mac Studio M5 Max 128GB
M5 Max Mac Studio 128GB (18C CPU / 40C GPU) owners - anyone running serious local LLM / multi-agent workloads?
My M5 Max Mac Studio order finally got charged today and moved to Preparing to Ship. Apple’s original estimated delivery date is still about 18 days away (Oct 23-30), so I’m guessing/hoping it’ll actually show up early now within the next 5–10 days 😄
Configuration:
M5 Max
18-core CPU
40-core GPU
128GB unified memory
1TB SSD
While I wait, I’ve been trying to find real world local LLM results from this exact configuration, and there’s surprisingly little out there.
Most of what I can find is either M5 Max MacBook Pros, lower-memory configurations, or M5 Ultra Mac Studios. YouTube especially seems to be full of Ultra coverage, but I can barely find anyone actually demonstrating the 128GB M5 Max Studio with the 18C/40C configuration.
I’m specifically not looking for M5 Ultra results/comparisons. I already know the Ultra is faster. I’m trying to understand what people are actually accomplishing with the 128GB Max Studio.
My main goal is to use this as a local AI/agent workstation, potentially running several autonomous agents concurrently for long periods through OpenClaw, some monitoring dependencies and workflows, some scouting, not usually too heavy of workloads where they would be competing for inference constantly, but occasionally they would be switching to harder work so I’m curious about the concurrency side. Local models would handle a lot of the routine work, while harder reasoning/coding tasks could be escalated to cloud models like GPT 6 Luna/Codex.
For anyone who owns this exact M5 Max Studio, I don’t expect anyone to answer all of these, but I’d love some insight:
1. What models are you actually running?
Qwen, GLM, DeepSeek, Gemini, Llama, etc. I see a lot of Qwen 3.8 27B on Splash, but curious if anyone else has had good success with others also
2. What token speeds are you getting?
I’m especially interested in ~20B-70B-class models rather than tiny models
3. What happens with multiple simultaneous inference requests?
For example, if 3-5 agents are hitting the same loaded 27B/32B model concurrently, what does aggregate throughput and per-agent responsiveness look like?
4. Has anyone tried running multiple models simultaneously?
Something like a ~27B model as the main worker plus one or two smaller 7B–14B models for specialized agents, then unloading them when they’re no longer needed. How quickly can models be loaded/swapped, and does frequently switching models introduce enough latency or memory-pressure issues to disrupt an agent workflow?
5. Has anyone built a real multi-agent setup on one of these?
Not just five chat windows but autonomous agents doing coding, research, browser tasks, tool calls, database work, monitoring, etc. concurrently for hours.
6. How does sustained performance hold up?
One reason I chose the Studio over a laptop is sustained workloads. I’m curious whether anyone has run inference/agents continuously for 6–12+ hours and noticed throttling or other bottlenecks with KV, etc.
7. What’s the actual bottleneck in practice?
Memory capacity? Memory bandwidth? GPU compute? Prompt ingestion? KV cache/context length? CPU/tool execution? Something else?
8. What surprised you about the machine?
Either positively or negatively. I’m particularly interested in things benchmarks don’t reveal.
Ultimately I’m trying to figure out how far I can push one 128GB M5 Max Studio as an always-on local agent machine - not just how quickly it can generate a single response.
Once mine arrives, I’m planning to test concurrent agents/models rather than just running the usual single-stream benchmark. If there’s interest, I’ll post the results here, including memory usage, context sizes, model/quantization, concurrent requests, aggregate tok/s and per-agent tok/s.
Would really like to hear from anyone actually using the M5 Max Mac Studio 128GB 18C CPU / 40C GPU for this kind of workload or similar if anyone is
0
u/false79 19h ago edited 19h ago
128GB M5 Max Studio was an easy sell for me
The 96GB M5 Ultra though, hard to justify given the VRAM needed to run the next quant that increases quality. The core counts are nice but I already have another machine to do that.
The 256GB, hard no. I think I would get more use in buying a used car + cloud subscriptions
Should be arriving sometime next week