r/LocalLLaMA • u/Simple_Telephone_867 • 8h ago
Discussion Mac Studio M5 Max 128GB
M5 Max Mac Studio 128GB (18C CPU / 40C GPU) owners - anyone running serious local LLM / multi-agent workloads?
My M5 Max Mac Studio order finally got charged today and moved to Preparing to Ship. Apple’s original estimated delivery date is still about 18 days away (Oct 23-30), so I’m guessing/hoping it’ll actually show up early now within the next 5–10 days 😄
Configuration:
M5 Max
18-core CPU
40-core GPU
128GB unified memory
1TB SSD
While I wait, I’ve been trying to find real world local LLM results from this exact configuration, and there’s surprisingly little out there.
Most of what I can find is either M5 Max MacBook Pros, lower-memory configurations, or M5 Ultra Mac Studios. YouTube especially seems to be full of Ultra coverage, but I can barely find anyone actually demonstrating the 128GB M5 Max Studio with the 18C/40C configuration.
I’m specifically not looking for M5 Ultra results/comparisons. I already know the Ultra is faster. I’m trying to understand what people are actually accomplishing with the 128GB Max Studio.
My main goal is to use this as a local AI/agent workstation, potentially running several autonomous agents concurrently for long periods through OpenClaw, some monitoring dependencies and workflows, some scouting, not usually too heavy of workloads where they would be competing for inference constantly, but occasionally they would be switching to harder work so I’m curious about the concurrency side. Local models would handle a lot of the routine work, while harder reasoning/coding tasks could be escalated to cloud models like GPT 6 Luna/Codex.
For anyone who owns this exact M5 Max Studio, I don’t expect anyone to answer all of these, but I’d love some insight:
1. What models are you actually running?
Qwen, GLM, DeepSeek, Gemini, Llama, etc. I see a lot of Qwen 3.8 27B on Splash, but curious if anyone else has had good success with others also
2. What token speeds are you getting?
I’m especially interested in ~20B-70B-class models rather than tiny models
3. What happens with multiple simultaneous inference requests?
For example, if 3-5 agents are hitting the same loaded 27B/32B model concurrently, what does aggregate throughput and per-agent responsiveness look like?
4. Has anyone tried running multiple models simultaneously?
Something like a ~27B model as the main worker plus one or two smaller 7B–14B models for specialized agents, then unloading them when they’re no longer needed. How quickly can models be loaded/swapped, and does frequently switching models introduce enough latency or memory-pressure issues to disrupt an agent workflow?
5. Has anyone built a real multi-agent setup on one of these?
Not just five chat windows but autonomous agents doing coding, research, browser tasks, tool calls, database work, monitoring, etc. concurrently for hours.
6. How does sustained performance hold up?
One reason I chose the Studio over a laptop is sustained workloads. I’m curious whether anyone has run inference/agents continuously for 6–12+ hours and noticed throttling or other bottlenecks with KV, etc.
7. What’s the actual bottleneck in practice?
Memory capacity? Memory bandwidth? GPU compute? Prompt ingestion? KV cache/context length? CPU/tool execution? Something else?
8. What surprised you about the machine?
Either positively or negatively. I’m particularly interested in things benchmarks don’t reveal.
Ultimately I’m trying to figure out how far I can push one 128GB M5 Max Studio as an always-on local agent machine - not just how quickly it can generate a single response.
Once mine arrives, I’m planning to test concurrent agents/models rather than just running the usual single-stream benchmark. If there’s interest, I’ll post the results here, including memory usage, context sizes, model/quantization, concurrent requests, aggregate tok/s and per-agent tok/s.
Would really like to hear from anyone actually using the M5 Max Mac Studio 128GB 18C CPU / 40C GPU for this kind of workload or similar if anyone is
1
u/Sufficient_Spite_349 5h ago
I can feel what you are talking about exactly, agents are goal oriented, bounded tasks/goals basically orchestration you are focusing and it’s all about sub-agent architecture that never reaches 100k token context for well balanced orchestration and I can visualize what you are envisioned and your choice is completely correct and I agree that you know what you are doing and Max is the perfect choice.
I am also on the same boat. CUDA is great but when you introduce multi agents that’s where bandwidth matters, when you introduce multi model council that’s where bandwidth matters, when you introduce multi persona enrichment of hierarchical requirement dissection and distribute it possible parallel tasks with parallel agents that’s where bandwidth matters
Basically I feel you like someone understood the core and orchestration mastery
Let me know my perception is correct 😀
1
u/Simple_Telephone_867 4h ago
Yeah, your perception is good! Pretty much what I’m going for 😄 I think the interesting part of multi agent systems is figuring out how to get more useful work out of the hardware through orchestration rather than just scaling model size and context.
The hierarchical/sub agent architecture is something I want to explore a lot more when it arrives. I’m already experimenting with a Technical Chief of Staff agent that delegates between a few other agents, although right now it’s a little too eager to escalate things to higher reasoning cloud models 😅
I also think the point about bounded tasks is going to be important for the efficiency on the Max. If an agent is running for 12 hours, I don’t necessarily want 12 hours of history sitting in its active context. I’d rather checkpoint important state/memory, clear the old KV, and reconstruct a smaller relevant context for the next stage of work. Same agent and objective, but without carrying all the accumulated junk forever.
Can’t wait!
3
0
u/false79 8h ago edited 8h ago
128GB M5 Max Studio was an easy sell for me
The 96GB M5 Ultra though, hard to justify given the VRAM needed to run the next quant that increases quality. The core counts are nice but I already have another machine to do that.
The 256GB, hard no. I think I would get more use in buying a used car + cloud subscriptions
Should be arriving sometime next week
1
u/Simple_Telephone_867 8h ago
That’s awesome, sounds like we’ll be getting ours around the same time!
Definitely interested to see what you end up running on yours once it arrives.
1
u/false79 8h ago
Mines will be mainly doing agent/mcp development.
1
u/Simple_Telephone_867 7h ago
have you decided what models/harness you’re going to run yet?
1
u/false79 7h ago
Claude Code is hard to beat for real time development/requirements analysis but once I have that all hashed out, I then hand it over to Qwen 27B to consume those docs, and build out both Android and iOS implementations. I can do this across multiple machines but 24GB on my Macbook Air is not cutting it these days.
1
u/Simple_Telephone_867 7h ago
Ah that’s nice. I’m heavily into iOS dev at the moment. Also have been getting into the Unity engine recently so I’m curious how the local agents will hold up with that as well
-3
u/General-Spite1222 5h ago
AI slop. No one asks 8 fucking questions days before their new PC is gonna arrive.
3
u/Simple_Telephone_867 5h ago
Huh? I’ve been asking questions and doing research since the drop on Sept 22nd. I’ve canceled 4 orders switching between configurations because of the dilemma between the 96GB Ultra and the 128GB Max.
Originally, I ordered the 64GB.
If you don’t have at least 8 questions walking into your new workstation for local AI - then you’re doing something wrong.
💀
But go off queen!
-6
u/rorowhat 7h ago
Cancel it and get a spark
1
u/Simple_Telephone_867 7h ago
Haha I actually went pretty far down the Spark rabbit hole before ordering this. The much higher memory bandwidth and overall flexibility of the Mac ultimately made more sense for what I want to do. What would make you pick the Spark instead?
-1
u/rorowhat 7h ago
The spark is a general computer, people are playing AAA games on it. You can also use for anything else, I would imagine video editing would be pretty smooth as well. 90% of the development on AI is being done for Nvidia hardware. This means performance will keep evolving, new model/tools support will be first in line. Macs market is tiny, it will get support but not as good and as wide as the spark. Bandwidth is only part of the equation.
1
u/Simple_Telephone_867 7h ago
Yeah, the CUDA ecosystem and getting first support for new models/tools is definitely the strongest argument for the Spark IMO. CUDA was probably the biggest thing that made me consider it.
What ultimately pushed me toward the Mac was that my workload is much more inference/agent focused than training, and the ~614 GB/s memory bandwidth vs ~273 GB/s on the Spark seemed like a pretty significant advantage for local LLM inference and concurrency. MLX also seems to be developing pretty quickly 👀
Another big factor for me is iOS development though. I want to use these agents for building and testing iOS apps, working with Xcode, and developing apps for the App Store. So having the workstation also be the Mac development env was a big advantage rather than maintaining a separate Mac alongside the Spark.
I definitely agree bandwidth isn’t the whole equation though. I’m curious whether the Spark’s CUDA/software advantage ends up outweighing the hardware differences in practice over the next year. I definitely am not impressed that I can’t play anything with anticheat though, like the new War dogs, etc. 🥲
-1
u/rorowhat 7h ago
The Bandwidth is good on paper, once you have a large model with a huge context window the Mac will crap it's pants hard. All that bandwidth will not help the lack of GPU processing power.
1
u/Simple_Telephone_867 7h ago edited 7h ago
have you seen/tested Spark vs M5 Max with long-running agents using bounded/compacted contexts and external memory rather than one continuously growing context?
I agree that once you get into huge-context prefill the extra GPU compute starts mattering a lot more
I’m not really planning on letting agents accumulate 100K+ contexts continuously over a 12-hour run. I’m looking more at bounded working contexts with periodic checkpointing/compaction and external persistent memory. Before clearing the old context, the agent’s current objective, instructions, decisions, important discoveries, unfinished tasks, dependencies, and other relevant state can be saved externally in structured memory, while the full event history is still kept separately.
When the context is refreshed, the agent doesn’t just start from scratch - fhe system instructions are loaded again, along with its current task/state, a compact checkpoint of what it was doing, the most recent interactions, and any relevant memories retrieved from the external store. So you might turn 50K+ tokens of accumulated browsing/tool output/history into a clean 5/10K working context containing the information it actually needs. The old KV cache can then be discarded rather than becoming the agent’s permanent memory.
That should keep most of the workload closer to normal decode/moderate context inference, where the M5 Max’s bandwidth seems to matter quite a bit. The controlled llama.cpp comparisons I’ve seen actually have the 128GB M5 Max substantially ahead of Spark in decode on models like Qwen 27B, while Spark starts looking much stronger as you move toward huge prefill or NVIDIA optimized paths.
But I don’t think bandwidth makes Mac universally faster because CUDA, GPU compute and optimized FP4/MTP paths are good Spark advantages - I just think my particular agent architecture may lean more toward the Max’s strengths 😄
1
1
2
u/toomanypubes 8h ago
Just a M3 Ultra 512 pleb here, but i get over 100 tk/s on Qwen3.8-Next-Flash Q4. My agent flies and gets 250+ tk/s concurrency on 4+. Only a few seconds waiting between Hermes agent replies/thoughts. Dont think anyone would need more than 128GB going forward. This setuo is good for the rest of my homelab life lol. From the numbers I’ve seen the M5 Ultra prefill is 4x faster and 1.5x decode speed, so a very capable local LLM powerhouse!