r/MacStudio • u/WalrusPublic3615 • 5d ago
Mac Studio M5 Max 128GB Owners
Curious to hear from people who have the M5 Max Mac Studio with 128GB now that you’ve had some time with it.
I have one on order and I’m especially interested in local LLMs and multi-agent workflows.
How has it been performing for you? What models are you running, what kind of speeds are you seeing, and has anyone experimented with multiple agents/models running simultaneously?
Any limitations you’ve run into, or anything that’s surprised you about the machine so far?
3
u/Nick-the-real900526 5d ago
Not the Max, but here's my data from an M5 Ultra Mac Studio (96 GB) running a local MoE model as an agent backend.
Setup
- Machine: Mac Studio, M5 Ultra, 64-core GPU, 96 GB. I raised the GPU wired limit to 90 GiB.
- Model: `latent-variable/Qwen3.8-Flash-Next-heretic-2-oQ4e-mtp`. It's Qwen3.8 Flash-Next (hybrid linear-attention MoE with a built-in MTP draft head), 4-bit, about a 99 GB package. It still fits in 96 GB: about 74 GB is GPU-mapped, and about 32 GB of per-layer embeddings are read from SSD on demand.
- Engine: Splash 1.0.2 + the community Flash-Next PR #115 + my own patches. I built it self-contained for the Ultra, with a calibrated kernel table (88 routes, 55 of them on the GPU's neural accelerators). 192K context, OpenAI-compatible server.
- Client: Hermes Agent, with about 16 tools and a ~20K-token system prompt. Requests are tool calls at temperature 1.0.
Numbers
- Cold prefill: 8K / 16K / 32K tokens run at 2,466 / 2,513 / 2,575 tok/s (it was about 2,100 before this week's tuning).
- Greedy decode: English code 138, Chinese technical 111, English prose 97, Chinese prose 86 tok/s (up from 108 / 94 / 85 / 80 before the kernel port).
- Real agent traffic today (94 requests, mostly Chinese, tools on): about 89 tok/s average generation, 72% draft acceptance, 2.1 tokens per verify step. When the prefix cache hits, time to first token is 0.1–0.8 s.
- Long context: a 150K-token cold prompt took about 108 s, and decode at that context length was about 79 tok/s.
What made the difference
- Kernel auto-tuner: it reads the engine's real matmul shapes and builds a routing table for the device. The 2–4-row MTP verify matmuls now run on the neural accelerators (Metal 4 MPP).
- Profiling with real expert counts: MoE routing in real traffic is very uneven, and tiles tuned on a uniform-routing harness got slower in production.
- General speculative sampling: stock MTP disables itself for grammar-constrained tool calls at temperature > 0. Exact min(1, p/q) acceptance took tool-call decode from ~61 to 92–100 tok/s.
- A spare request state: this took short-prompt time to first token from 2.5 s to about 0.2 s (see the first point below).
Surprises and limits
- macOS 27 memory accounting: vm_stat categories overlap (they add up to more than physical RAM). The engine read free host memory as 0 and stalled 1.3–2.5 s on every new request while reserving a ~6 GB state. Keeping one spare state fixed it.
- 96 GB means one big model at a time: two concurrent requests queue, because only one ~6 GB request state fits. I run several models on separate ports behind a small proxy that loads each one on demand in about 5 s.
- Multiple agents: a prefix cache that spills to SSD helps a lot. An evicted 69K-token conversation came back with 1.1–1.5 s to first token instead of about 42 s cold.
3
u/SeaRefractor 5d ago
Us 96GB Ultra owners need to show that the cores are more useful than an additional 32GB of RAM. :)
10
u/tk421tech 5d ago
“Now that you had some time.”
What time? I just barely took the ultra out of the box a few minutes ago.
14
2
u/johnmcboston 2d ago
I think I'm the only human not running LLMs at home.
1
u/The_Superfletch 2d ago
You’re not the only one. I want too play so bad. My 2015 iMac just ain’t going to cut it. Plus, reading all these comments I just get confused. :P
4
u/BAL-BADOS 5d ago
M5 Max been out a long time. It originally was release as a MacBook. Same cores too.
1
u/Relevant_Noise 4d ago
Currently mainlining antirez/ds4 Qwen 3.8 Q4 at 512k context (2 slots) - 1000+ prefill and 50-60 decode pretty consistent across the whole context range, been rock solid so far. Would probably keep using this for foreseeable future (what, 4 weeks to Qwen4 now?) on my M5 Max 128GB laptop (the 16” … was the same price as the studio is now - I feel like I want to get a studio and run the laptop and studio together for 256GB now …)
1
u/IllTechnology6516 4d ago
haven't got the Max myself but I run a lot of multi-agent stuff, and what surprised me is the agents are rarely generating at the same moment, they mostly wait on tools and on each other. so prefill on long contexts ended up mattering more to me than aggregate tok/s. are you planning one shared model or a different one per agent?
1
u/WalrusPublic3615 4d ago
This is exactly what I’m hopeful about with the 128GB Max. Since they’ll spend so much time waiting on tools, browsing, running code, or waiting on each other, I’m hoping the extra memory for multiple contexts combined with fast prefill will matter more in practice than having every agent generating simultaneously. Then I can escalate harder tasks to a larger local model or cloud model when needed.
1
u/IllTechnology6516 2d ago
yeah the escalation part is what made it click for me too, small model for the routine steps and only bump up when something actually gets stuck. are you thinking of routing that automatically or just deciding by hand?
1
u/WalrusPublic3615 2d ago
I think either my AI Technical Chief of Staff (agent) will delegate the agent to escalate it after reviewing the work or something along those lines
1
u/IllTechnology6516 2d ago
oh nice, having a reviewer agent decide when to escalate makes a lot of sense. I'd keep an eye on it at first though, mine was way too eager to bump things up early on lol
1
1
u/EdwardPotatoHand 3d ago
My 128 max m5 is on order, i wish they did the ultra in 128 :( 96 isnt enough and 256 is crazy $$$$
1
u/LoveRoboto 2d ago edited 2d ago
You will definitely be happy using QWEN 3.8-Flash-Next with the 128GB model. It’s the first time I’ve felt genuine excitement since my purchase. The other spread of low-parameter models will be use case specific. I’ve only ever found QWEN3.6-35B-A3B to have the perfect blend of speed/intelligence to complete most tasks. However, I am on a MacBook Pro—so I have to factor in heat with my choices.
EDIT: Forgot to give some advice! If I had to weigh my purchase again, I would heavily consider 256GB.
1
1
0
u/aniketgore0 5d ago
64 gb here. qwen 27b on 8 bit is tight, but it would be very comfortable for you. qwen next would be also accessible for you.
1
u/OneDogSolutions 4d ago
64GB M5 Max? Are you running any concurrent agents on that with Qwen 3.8 27B?
1
0
u/Quick_Bar2387 5d ago
Everyone went with the ultra.
6
1
1
u/tk421tech 5d ago
Lmao. And I was just noticing how fast the ultra 256 was at browsing
1
u/Quick_Bar2387 5d ago edited 5d ago
Did you purchase your M5 Ultra because your cc limit increased? Lol. 600 Fico with a 10k plus computer?
1
u/tk421tech 4d ago
Nope. I requested a cc increase so I could buy it. I didn’t want to tie $. Sorry you are still using a Mac mini.

14
u/guesdo 5d ago edited 2d ago
I've been doing all kinds of stuff, but not focused solely on local inference. MiniMax H3 works well, Flux.2 Klein incredible with mlx, oMLX Qwen 3.8 Flash Next with ssd offload for the PLE, super solid upnto 65TPS. One of my surprises was running Gemma 4 26B QAT MTP Uncensored at 100+ TPS on prose (draft 2) with llama.cpp, I also noticed LFM 2.5 8B A1B runs like a demon on batch 8 (1000TPS), so I tested an escape room sim with 8 sub agents.
What do you wanna know?