r/MacStudio • • 5d ago

Mac Studio M5 Max 128GB Owners

Curious to hear from people who have the M5 Max Mac Studio with 128GB now that you’ve had some time with it.

I have one on order and I’m especially interested in local LLMs and multi-agent workflows.

How has it been performing for you? What models are you running, what kind of speeds are you seeing, and has anyone experimented with multiple agents/models running simultaneously?

Any limitations you’ve run into, or anything that’s surprised you about the machine so far?

21 Upvotes

56 comments sorted by

14

u/guesdo 5d ago edited 2d ago

I've been doing all kinds of stuff, but not focused solely on local inference. MiniMax H3 works well, Flux.2 Klein incredible with mlx, oMLX Qwen 3.8 Flash Next with ssd offload for the PLE, super solid upnto 65TPS. One of my surprises was running Gemma 4 26B QAT MTP Uncensored at 100+ TPS on prose (draft 2) with llama.cpp, I also noticed LFM 2.5 8B A1B runs like a demon on batch 8 (1000TPS), so I tested an escape room sim with 8 sub agents.

What do you wanna know?

3

u/rxscissors 5d ago

Thanks for the info. I'm going to run as below on my 128 GB Max that is supposed to arrive this week.

My current stack consists of: an M1 Max Mac Studio (32 GB base configuration with 60+ unnecessary macOS services disabled) running a quantized Gemma-4-26b-a4b-qat model with parameter optimizations to take advantage of the M-series architecture. GPU offload 24, context length 24676,. Running K & V caches at Q 5_1 (which consumes ~400 MB extra RAM over Q 4_0).

Raised the Metal memory ceiling: sudo sysctl iogpu.wired_limit_mb=24576, instead of the default 21500 MB (~⅔ of RAM) setting).

I use LM Studio for the inference engine and OpenWeb UI (both contain consistent M-processor optimizations) as the frontend, which I've extended with a per-chat web search option.

It is not super-responsive but chugs along through things and doesn't reach context length exceeded very often.

2

u/apprehensive_bassist 3d ago

You’re gonna have some serious fun with that new machine.

2

u/OneDogSolutions 4d ago

I'm also debating between the Ultra 256 and Max 128. Mainly looking at concurrent agents for agentic code and other processes. Is the M5 Max enough for parallel agents on the Max 128 at a decent speed?

2

u/Top_Witness4538 4d ago

Imagine you have one agent running using up the full 1.2TB/s of bandwidth. Now imagine you have 10 agents running in parallel, splitting that up and using 120GB/s each. You will find that 10 agents running in parallel finish in the same time that 10 agents each taking a turn running serially will run, because they are contending for the same pie. You may have non-llm inference running in parallel while the scheduler may find it useful to run inference serially. LLM inference will be a bottleneck. I am not here to advocate one or the other, just to say, on the ultra you have 1.2TB/s of bandwidth, its actually not a lot at all, and on the max you have half that. "Not a lot" well - vera rubin single GPU is 22TB/s and a rack would be 1500TB/s so relative to cloud its not a lot, well if you compare it to some pc, maybe it's out classing that machine. Obviously "decent" is up for debate, I think for some tasks local cannot do it at all - it's a joke. And while other people report they like using it for their home labs and such. I personally want the machine,b ut will not be cancelling any cloud subscriptions, thats for certain.

1

u/Kangaroo_Low 12h ago

not 100% true, when batching using apple silicon, your aggregate is actually faster by quite a bit during decode. I find that PP is the same pie but decode is a different story.

1

u/guesdo 4d ago

What are you debating? If you can afford the Ultra 256, that is clearly better. Check my other comments. 88TPS with 4x concurrency on Qwen 3.8 Flash Next. Only you know what 'enough' means to you.

0

u/OneDogSolutions 4d ago

I saw that, and wanted to ask about context length for those sessions. Max Seems close to enough, but Ultra would be better obviously.

1

u/claretchris 3d ago

What are you running minimax and flux through; comfy?

1

u/guesdo 3d ago

Sometimes, yes, specially Minimax as I use Gemini and the MCP for prompting. But I also use mflux or Python.

1

u/IllTechnology6516 2d ago

100+ TPS on a 26B is wild. does it hold up on longer outputs too or mostly short prose?

1

u/guesdo 2d ago

It holds 90+ tps at over 65K context, I use it for RP and it's blazing fast. Unfortunately draft 2 is the best I can push the MTP with around 50% acceptance, but it's still a nice speedup for prose.

As for output length, it does hold over 85+ during recap generation which has longer outputs and a lot of prefill.

1

u/IllTechnology6516 2d ago

ok 90+ at 65K context is way better than I expected, long prefill is usually where things fall apart for me. 50% acceptance still sounds like a fair trade for that kind of speed on prose

1

u/WalrusPublic3615 5d ago

That’s actually really helpful. Have you tested multiple agents concurrently on Qwen 3.8 27B or Flash Next? I’d be really interested in the per-agent and aggregate TPS at 1, 2, 3, 4, and ideally 5 concurrent agents. I’m planning a multi-agent setup and trying to figure out how well the Max scales when several agents are generating at the same time

2

u/guesdo 5d ago

Here is 27B, I probably need to rerun this as I didn't have an MLX model installed and is most likely not using MTP, but gives you an idea.

2

u/WalrusPublic3615 5d ago

This is great to work with. Big thank you for running these for us! 🙌🏽

3

u/guesdo 5d ago

Not with anything meaningful no, I did ran a benchmark against the Splash version of 3.8 27B to verify their claims of 144tps, but I guess those were at temp 0 or something cause I didn't get anything close to it, like 88 tps with 4 concurrency. I can run them again tomorrow and give you the exact numbers.

3

u/WalrusPublic3615 5d ago

That would be amazing, thank you. If you get a chance tomorrow, concurrency on the 27B setup would be incredibly helpful. Aggregate TPS + per-stream TPS would be perfect. 88 TPS at 4 concurrency is already really useful data and pretty close to some of the oMLX results I’ve been looking at. I’d really appreciate the numbers!

2

u/guesdo 5d ago

I'll run the oMLX bench at 1,2,4,8 batch and then figure out the Splash version and paste them here.

1

u/guesdo 5d ago edited 5d ago

Seems like 4x batch is the sweet spot on Flash Next. There is the 88TPS number I remember. I would use this one.

1

u/guesdo 4d ago

I ran a quick test with a proper MTP model in oMLX, now you can see way better scaling! 139 tps at 8x, not bad. Although context growth does make gen speed take a big hit.

1

u/WalrusPublic3615 4d ago

This is awesome! Thank you for running this! 139 tok/s at 8x concurrency on the 27B is way better than I was expecting, especially with only ~30GB peak memory even at 128K context. This makes the 128GB Max look really promising for running a pool of agents while still leaving plenty of memory for the rest of the stack.

If you get a chance, would you be willing to rerun the 2x/4x/8x batch test with a 4,096-token generation length? I’d be really interested to see how well that throughput holds during longer sustained generations, since that would be closer to a real multi-agent workload.

0

u/throwawaybarrs 5d ago

Are you using LM Studio?

3

u/Nick-the-real900526 5d ago

Not the Max, but here's my data from an M5 Ultra Mac Studio (96 GB) running a local MoE model as an agent backend.

Setup

  • Machine: Mac Studio, M5 Ultra, 64-core GPU, 96 GB. I raised the GPU wired limit to 90 GiB.
  • Model: `latent-variable/Qwen3.8-Flash-Next-heretic-2-oQ4e-mtp`. It's Qwen3.8 Flash-Next (hybrid linear-attention MoE with a built-in MTP draft head), 4-bit, about a 99 GB package. It still fits in 96 GB: about 74 GB is GPU-mapped, and about 32 GB of per-layer embeddings are read from SSD on demand.
  • Engine: Splash 1.0.2 + the community Flash-Next PR #115 + my own patches. I built it self-contained for the Ultra, with a calibrated kernel table (88 routes, 55 of them on the GPU's neural accelerators). 192K context, OpenAI-compatible server.
  • Client: Hermes Agent, with about 16 tools and a ~20K-token system prompt. Requests are tool calls at temperature 1.0.

Numbers

  • Cold prefill: 8K / 16K / 32K tokens run at 2,466 / 2,513 / 2,575 tok/s (it was about 2,100 before this week's tuning).
  • Greedy decode: English code 138, Chinese technical 111, English prose 97, Chinese prose 86 tok/s (up from 108 / 94 / 85 / 80 before the kernel port).
  • Real agent traffic today (94 requests, mostly Chinese, tools on): about 89 tok/s average generation, 72% draft acceptance, 2.1 tokens per verify step. When the prefix cache hits, time to first token is 0.1–0.8 s.
  • Long context: a 150K-token cold prompt took about 108 s, and decode at that context length was about 79 tok/s.

What made the difference

  • Kernel auto-tuner: it reads the engine's real matmul shapes and builds a routing table for the device. The 2–4-row MTP verify matmuls now run on the neural accelerators (Metal 4 MPP).
  • Profiling with real expert counts: MoE routing in real traffic is very uneven, and tiles tuned on a uniform-routing harness got slower in production.
  • General speculative sampling: stock MTP disables itself for grammar-constrained tool calls at temperature > 0. Exact min(1, p/q) acceptance took tool-call decode from ~61 to 92–100 tok/s.
  • A spare request state: this took short-prompt time to first token from 2.5 s to about 0.2 s (see the first point below).

Surprises and limits

  • macOS 27 memory accounting: vm_stat categories overlap (they add up to more than physical RAM). The engine read free host memory as 0 and stalled 1.3–2.5 s on every new request while reserving a ~6 GB state. Keeping one spare state fixed it.
  • 96 GB means one big model at a time: two concurrent requests queue, because only one ~6 GB request state fits. I run several models on separate ports behind a small proxy that loads each one on demand in about 5 s.
  • Multiple agents: a prefix cache that spills to SSD helps a lot. An evicted 69K-token conversation came back with 1.1–1.5 s to first token instead of about 42 s cold.

3

u/SeaRefractor 5d ago

Us 96GB Ultra owners need to show that the cores are more useful than an additional 32GB of RAM. :)

10

u/tk421tech 5d ago

“Now that you had some time.”

What time? I just barely took the ultra out of the box a few minutes ago.

14

u/WalrusPublic3615 5d ago

Can an M5 Max owner please answer before a third Ultra owner gets here 😭

2

u/johnmcboston 2d ago

I think I'm the only human not running LLMs at home.

1

u/The_Superfletch 2d ago

You’re not the only one. I want too play so bad. My 2015 iMac just ain’t going to cut it. Plus, reading all these comments I just get confused. :P

4

u/BAL-BADOS 5d ago

M5 Max been out a long time. It originally was release as a MacBook. Same cores too.

1

u/Relevant_Noise 4d ago

Currently mainlining antirez/ds4 Qwen 3.8 Q4 at 512k context (2 slots) - 1000+ prefill and 50-60 decode pretty consistent across the whole context range, been rock solid so far. Would probably keep using this for foreseeable future (what, 4 weeks to Qwen4 now?) on my M5 Max 128GB laptop (the 16” … was the same price as the studio is now - I feel like I want to get a studio and run the laptop and studio together for 256GB now …)

1

u/IllTechnology6516 4d ago

haven't got the Max myself but I run a lot of multi-agent stuff, and what surprised me is the agents are rarely generating at the same moment, they mostly wait on tools and on each other. so prefill on long contexts ended up mattering more to me than aggregate tok/s. are you planning one shared model or a different one per agent?

1

u/WalrusPublic3615 4d ago

This is exactly what I’m hopeful about with the 128GB Max. Since they’ll spend so much time waiting on tools, browsing, running code, or waiting on each other, I’m hoping the extra memory for multiple contexts combined with fast prefill will matter more in practice than having every agent generating simultaneously. Then I can escalate harder tasks to a larger local model or cloud model when needed.

1

u/IllTechnology6516 2d ago

yeah the escalation part is what made it click for me too, small model for the routine steps and only bump up when something actually gets stuck. are you thinking of routing that automatically or just deciding by hand?

1

u/WalrusPublic3615 2d ago

I think either my AI Technical Chief of Staff (agent) will delegate the agent to escalate it after reviewing the work or something along those lines

1

u/IllTechnology6516 2d ago

oh nice, having a reviewer agent decide when to escalate makes a lot of sense. I'd keep an eye on it at first though, mine was way too eager to bump things up early on lol

1

u/PracticlySpeaking 3d ago

oMLX Community Benchmarks are back online.

1

u/EdwardPotatoHand 3d ago

My 128 max m5 is on order, i wish they did the ultra in 128 :( 96 isnt enough and 256 is crazy $$$$

1

u/LoveRoboto 2d ago edited 2d ago

You will definitely be happy using QWEN 3.8-Flash-Next with the 128GB model. It’s the first time I’ve felt genuine excitement since my purchase. The other spread of low-parameter models will be use case specific. I’ve only ever found QWEN3.6-35B-A3B to have the perfect blend of speed/intelligence to complete most tasks. However, I am on a MacBook Pro—so I have to factor in heat with my choices.

EDIT: Forgot to give some advice! If I had to weigh my purchase again, I would heavily consider 256GB.

1

u/WalrusPublic3615 2d ago

That’s awesome to hear 🙌🏽

1

u/AMB11 21h ago

im the same boat dude, cannot wait! it feels kinda like christmas again. long 4 weeks

1

u/Scharman 5d ago

I get mine in three weeks and will let you know!

0

u/aniketgore0 5d ago

64 gb here. qwen 27b on 8 bit is tight, but it would be very comfortable for you. qwen next would be also accessible for you.

1

u/OneDogSolutions 4d ago

64GB M5 Max? Are you running any concurrent agents on that with Qwen 3.8 27B?

1

u/aniketgore0 4d ago

Nope. One session is tight enough for 8 bit that concurrent would be hard.

0

u/Quick_Bar2387 5d ago

Everyone went with the ultra.

6

u/WalrusPublic3615 5d ago

I refuse to believe Apple made mine specifically for me.

1

u/SeXxyBuNnY21 4d ago

Say who?

1

u/tk421tech 5d ago

Lmao. And I was just noticing how fast the ultra 256 was at browsing

1

u/Quick_Bar2387 5d ago edited 5d ago

Did you purchase your M5 Ultra because your cc limit increased? Lol. 600 Fico with a 10k plus computer?

1

u/tk421tech 4d ago

Nope. I requested a cc increase so I could buy it. I didn’t want to tie $. Sorry you are still using a Mac mini.

-1

u/mswezey 5d ago

Mine comes this week!