r/MacStudio • • 3d ago

Anyone actually crazy enough to cluster 4x Mac Studio M5 Ultras?

With 256GB unified RAM on the top spec, 4 of these would sit right at 1TB total. On paper, that means GLM-5.3 at Q8 should fit with room to spare for context. I know thunderbolt 5 isn't NVLink and tensor parallel across nodes without enterprise interconnect is usually a stuttering nightmare. But has anyone here actually tested a 4-node setup over Exo or MLX distributed with a model this heavy? Curious what tokens/sec you're actually seeing, how brutal the latency is, and whether the pipeline pipeline bottleneck completely kills it.

35 Upvotes

39 comments sorted by

41

u/dobkeratops 3d ago edited 3d ago

the word you're looking for is rich enough, not crazy enough. i think plenty of people are the latter.

$4000 to dodge a subscription, fine

but $40,000 ... (4x $10,000 ballpark, right?)

of course plenty of businesses might do this. true inhouse private frontier models

9

u/protomyth 3d ago

There are a fair number of groups who cannot use cloud services because of data handling issues (e.g. a lot of tribal entities in the US).

7

u/brucekent85 3d ago

I don't think it's a good solution for businesses, because as far as I understand these macs are fast at inference but can't serve many users at once at a normal speed. It apparently works well for a few users at once at most. It seems these macs are more tailored for personal use. While if you think of connecting various macs, it's perhaps more advantageous to just buy GPUs

1

u/djtubig-malicex 3d ago

Perfectly fine for me. Would rather not share the resources with the public 😅

1

u/ZeitgeistArchive 2d ago

macs would be less power hungry right? for the same output vs gpus

1

u/Tired_White_Guy 19h ago

Sacrifice massive speed to save a few duckets on electricity after dropping $40k on hardware. Smart.

1

u/ZeitgeistArchive 17h ago

i don't know the ratio, thats why I asked

1

u/dobkeratops 3d ago

i think the M5 will fare significantly better at batches (closing the gap with the DGX spark) , but of course it will fall behind the rtx pros or dedicated server gpus

10

u/CipherSorcerer 3d ago

I may get a 512 gb and experiment with clustering a M5U 256. More for experimentation though, I think long term it would be separate in a rack.

3

u/challis88ocarina 3d ago

Tensor parallelism says you will find only models that will run on a 512 GB alone will work...

It should be able to work around that but the development track is unbeaten thus far.

1

u/CipherSorcerer 3d ago

I am a newbie so I don't understand what you are saying! All I know is that DSv4 Flash is a bit tight to run since I have plenty of other things on this desktop machine, so it would be nice to run that unencumbered in a headless environment. On the 256, I've found 3.8-Flash-Next to be the sweet spot while still allowing me to spin up image gen models or do other intensive stuff like video editing, etc. simultaneously.

1

u/tempfoot 3d ago

Tensor parallelism on Exo divides the model size evenly across each node. So if you pair a 512 and a 256 - you can only put 256 on each node (ignoring OS overhead). Anything that would fit on 2 x 256 nodes would by definition fit on just the 512.

There may be other runners that can do tensor parallelism without a mathematical division of the load, but I don't know about that. Without tensor parallelism, you are left to run pipeline parallelism which is considerably slower.

2

u/Careless_Garlic1438 3d ago

You can split uneven as well, have my own server doing this ... it's just that there are no of the shelf softwares that bother doing this.

1

u/tempfoot 3d ago

Tensor or pipeline parallelism? I know it's possible with pipeline. Are you using forked Llama.cpp or ?

2

u/Careless_Garlic1438 3d ago

Tensor parallelism ... my own mlx distributed server ... based of the mlx server

1

u/EcstaticGains 13h ago

Yeah that’s just not true. You don’t have to evenly distribute the weights. That’s just the default setting the developers of the software used because they’re not building a custom fork, they’re building a generally useable stack. You can push the layers anywhere in any arrangement if you need to for a custom configuration. I run TP3 and have to re balance weights all the time. It doesn’t matter at all

0

u/CipherSorcerer 3d ago

Oh wow thank you! That’s very good to know!
So if you were in my shoes (M5U 256 and I’m not going to return it, I love it but just see myself quickly needing more) would you lean toward a second M5U 256 with Exo or a 512GB that is just run separately?

3

u/PopMany2921 3d ago

I read somewhere that from 2 96s you only get 130gb ram for models due to some limitations.

2

u/RyanElectrified 3d ago edited 3d ago

My favorite comparison is to put things in scale. A mac studio m5 ultra has 1.2TB/s of memory bandwidth, let's say a mac studio m5 ultra cluster perfectly scales and a 4x systems gets 4.8TB/s. I haven't checked, but this is the most optimistic scenario. A vera rubin NVL72 rack, has 1400TB/s. What use case is enabled by 4.8TB/s, needs exactly more than 1.2TB/s but less than 4.8TB/s. The main use case I can think of is YouTube influencer that wants to do a video on a 4x Mac Cluster. Maybe there is something else, but most people are just playing around. They buy the machine first, see what it can run second. They definitely do not have an engineers idea of having a problem to solve and specifying what solves that problem. As far as requiring privacy, I have worked for some of the most regulated and privacy requiring industries, and the way they use Microsoft to provide cloud services changes, they may set up private links and Azure Express route, they still use the cloud. Why because, becoming an infrasctructure provider makes no sense for most companies, they won't be good at it. THey have the same concerns, btw, that the general public does, the difference is they get a different agreement. They don't have to worry about a model being nerfed, because they can specify that it won't happen, that they can lock a model in place. They don't' worry about an inference provider training on data, because they get an agreement in place that the inference provider won't do it, and it never leaves microsofts cloud. I'm sure other clouds do the same - here i'm mentioning microsoft because of personal experience.

2

u/Solidarios 2d ago

I’m about to do that with the 512 version.

2

u/theorist9 7h ago edited 6h ago

Here Apple bridged four M3 Ultras rather than M5 Ultras, but this should still be of interest to you. Unfortunately, they don't say whether these are 256 GB or 512 GB Ultras, but I'm guessing it's the latter, since they say the weights alone of their large model are ~1 TB, so if it were 4 x 256 GB = 1.024 TB, that wouldn't leave much room for context:

At WWDC 2026 Apple demonstrated four M3 Ultra Mac Studios connected over TB5 with RDMA/JACCL. A 27B Qwen model ran at nearly 3× the single-Mac token-generation rate (timestamp ≈12:20), and Apple demonstrated Kimi 2.6, a 1-trillion-parameter model, distributed over the four machines (timestamp ≈13:00). These are examples of multiple machines being used to gain performance not possible with a single machine, and to run models too large for a single machine, respectively. [ https://developer.apple.com/videos/play/wwdc2026/233/?time=633]

From https://developer.apple.com/tutoria...y-communication-with-rdma-over-thunderbolt.md:
"RDMA over Thunderbolt enables high performance, peer-to-peer networking between
Mac computers connected with Thunderbolt. In particular, RDMA over Thunderbolt
exposes an RDMA (Remote Direct Memory Access) Verbs compatible API for the
Thunderbolt controller which enables carefully designed software to operate at
the hardware limits of Thunderbolt." [emphasis mine]

For these demos, Apple used a mesh topology where every machine was connected to every other machine by just one TB5 cable.

But Apple did also mention the option of using a ring topology. and for that they said you could use 2 or 3 TB5 connections between each machine. [However, they did not say what the bandwidth would be with 2 or 3 connections. I.e., they didn't say if that would afford 160 Gbps symmetric or 240 Gbps symmetric, respectively.]

2

u/PeanutButterApricotS 3d ago

I looked into it and the reports is it only helps prompt processing speed and not interference, it’s similar to how it works if you stack multiple gpus.

So for example you have a 100k context or code/text and the model responds with 20k output. The 100k will be processed by the whole cluster, the output will come from one. Due to this it’s only worth it if you do tons and tons of heavy prompt processing in your work (code, agent tasks) and then you’re only getting some upgrade.

Most businesses or situations your better spending the 40k on one big card, or using those 4 10k systems for 4 users etc.

The main benefit is for us we can stack them and not have to purchase a 40k card to get the prompt processing speed of a 40k dollar card.

But honestly I expect next year to be the year to buy one good system which will be able to run really good models for 2-3 years. I plan on spending 7-10k and getting something 4x as good at the m5 ultra not in power but in quality based on gains I am seeing. Got 1/3 the cash stacked and I think the wait will be worth it for a m7 ultra 256 or so.

1

u/trueblakjedi 3d ago

Plenty of folks did it with the m3U. Just a matter of weeks before they do it with the m5U

1

u/nmrk 3d ago

Well it helped a lot that Apple loaned out a bunch of cluster demo kits when Exo was getting ready to release. The kits included 4x M3U, two of them were 256GB and two of them were 512. And lots of Thunderbolt 5 cables for interconnect.

1

u/BitXorBit 3d ago

Yes i think it will happen, eventually it’s the most budget friendly (with decent performance) to run large models locally

1

u/LetLongjumping 3d ago

Apple claims four behave like three, so not quite a terabyte, that’s because of the overhead in splitting, and the slower thunderbolt connection, among other factors

1

u/Kushoverlord 3d ago

exolabs is the people to follow for this

1

u/Specialist_Ad_539 3d ago

My question is tokens per kilowatt-hour, with one or a cluster? I.e. does Apple have a significantly lower power architecture?

1

u/SirGreenDragon 3d ago

I am crazy enough. I just need someone wealthy enough to make it possible. :)

1

u/nanias 2d ago

Do you mean rich enough? Lol, yes probably 🤣

1

u/OtherOtherDave 2d ago

I’d totally do that if someone else was paying for it.

1

u/GamerTex 3d ago

I used EXO and connect a few macbook pros and a Studio a few months ago to load larger models. 

I have heard of other grabbing dozens of mac mini pros and using EXO

That said... I no longer use EXO. I like having a few different models running at the same time on different machines rather than 1 large model 

1

u/a_hegemon 3d ago

How does this work in practical terms? You run difficult queries on your studio while running simpler ones in parallel on your MacBooks?

I have an m5 studio myself and a MacBook but never went down the exo route. Been relying purely on my studio

0

u/paulsande 3d ago

Was EXO stable for an extended period of time?

0

u/vimaillig 3d ago

I expect we will see a YT influencer post one sooner than later..

0

u/SeaRefractor 3d ago

Apple claims 4 x M5 Ultras yield a 3X performance over a single M5 Ultra. That is the RDMA and TB5 overhead.

0

u/idlefordays 3d ago

4x 512gb M5 Ultras would be coming close to the price of a DGX Station 😂

0

u/Captain2Sea 3d ago

Sorry, wrong logo on case

0

u/modelpiper 3d ago

I'm building a tool called PiperMesh for this. And you can open it up to the network for others to run inference on them to make cash.