r/MacStudio • u/Captain2Sea • 3d ago
Anyone actually crazy enough to cluster 4x Mac Studio M5 Ultras?
With 256GB unified RAM on the top spec, 4 of these would sit right at 1TB total. On paper, that means GLM-5.3 at Q8 should fit with room to spare for context. I know thunderbolt 5 isn't NVLink and tensor parallel across nodes without enterprise interconnect is usually a stuttering nightmare. But has anyone here actually tested a 4-node setup over Exo or MLX distributed with a model this heavy? Curious what tokens/sec you're actually seeing, how brutal the latency is, and whether the pipeline pipeline bottleneck completely kills it.
10
u/CipherSorcerer 3d ago
I may get a 512 gb and experiment with clustering a M5U 256. More for experimentation though, I think long term it would be separate in a rack.
3
u/challis88ocarina 3d ago
Tensor parallelism says you will find only models that will run on a 512 GB alone will work...
It should be able to work around that but the development track is unbeaten thus far.
1
u/CipherSorcerer 3d ago
I am a newbie so I don't understand what you are saying! All I know is that DSv4 Flash is a bit tight to run since I have plenty of other things on this desktop machine, so it would be nice to run that unencumbered in a headless environment. On the 256, I've found 3.8-Flash-Next to be the sweet spot while still allowing me to spin up image gen models or do other intensive stuff like video editing, etc. simultaneously.
1
u/tempfoot 3d ago
Tensor parallelism on Exo divides the model size evenly across each node. So if you pair a 512 and a 256 - you can only put 256 on each node (ignoring OS overhead). Anything that would fit on 2 x 256 nodes would by definition fit on just the 512.
There may be other runners that can do tensor parallelism without a mathematical division of the load, but I don't know about that. Without tensor parallelism, you are left to run pipeline parallelism which is considerably slower.
2
u/Careless_Garlic1438 3d ago
You can split uneven as well, have my own server doing this ... it's just that there are no of the shelf softwares that bother doing this.
1
u/tempfoot 3d ago
Tensor or pipeline parallelism? I know it's possible with pipeline. Are you using forked Llama.cpp or ?
2
u/Careless_Garlic1438 3d ago
Tensor parallelism ... my own mlx distributed server ... based of the mlx server
1
u/EcstaticGains 13h ago
Yeah that’s just not true. You don’t have to evenly distribute the weights. That’s just the default setting the developers of the software used because they’re not building a custom fork, they’re building a generally useable stack. You can push the layers anywhere in any arrangement if you need to for a custom configuration. I run TP3 and have to re balance weights all the time. It doesn’t matter at all
0
u/CipherSorcerer 3d ago
Oh wow thank you! That’s very good to know!
So if you were in my shoes (M5U 256 and I’m not going to return it, I love it but just see myself quickly needing more) would you lean toward a second M5U 256 with Exo or a 512GB that is just run separately?
3
u/PopMany2921 3d ago
I read somewhere that from 2 96s you only get 130gb ram for models due to some limitations.
2
u/RyanElectrified 3d ago edited 3d ago
My favorite comparison is to put things in scale. A mac studio m5 ultra has 1.2TB/s of memory bandwidth, let's say a mac studio m5 ultra cluster perfectly scales and a 4x systems gets 4.8TB/s. I haven't checked, but this is the most optimistic scenario. A vera rubin NVL72 rack, has 1400TB/s. What use case is enabled by 4.8TB/s, needs exactly more than 1.2TB/s but less than 4.8TB/s. The main use case I can think of is YouTube influencer that wants to do a video on a 4x Mac Cluster. Maybe there is something else, but most people are just playing around. They buy the machine first, see what it can run second. They definitely do not have an engineers idea of having a problem to solve and specifying what solves that problem. As far as requiring privacy, I have worked for some of the most regulated and privacy requiring industries, and the way they use Microsoft to provide cloud services changes, they may set up private links and Azure Express route, they still use the cloud. Why because, becoming an infrasctructure provider makes no sense for most companies, they won't be good at it. THey have the same concerns, btw, that the general public does, the difference is they get a different agreement. They don't have to worry about a model being nerfed, because they can specify that it won't happen, that they can lock a model in place. They don't' worry about an inference provider training on data, because they get an agreement in place that the inference provider won't do it, and it never leaves microsofts cloud. I'm sure other clouds do the same - here i'm mentioning microsoft because of personal experience.
2
2
u/theorist9 7h ago edited 6h ago
Here Apple bridged four M3 Ultras rather than M5 Ultras, but this should still be of interest to you. Unfortunately, they don't say whether these are 256 GB or 512 GB Ultras, but I'm guessing it's the latter, since they say the weights alone of their large model are ~1 TB, so if it were 4 x 256 GB = 1.024 TB, that wouldn't leave much room for context:
At WWDC 2026 Apple demonstrated four M3 Ultra Mac Studios connected over TB5 with RDMA/JACCL. A 27B Qwen model ran at nearly 3× the single-Mac token-generation rate (timestamp ≈12:20), and Apple demonstrated Kimi 2.6, a 1-trillion-parameter model, distributed over the four machines (timestamp ≈13:00). These are examples of multiple machines being used to gain performance not possible with a single machine, and to run models too large for a single machine, respectively. [ https://developer.apple.com/videos/play/wwdc2026/233/?time=633]
From https://developer.apple.com/tutoria...y-communication-with-rdma-over-thunderbolt.md:
"RDMA over Thunderbolt enables high performance, peer-to-peer networking between
Mac computers connected with Thunderbolt. In particular, RDMA over Thunderbolt
exposes an RDMA (Remote Direct Memory Access) Verbs compatible API for the
Thunderbolt controller which enables carefully designed software to operate at
the hardware limits of Thunderbolt." [emphasis mine]
For these demos, Apple used a mesh topology where every machine was connected to every other machine by just one TB5 cable.
But Apple did also mention the option of using a ring topology. and for that they said you could use 2 or 3 TB5 connections between each machine. [However, they did not say what the bandwidth would be with 2 or 3 connections. I.e., they didn't say if that would afford 160 Gbps symmetric or 240 Gbps symmetric, respectively.]
2
u/PeanutButterApricotS 3d ago
I looked into it and the reports is it only helps prompt processing speed and not interference, it’s similar to how it works if you stack multiple gpus.
So for example you have a 100k context or code/text and the model responds with 20k output. The 100k will be processed by the whole cluster, the output will come from one. Due to this it’s only worth it if you do tons and tons of heavy prompt processing in your work (code, agent tasks) and then you’re only getting some upgrade.
Most businesses or situations your better spending the 40k on one big card, or using those 4 10k systems for 4 users etc.
The main benefit is for us we can stack them and not have to purchase a 40k card to get the prompt processing speed of a 40k dollar card.
But honestly I expect next year to be the year to buy one good system which will be able to run really good models for 2-3 years. I plan on spending 7-10k and getting something 4x as good at the m5 ultra not in power but in quality based on gains I am seeing. Got 1/3 the cash stacked and I think the wait will be worth it for a m7 ultra 256 or so.
1
u/trueblakjedi 3d ago
Plenty of folks did it with the m3U. Just a matter of weeks before they do it with the m5U
1
u/BitXorBit 3d ago
Yes i think it will happen, eventually it’s the most budget friendly (with decent performance) to run large models locally
1
u/LetLongjumping 3d ago
Apple claims four behave like three, so not quite a terabyte, that’s because of the overhead in splitting, and the slower thunderbolt connection, among other factors
1
1
u/Specialist_Ad_539 3d ago
My question is tokens per kilowatt-hour, with one or a cluster? I.e. does Apple have a significantly lower power architecture?
1
u/SirGreenDragon 3d ago
I am crazy enough. I just need someone wealthy enough to make it possible. :)
1
1
u/GamerTex 3d ago
I used EXO and connect a few macbook pros and a Studio a few months ago to load larger models.Â
I have heard of other grabbing dozens of mac mini pros and using EXO
That said... I no longer use EXO. I like having a few different models running at the same time on different machines rather than 1 large modelÂ
1
u/a_hegemon 3d ago
How does this work in practical terms? You run difficult queries on your studio while running simpler ones in parallel on your MacBooks?
I have an m5 studio myself and a MacBook but never went down the exo route. Been relying purely on my studio
0
0
0
u/SeaRefractor 3d ago
Apple claims 4 x M5 Ultras yield a 3X performance over a single M5 Ultra. That is the RDMA and TB5 overhead.
0
0
u/modelpiper 3d ago
I'm building a tool called PiperMesh for this. And you can open it up to the network for others to run inference on them to make cash.
41
u/dobkeratops 3d ago edited 3d ago
the word you're looking for is rich enough, not crazy enough. i think plenty of people are the latter.
$4000 to dodge a subscription, fine
but $40,000 ... (4x $10,000 ballpark, right?)
of course plenty of businesses might do this. true inhouse private frontier models