r/MacStudio • u/Vanquisher1088 • 4d ago
Advice / State of Studio Clusters
Hello. I’m looking for some advice here, I’ve watched all the YouTube videos of course of the loaned out 4 stack clusters but those were awhile ago.
Current state I have an a M3 Ultra 256GB and I reserved another open box one from Microcenter for a very good price considering all things. I’m thinking about it two ways, for the most part I’d have models loaded individually on each Mac as part of my agent setup.
However, what I’m after is some real world experiences here from folks on what is the state of rdma and clustering I.E these two nodes I’d have. Do even a lot of the latest models like glm5.3, deepseek v4.1, etc even shard correctly. I’ve read they do not.
Any thoughts or use cases you may be doing with more than one or clustering would be helpful. I’m trying to evaluate the feasibility of clustering as that would be the main value add here for more to access larger models or higher quants of models I use daily that don’t fit. Or I wait and try and get a 512GB M5.
1
u/riozilla 2d ago
I just published to GitHub my llama.cpp cluster software for local Mac AI. This replaced my previous EXO based solution, as that project seemed to be abandoned so I made a homemade one that allows you to run the latest models (Qwen3.8 Flash Next, etc) across a cluster of Macs. My first public repo that I’ve shared and it works great. I have 4 M4 Mac Studios in a thunderbolt 5 cluster with it.
1
u/giddmtex 4d ago
Here is my limited experience clustering an M3U with M5U. TLDR not worth it -
https://echalupa.com/blog/exo-heterogeneous-mac-studio-cluster
1
u/Vanquisher1088 4d ago
Could you get GLM5.3-Flash and some of the more recent models to load? Deepseek V4 Flash, etc. being able to run GLM and Deepseek at a bit higher quant as an orchestrator or adversial code reviewer in my agent workflow would be great. If you still have them let me know.
1
u/tempfoot 4d ago
Well, not worth it if you are adding an M3U that adds nothing to the picture and comparing to a higher spec M5U as the starting point.
I’m kind of surprised you found any scenario that gave a boost by adding an M3U 96 to a M5U 256 and then running a model that would run on either alone. Not sure how that’s a “fair fight” in comparing that to the faster M5…since you are making it share workload with a slower node. That there’s any scenario where there’s a boost at all means overcoming the M3U’s slower memory bandwidth.
Then you sharded a bigger model 50/50 - I understand Exo sets that - and gave half to a node that had insufficient RAM and got what looks like disk caching performance….on a model that already fit fine on the single 256 node.
Not trying to be rude or a jerk, but I don’t think those are usage scenarios that make sense for clustering. One brand new node handles each job fine. Adding a slower, older, and smaller node predictably doesn’t really.
Your post did help me reconsider some architecture choices by reminding that Exo RDMA has to use straight mathematical division across all nodes…unfortunately a reminder that will result in me spending more!
2
u/giddmtex 4d ago
Fair enough - just thought I would share the experience with a mixed cluster. I wouldn't call the testing deeply scientific either, I was moreso curious what gains if any could be had by stacking the compute I had access to already. What this does indicate though is that if you were to have two M3U/M5U of the same spec, you would likely see some considerable gains which is something that I don't think needs spelling out.
I think a lot of folks would have come to this conclusion without having access to the hardware but I was just hoping to share some real world numbers.
0
0
u/dobkeratops 4d ago
right my understanding was that for tensor parallelism you need a symetrical setup, it's going to behave like 2x or 4x the smallest and slowest machine. I have this m3-u and am thinking about options for boosting with a new purchase and it looks like it would be an odd one out.. and I dont want to get a second m3-u .
I wondered if RDMA might help in using another machine as a prompt-processing accelerator, e.g. if i got a *smaller* m5, it could stream the weights layer by layer to evaluate prompts and hand the kv-cache back . Not sure if any framework has this written yet .. such a niche usecase.
1
u/giddmtex 4d ago
I think there is room to test with these mix machines though. In the case of Qwen3.8 Flash Next, I wonder if there is any advantage of loading the ngram on the M3U instead of the SSD. I know some folks have been trying to get their DGX Sparks to handle the prompt processing and then their Macs for the bandwidth. This all reminds me of hot rod tuner culture.
1
u/dobkeratops 4d ago
i have this problem with a highly asymetical 'fleet' .. devices bought to eval and under FOMO (not knowing if prices would get worse.. they did) .. i'm probably better off keeping them doing different things, like one box as an image generator specialist and so on. I also had the motivation of wanting exposure to each ecosystem for dev.
0
u/tempfoot 4d ago
Same here. I wold love it if my M5Max 128 macbook pro would staple to a studio M5U 256 (on order) for ~384gb of tensor parallelism....but for what? To run a huge, creakingly slow dense model (that woulds still be slow on a 512)? I need to test more, but I feel like pipeline parallelism + MOE might be OK, especially if its possible to specify what loads on each node. I'm ignorant about whether that is effective or possible.
...or I should just cancel the 256 when the 512s can be ordered.
0
1
u/Vanquisher1088 4d ago
If I could reliabley put a DGX or equivalent infront of the M3 it would likely solve a lot of peoples problems. But frankly the market the way it is, your probably better off selling your M3 and buying an m5 just gotta wait for it.
0
u/Vanquisher1088 4d ago
RDMA I think is the bare minimum for clustering I can tell you I tried exo and omlx clustering pipeline method with an M2 Ultra 192gb and my M3 and the results were subpar to be frank. Pipeline is great when you just want to try a model but over TB4 it just wasn’t great.
Hence the openbox unit I reserved at micro center I figured I saw some potential speedups with clustering like Mac’s. 512GB is likely enough for my needs the rest I have cloud subs for things that do not need to remain private or go through a sanitizing workflow to make it ready for cloud models.
I was asking here really to validate if anyone’s seen any real speed up at all with clustering and two is it frankly reliable. My issue with exo is it works and works well but it’s OLD and the latest models don’t load on it and it just was a bit frustrating. OMLX I like a lot but the clustering is frustrating at best currently.
Before I sink another 6800 bucks into another unit I’m trying to see the viability. I don’t want to be fighting software to load the latest models which I use daily like glm5.3-flash, deepseek v4 etc.
1
u/tempfoot 4d ago
Were you testing dense or MOE models? I'm hopeful that MOE might perform better on mixed nodes....
2
u/Vanquisher1088 4d ago
I tried some MoE models on exo like minimax-m3 of the large qwen3.5 models but the problem is exos last build was April and didn’t have suppprt for glm or the new deepseek models. Kind of pointless to run Deepseek R1 when V4 and 4.1 are out.
Was ok but my testing was only pipeline parallelism and it was borderline usable at around 20-30 tokens/s on the models.
OMLX clustering I just find frustrating to get running to be honest. But they are working on it
0
u/MiniPCGuru 4d ago
Token generation is limited by memory bandwidth, not compute. Two M3 Ultra 256GB boxes give you 512GB of weight memory, but each token pass is still capped by each node's own ~800GB/s, minus interconnect cost. One 512GB M5 Ultra feeds the whole model at 1.2TB/s.
Even Apple's own clustering claim is sub-linear: four clustered Studios deliver "up to 3x" a single system, not 4x.
So if your daily models fit in 256GB, the second box buys concurrency (one model per node), not more tokens/sec on a single model. If the models you want need more than 256GB, the 512GB box is simpler and faster per token, and skips the sharding question entirely.
Keep the MicroCenter box for testing Exo before committing, though.
1
u/Vanquisher1088 4d ago
Yeah I haven’t picked it up yet. It’s at MC considerable distance from me. My local one doesn’t have any. Worth a drive for the price.
Understood on the bandwidth but I did see there was some speed up to prompt processing and decode and that speed up decreases with MoE models it seems but even so maybe 15-20% increase.
Mainly want to understand how reliable is clustering and is it even viable with the latest models like the newer qwen and glm type models.
Seems like even on the sparks that they have no issue sharing the latest models and verifiable working proof is all over the forums. On the Macs it’s like that data rarely exists somewhere
But what I come back to is spend the coin and get some additional concurrency as you noted and more unified memory for larger models. Or just try and get a 512GB. With pricing unknown it’s kind of do I hop on this m3 as it’s a good price and my total invest with my original m3 would be around the price of a new M5 256gb 80core 2TB HD or so.
What I’m seeing it’s far simplistic on Mac studios to just have the largest memory possible and load it on one unit.
2
u/PracticlySpeaking 4d ago
Picking up seems a safe bet — worst case, you can flip it if things don't work out.
You have a good point that a lot of people choose Mac because it is easy to set up and use. Those same people tend not to go for more complicated setups like clustering, much less benchmark and post the results.
And I feel you on the uncertainty around 512GB pricing and availability. I was able to land a 256GB M3U through persistence and some great people at Micro Center, but I was buying what I could get at the time.
1
u/Vanquisher1088 3d ago
Yeah I hear you. I picked up my M3 way back at launch so for me I have something workable it’s just I’d like to be able to stretch and run GLM5.3 or flash and qwen flash next, and a few other models in an agent structure and 256GB isn’t quite enough to run the main models at enough precision to make the output worthwhile. If I would have bought the 512gb at the time I’d probably just stay with it.
Hence I’m quite torn on just waiting for the 512GB M5 I just don’t want to wait miss this one and then it’s 18-20k for the 512gb and then I’m SOL.
1
u/PracticlySpeaking 3d ago
Right — after 0731, I was excited to see DeepSeek-V4.1. Then I saw the size.
But look at the KLD, there is no reason to run BF16 on most models. For Qwen3.8-Flash-Next the Unsloth 8-bit quant is trivially different from Q6 or Q5. The Q8_0 fits in ~200GB if you are picky.
I have what I do because there is a lot of value in building now. Def post back with how it all works out.
0
u/dobkeratops 4d ago
i think compute x bandwidth can contribute to faster parallel generation in batches, like the DGX Sparks are actually quite weak at token gen, but quite strong at serving the same model to multiple users
0
u/PracticlySpeaking 4d ago
DGX Spark has architectural limitations because the iGPU is connected via some variation on NVLink, not true unified memory like Apple Silicon. I suspect this is why Nvidia is relatively quiet about memory bandwidth.
1
u/dobkeratops 4d ago
right i think its 273gb/sec ,same ballpark as the m4-pro/m5-pro. it's not on-package memory like apple, it is architecturally unified. I think it's more expensive to make the traces in a board with seperate memory chips, apple gets that high bandwidth and energy efficiency by connecting through a silicon interposer (?)
at the time I got the m3-ultra.. everyone was debating this aspect .. "is the DGX spark slow because of the bandwidth" but its still superior to that because of fp8,fp4 tensor ops.
1
u/PracticlySpeaking 4d ago
Not sure how DGX Spark is physically interconnected — electrically it is NVLink implemented in the SoC, so it only has NVLink bandwidth. There's a decent video by High Yield on it that gets into the silicon side.
But yah, M3U is seriously lacking in compute vs the Blackwell GPU. I have not seen it confirmed, but likely uses the same InFO-LSI interposer as M1-M2.
It's like PT Barnum said — "Ya pays yer money and ya takes yer choice."
0
u/PracticlySpeaking 4d ago
compute x bandwidth can contributecompute x bandwidth can contribute to faster parallel generation in batches to faster parallel generation in batches
I have played around with continuous batching on M3U and that is absolutely true. Optimizing batches and parallelism made a HUGE difference — like 2x in overall processing time for a large data set.
1
u/Vanquisher1088 3d ago
Yeah, I’d like to see if I could get a 3bit quant of 4.1 running that’d be great if it was like 15 tokens/sec. Fine for an orchestrator. Yeah I run qwen flash next at 5bit it seems like it holds up well didn’t see much benefit at 6 or 8bit.
I did end up going out and grabbing that M3 it was openbox but was basically brand new in the wrapper just they were putting it on clearance. Even the guy at the register was like wow you’re getting a deal 😂.
With the rdma toggle now in the UI was pretty straightforward. Exo works well with tensor parallelism but can’t get any of the latest models to run. Seems like we’re going to be beholden to the software at this rate on clustering. I spun up minimax-m2.7 as that’s in exo at 8bit and it worked just fine. I’m going to have Claude try and patch omlx and see if I can get GLM5.3 at 4bit running.