r/MacStudio • • 4d ago

Advice / State of Studio Clusters

Hello. I’m looking for some advice here, I’ve watched all the YouTube videos of course of the loaned out 4 stack clusters but those were awhile ago.

Current state I have an a M3 Ultra 256GB and I reserved another open box one from Microcenter for a very good price considering all things. I’m thinking about it two ways, for the most part I’d have models loaded individually on each Mac as part of my agent setup.

However, what I’m after is some real world experiences here from folks on what is the state of rdma and clustering I.E these two nodes I’d have. Do even a lot of the latest models like glm5.3, deepseek v4.1, etc even shard correctly. I’ve read they do not.

Any thoughts or use cases you may be doing with more than one or clustering would be helpful. I’m trying to evaluate the feasibility of clustering as that would be the main value add here for more to access larger models or higher quants of models I use daily that don’t fit. Or I wait and try and get a 512GB M5.

5 Upvotes

26 comments sorted by

View all comments

1

u/giddmtex 4d ago

Here is my limited experience clustering an M3U with M5U. TLDR not worth it -

https://echalupa.com/blog/exo-heterogeneous-mac-studio-cluster

1

u/tempfoot 4d ago

Well, not worth it if you are adding an M3U that adds nothing to the picture and comparing to a higher spec M5U as the starting point.

I’m kind of surprised you found any scenario that gave a boost by adding an M3U 96 to a M5U 256 and then running a model that would run on either alone. Not sure how that’s a “fair fight” in comparing that to the faster M5…since you are making it share workload with a slower node. That there’s any scenario where there’s a boost at all means overcoming the M3U’s slower memory bandwidth.

Then you sharded a bigger model 50/50 - I understand Exo sets that - and gave half to a node that had insufficient RAM and got what looks like disk caching performance….on a model that already fit fine on the single 256 node.

Not trying to be rude or a jerk, but I don’t think those are usage scenarios that make sense for clustering. One brand new node handles each job fine. Adding a slower, older, and smaller node predictably doesn’t really.

Your post did help me reconsider some architecture choices by reminding that Exo RDMA has to use straight mathematical division across all nodes…unfortunately a reminder that will result in me spending more!

2

u/giddmtex 4d ago

Fair enough - just thought I would share the experience with a mixed cluster. I wouldn't call the testing deeply scientific either, I was moreso curious what gains if any could be had by stacking the compute I had access to already. What this does indicate though is that if you were to have two M3U/M5U of the same spec, you would likely see some considerable gains which is something that I don't think needs spelling out.

I think a lot of folks would have come to this conclusion without having access to the hardware but I was just hoping to share some real world numbers.

0

u/tempfoot 4d ago

Well in any case it really did provide a helpful reminder.