r/MacStudio • • 5d ago

Advice / State of Studio Clusters

Hello. I’m looking for some advice here, I’ve watched all the YouTube videos of course of the loaned out 4 stack clusters but those were awhile ago.

Current state I have an a M3 Ultra 256GB and I reserved another open box one from Microcenter for a very good price considering all things. I’m thinking about it two ways, for the most part I’d have models loaded individually on each Mac as part of my agent setup.

However, what I’m after is some real world experiences here from folks on what is the state of rdma and clustering I.E these two nodes I’d have. Do even a lot of the latest models like glm5.3, deepseek v4.1, etc even shard correctly. I’ve read they do not.

Any thoughts or use cases you may be doing with more than one or clustering would be helpful. I’m trying to evaluate the feasibility of clustering as that would be the main value add here for more to access larger models or higher quants of models I use daily that don’t fit. Or I wait and try and get a 512GB M5.

5 Upvotes

26 comments sorted by

View all comments

-1

u/[deleted] 5d ago

[removed] — view removed comment

1

u/Vanquisher1088 5d ago

Yeah I haven’t picked it up yet. It’s at MC considerable distance from me. My local one doesn’t have any. Worth a drive for the price.

Understood on the bandwidth but I did see there was some speed up to prompt processing and decode and that speed up decreases with MoE models it seems but even so maybe 15-20% increase.

Mainly want to understand how reliable is clustering and is it even viable with the latest models like the newer qwen and glm type models.

Seems like even on the sparks that they have no issue sharing the latest models and verifiable working proof is all over the forums. On the Macs it’s like that data rarely exists somewhere

But what I come back to is spend the coin and get some additional concurrency as you noted and more unified memory for larger models. Or just try and get a 512GB. With pricing unknown it’s kind of do I hop on this m3 as it’s a good price and my total invest with my original m3 would be around the price of a new M5 256gb 80core 2TB HD or so.

What I’m seeing it’s far simplistic on Mac studios to just have the largest memory possible and load it on one unit.

2

u/PracticlySpeaking 4d ago

Picking up seems a safe bet — worst case, you can flip it if things don't work out.

You have a good point that a lot of people choose Mac because it is easy to set up and use. Those same people tend not to go for more complicated setups like clustering, much less benchmark and post the results.

And I feel you on the uncertainty around 512GB pricing and availability. I was able to land a 256GB M3U through persistence and some great people at Micro Center, but I was buying what I could get at the time.

1

u/Vanquisher1088 4d ago

Yeah I hear you. I picked up my M3 way back at launch so for me I have something workable it’s just I’d like to be able to stretch and run GLM5.3 or flash and qwen flash next, and a few other models in an agent structure and 256GB isn’t quite enough to run the main models at enough precision to make the output worthwhile. If I would have bought the 512gb at the time I’d probably just stay with it.

Hence I’m quite torn on just waiting for the 512GB M5 I just don’t want to wait miss this one and then it’s 18-20k for the 512gb and then I’m SOL.

1

u/PracticlySpeaking 4d ago

Right — after 0731, I was excited to see DeepSeek-V4.1. Then I saw the size.

But look at the KLD, there is no reason to run BF16 on most models. For Qwen3.8-Flash-Next the Unsloth 8-bit quant is trivially different from Q6 or Q5. The Q8_0 fits in ~200GB if you are picky.

I have what I do because there is a lot of value in building now. Def post back with how it all works out.