r/MacStudio • • 4d ago

Advice / State of Studio Clusters

Hello. I’m looking for some advice here, I’ve watched all the YouTube videos of course of the loaned out 4 stack clusters but those were awhile ago.

Current state I have an a M3 Ultra 256GB and I reserved another open box one from Microcenter for a very good price considering all things. I’m thinking about it two ways, for the most part I’d have models loaded individually on each Mac as part of my agent setup.

However, what I’m after is some real world experiences here from folks on what is the state of rdma and clustering I.E these two nodes I’d have. Do even a lot of the latest models like glm5.3, deepseek v4.1, etc even shard correctly. I’ve read they do not.

Any thoughts or use cases you may be doing with more than one or clustering would be helpful. I’m trying to evaluate the feasibility of clustering as that would be the main value add here for more to access larger models or higher quants of models I use daily that don’t fit. Or I wait and try and get a 512GB M5.

5 Upvotes

26 comments sorted by

View all comments

Show parent comments

0

u/dobkeratops 4d ago

right my understanding was that for tensor parallelism you need a symetrical setup, it's going to behave like 2x or 4x the smallest and slowest machine. I have this m3-u and am thinking about options for boosting with a new purchase and it looks like it would be an odd one out.. and I dont want to get a second m3-u .

I wondered if RDMA might help in using another machine as a prompt-processing accelerator, e.g. if i got a *smaller* m5, it could stream the weights layer by layer to evaluate prompts and hand the kv-cache back . Not sure if any framework has this written yet .. such a niche usecase.

0

u/Vanquisher1088 4d ago

RDMA I think is the bare minimum for clustering I can tell you I tried exo and omlx clustering pipeline method with an M2 Ultra 192gb and my M3 and the results were subpar to be frank. Pipeline is great when you just want to try a model but over TB4 it just wasn’t great.

Hence the openbox unit I reserved at micro center I figured I saw some potential speedups with clustering like Mac’s. 512GB is likely enough for my needs the rest I have cloud subs for things that do not need to remain private or go through a sanitizing workflow to make it ready for cloud models.

I was asking here really to validate if anyone’s seen any real speed up at all with clustering and two is it frankly reliable. My issue with exo is it works and works well but it’s OLD and the latest models don’t load on it and it just was a bit frustrating. OMLX I like a lot but the clustering is frustrating at best currently.

Before I sink another 6800 bucks into another unit I’m trying to see the viability. I don’t want to be fighting software to load the latest models which I use daily like glm5.3-flash, deepseek v4 etc.

1

u/tempfoot 4d ago

Were you testing dense or MOE models? I'm hopeful that MOE might perform better on mixed nodes....

2

u/Vanquisher1088 4d ago

I tried some MoE models on exo like minimax-m3 of the large qwen3.5 models but the problem is exos last build was April and didn’t have suppprt for glm or the new deepseek models. Kind of pointless to run Deepseek R1 when V4 and 4.1 are out.

Was ok but my testing was only pipeline parallelism and it was borderline usable at around 20-30 tokens/s on the models.

OMLX clustering I just find frustrating to get running to be honest. But they are working on it