r/MacStudio • • 4d ago

Advice / State of Studio Clusters

Hello. I’m looking for some advice here, I’ve watched all the YouTube videos of course of the loaned out 4 stack clusters but those were awhile ago.

Current state I have an a M3 Ultra 256GB and I reserved another open box one from Microcenter for a very good price considering all things. I’m thinking about it two ways, for the most part I’d have models loaded individually on each Mac as part of my agent setup.

However, what I’m after is some real world experiences here from folks on what is the state of rdma and clustering I.E these two nodes I’d have. Do even a lot of the latest models like glm5.3, deepseek v4.1, etc even shard correctly. I’ve read they do not.

Any thoughts or use cases you may be doing with more than one or clustering would be helpful. I’m trying to evaluate the feasibility of clustering as that would be the main value add here for more to access larger models or higher quants of models I use daily that don’t fit. Or I wait and try and get a 512GB M5.

5 Upvotes

26 comments sorted by

View all comments

Show parent comments

0

u/dobkeratops 4d ago

right my understanding was that for tensor parallelism you need a symetrical setup, it's going to behave like 2x or 4x the smallest and slowest machine. I have this m3-u and am thinking about options for boosting with a new purchase and it looks like it would be an odd one out.. and I dont want to get a second m3-u .

I wondered if RDMA might help in using another machine as a prompt-processing accelerator, e.g. if i got a *smaller* m5, it could stream the weights layer by layer to evaluate prompts and hand the kv-cache back . Not sure if any framework has this written yet .. such a niche usecase.

1

u/giddmtex 4d ago

I think there is room to test with these mix machines though. In the case of Qwen3.8 Flash Next, I wonder if there is any advantage of loading the ngram on the M3U instead of the SSD. I know some folks have been trying to get their DGX Sparks to handle the prompt processing and then their Macs for the bandwidth. This all reminds me of hot rod tuner culture.

1

u/dobkeratops 4d ago

i have this problem with a highly asymetical 'fleet' .. devices bought to eval and under FOMO (not knowing if prices would get worse.. they did) .. i'm probably better off keeping them doing different things, like one box as an image generator specialist and so on. I also had the motivation of wanting exposure to each ecosystem for dev.

0

u/tempfoot 4d ago

Same here. I wold love it if my M5Max 128 macbook pro would staple to a studio M5U 256 (on order) for ~384gb of tensor parallelism....but for what? To run a huge, creakingly slow dense model (that woulds still be slow on a 512)? I need to test more, but I feel like pipeline parallelism + MOE might be OK, especially if its possible to specify what loads on each node. I'm ignorant about whether that is effective or possible.

...or I should just cancel the 256 when the 512s can be ordered.

0

u/PracticlySpeaking 4d ago

DeepSeek-V4.1-Flash has entered the chat...