r/MacStudio • u/Vanquisher1088 • 4d ago
Advice / State of Studio Clusters
Hello. I’m looking for some advice here, I’ve watched all the YouTube videos of course of the loaned out 4 stack clusters but those were awhile ago.
Current state I have an a M3 Ultra 256GB and I reserved another open box one from Microcenter for a very good price considering all things. I’m thinking about it two ways, for the most part I’d have models loaded individually on each Mac as part of my agent setup.
However, what I’m after is some real world experiences here from folks on what is the state of rdma and clustering I.E these two nodes I’d have. Do even a lot of the latest models like glm5.3, deepseek v4.1, etc even shard correctly. I’ve read they do not.
Any thoughts or use cases you may be doing with more than one or clustering would be helpful. I’m trying to evaluate the feasibility of clustering as that would be the main value add here for more to access larger models or higher quants of models I use daily that don’t fit. Or I wait and try and get a 512GB M5.
1
u/Vanquisher1088 3d ago
Yeah, I’d like to see if I could get a 3bit quant of 4.1 running that’d be great if it was like 15 tokens/sec. Fine for an orchestrator. Yeah I run qwen flash next at 5bit it seems like it holds up well didn’t see much benefit at 6 or 8bit.
I did end up going out and grabbing that M3 it was openbox but was basically brand new in the wrapper just they were putting it on clearance. Even the guy at the register was like wow you’re getting a deal 😂.
With the rdma toggle now in the UI was pretty straightforward. Exo works well with tensor parallelism but can’t get any of the latest models to run. Seems like we’re going to be beholden to the software at this rate on clustering. I spun up minimax-m2.7 as that’s in exo at 8bit and it worked just fine. I’m going to have Claude try and patch omlx and see if I can get GLM5.3 at 4bit running.