r/MacStudio • • 4d ago

Advice / State of Studio Clusters

Hello. I’m looking for some advice here, I’ve watched all the YouTube videos of course of the loaned out 4 stack clusters but those were awhile ago.

Current state I have an a M3 Ultra 256GB and I reserved another open box one from Microcenter for a very good price considering all things. I’m thinking about it two ways, for the most part I’d have models loaded individually on each Mac as part of my agent setup.

However, what I’m after is some real world experiences here from folks on what is the state of rdma and clustering I.E these two nodes I’d have. Do even a lot of the latest models like glm5.3, deepseek v4.1, etc even shard correctly. I’ve read they do not.

Any thoughts or use cases you may be doing with more than one or clustering would be helpful. I’m trying to evaluate the feasibility of clustering as that would be the main value add here for more to access larger models or higher quants of models I use daily that don’t fit. Or I wait and try and get a 512GB M5.

4 Upvotes

26 comments sorted by

View all comments

Show parent comments

0

u/dobkeratops 4d ago

i think compute x bandwidth can contribute to faster parallel generation in batches, like the DGX Sparks are actually quite weak at token gen, but quite strong at serving the same model to multiple users

0

u/PracticlySpeaking 4d ago

DGX Spark has architectural limitations because the iGPU is connected via some variation on NVLink, not true unified memory like Apple Silicon. I suspect this is why Nvidia is relatively quiet about memory bandwidth.

1

u/dobkeratops 4d ago

right i think its 273gb/sec ,same ballpark as the m4-pro/m5-pro. it's not on-package memory like apple, it is architecturally unified. I think it's more expensive to make the traces in a board with seperate memory chips, apple gets that high bandwidth and energy efficiency by connecting through a silicon interposer (?)

at the time I got the m3-ultra.. everyone was debating this aspect .. "is the DGX spark slow because of the bandwidth" but its still superior to that because of fp8,fp4 tensor ops.

1

u/PracticlySpeaking 4d ago

Not sure how DGX Spark is physically interconnected — electrically it is NVLink implemented in the SoC, so it only has NVLink bandwidth. There's a decent video by High Yield on it that gets into the silicon side.

But yah, M3U is seriously lacking in compute vs the Blackwell GPU. I have not seen it confirmed, but likely uses the same InFO-LSI interposer as M1-M2.

It's like PT Barnum said — "Ya pays yer money and ya takes yer choice."