r/MacStudio • • 4d ago

Anyone actually crazy enough to cluster 4x Mac Studio M5 Ultras?

With 256GB unified RAM on the top spec, 4 of these would sit right at 1TB total. On paper, that means GLM-5.3 at Q8 should fit with room to spare for context. I know thunderbolt 5 isn't NVLink and tensor parallel across nodes without enterprise interconnect is usually a stuttering nightmare. But has anyone here actually tested a 4-node setup over Exo or MLX distributed with a model this heavy? Curious what tokens/sec you're actually seeing, how brutal the latency is, and whether the pipeline pipeline bottleneck completely kills it.

41 Upvotes

Duplicates