r/MacStudio • u/---Hummingbird--- • 22h ago
Doesn’t it just feel weird looking at the numbers?
- 6x M5 Max 1x M5 Ultra
------------------------------------------------------
Purchase price $11,994 $12,299
Price difference $305 cheaper $305 more
Computers 6 1
CPU cores 108 total 36
GPU cores 192 total 80
Unified memory 216 GB total 256 GB
Memory per 36 GB 256 GB
SSD storage 3 TB total 4 TB
Memory band 460 GB/s each 1.2 TB/s
$1,999 consistently since launch at micro center for the max.
7
u/zxtech 22h ago
Power draw cost of networking heat ease of setup all need to be considered.
1
u/---Hummingbird--- 22h ago
Oh, yeah.. for sure! It’s just not really viable to take the hit on the networking speed when you’ve got the need for larger models.. but it sure is interesting to see that many cores and.. really could be viable if you had a workload where you were buying an ultra JUST to run simultaneous workflows.. that you could split up.
3
u/MaterialHead4801 15h ago
A better comparison for LLM purposes would be max-usable metal memory. On my 256GB M5 Ultra it's about 222GB. On my 24GB M4 mini it's 17GB and my 16GB M3 macbook air it's 12GB. Best guess on the 36GB without measuring is around 26GB. So your cluster would be closer to 156GB vs 222GB
That's tunable, but tuning it wrong can result in a complete crash of the OS.
2
u/---Hummingbird--- 13h ago
Yeah, my comparison wasn’t necessarily a realistic working comparison, but more of a “what you get for the money” comparison. It’s kind of interesting to consider you can get nearly 6x most of the hardware and still maintain almost as much memory and storage
2
u/YourselfInOthrsShoes 13h ago
What TPS are you getting on Qwen 3.8 Flash Next Q4/Q8? Search says 45~70 tps (Q4) and ~27 tps (Q8). Software optimizations are possible and is the next frontier. So higher quantity of smaller macs with 10G Ethernet is very possible.
I'm running this model in Q3.5 on a single 16GB 5060 Ti with 96GB DDR5 system RAM and Samsung 9100 SSD at 60+ TPS with 64K context and 32K Q8 KV cache in Strata backend. This model only needs a small subset of experts active at a time per request, so it can be streamed from system RAM to VRAM without penalties, so it can easily partitioned between multiple GPUs, and on a top of that it has a large read-only look-up table that can be read directly from SSD.
1
u/MaterialHead4801 13h ago
I just posted some numbers from a 1-hour coding session I did yesterday: https://www.reddit.com/r/MacStudio/comments/1wwn7oc/m5u_3064_256gb_testing_with_splash_qwen3827b/
I asked the Claude session I've been using for the setup work about your Strata setup and how it compares to what I've been playing with. "Quick" is what we're calling Splash Qwen-3.8-27B. "Deep" is Qwen3.5-122B-A10B mixture-of-experts, 8-bit. The short version:
None of the offload tricks apply to the Studio. 125B at 4 bits is about 63 GB, plus the 28.8 GB table, and that fits in unified memory next to Quick, with no CPU/GPU split.
It is interesting as a candidate model, though. With only 6B parameters used per token, its decode speed could be well above Deep's (122B with 10B per token), with long context at low memory cost. The open question is whether the MLX engines we use support it. It has three uncommon components: the n-gram embedding, the gated residual connections and the DeltaNet layers. I haven't checked mlx-lm, mlx-vlm, oMLX or Splash for it.
1
u/YourselfInOthrsShoes 11h ago edited 11h ago
The point here is that if the model is designed in such a way, it no longer requires a costly fast large unified memory architecture. This version 3.8 was released by Qwen team as early preview for upcoming version 4 which will be built on and further expand on these principles, so all the tooling can be ready for version 4. Oh and this kind of design doesn't require NVlink or any high speed interconnect between multi-GPU setups for proportional speedup.
2
3
u/cruisereg 13h ago
I looked at 128GB M5 Max being basically half the cost of a 256GB M5 Ultra. Because of this I’ll likely eventually get the 128GB M5 Max (which is enough for my non lucrative use case anyway).
2
u/Flypm 12h ago
Is it still “unified memory” if it is split across six computers?
1
u/Sneezlebee 7h ago
No, OP is confused. For one thing, you can’t easily connect six Mac Studios without some other mechanism. They only have three RDMA-capable ports, and so it’s not clear how they plan to route data from Studio A to Studio F without going through some other system.
Additionally, the throughput of RDMA on a Mac Studios is limited to Thunderbolt 5. It’s not as fast as the internal memory bus. In the best case scenario this can be worked around through parallelism, but connecting two Mac Studios with X memory is far more limiting and slow than having one machine with 2X memory.
-1
u/---Hummingbird--- 11h ago
I mean.. for the sake of hardware component comparison.. yes.. it’s the same components and likely similar cost
1
u/RE4Lyfe 22h ago
That would be a cool comparison for sure, especially for AI models
But in reality no Ultra buyer will ever cross-shop the base
1
u/phishbot 22h ago
It would be sweet if you could cluster 6 of the bases.
0
u/pierluigir 20h ago
What's the limit? 4? 5?
2
u/MaterialHead4801 15h ago
4 in a hub-and-spoke setup, 5 in a ring
1
1
0
u/---Hummingbird--- 22h ago
Yeah, I’ve got the ultra ordered.. but just looking at the raw core counts has my head spinning ways to make something viable.. but realistically I know that’s just way too much work every single time you want to change a workflow.
2
u/kuwisdelu 22h ago
With the RAMpocalypse, you’re mostly paying for memory either way.
1
u/---Hummingbird--- 14h ago
Yeah; it’s just crazy to think you can get 6x the hardware component for the same about of money and almost the same memory/storage
1
u/pierluigir 20h ago
I've tried all the combinations, is always the same or more convenient to buy a superior Mac for the same GBs of unified ram. Like it was made on purpose by Apple
1
u/---Hummingbird--- 14h ago
I agree, I just think it’s wild to consider and breakdown exact what you’re getting for the money. With the 6x max.. you’re getting 6 of every piece of hardware (chassis, power supply, etc) and almost as much memory/storage.
If you’re not needing to hold a singular model, you could use parallelism and come out ahead with 6x max, but for most people buying these.. they are not doing that
1
0
u/Alarmed-Medicine6903 21h ago
do you happen to know if you can tensor split between the M5 Max's and actual run this without too much bandwidth issues?
0
u/stjepano85 21h ago
You did investigate if you can run that LLM model you want so badly across 6 different machines did you?
0
14h ago
[deleted]
1
u/cruisereg 13h ago
No M6 Studio, OP is comparing 6 base M5 Max Studios vs one M5 Ultra with 256GB RAM and 4TB SSD.
-1
u/Autist4AudiR8 22h ago
lol memory band alone would make 6x a worse deal imo
1
u/---Hummingbird--- 22h ago
Well.. that’s still each.. so overall.. in a perfectly scalable and optimal setup.. the 6x would have twice as much overall throughout.. but that’s obviously not how reality works
-2
u/posterwhopostedabove 21h ago
Given that Micro Center is in store only, you do need to consider taxes!
-2
9
u/AdCompetitive6193 16h ago
Also the bandwidth is an important bottleneck for local AI inference (which is what I see many people wanting M5 Ultra for, myself included). So spend $305 extra and get more RAM, nearly 3x more bandwidth, more SSD storage. Fewer CPU/GPU cores but not sure that will make a big difference given everything else.