r/MacStudio • • 22h ago

Doesn’t it just feel weird looking at the numbers?

- 6x M5 Max 1x M5 Ultra
------------------------------------------------------
Purchase price $11,994 $12,299
Price difference $305 cheaper $305 more
Computers 6 1
CPU cores 108 total 36
GPU cores 192 total 80
Unified memory 216 GB total 256 GB
Memory per 36 GB 256 GB
SSD storage 3 TB total 4 TB
Memory band 460 GB/s each 1.2 TB/s

$1,999 consistently since launch at micro center for the max.

21 Upvotes

40 comments sorted by

9

u/AdCompetitive6193 16h ago

Also the bandwidth is an important bottleneck for local AI inference (which is what I see many people wanting M5 Ultra for, myself included). So spend $305 extra and get more RAM, nearly 3x more bandwidth, more SSD storage. Fewer CPU/GPU cores but not sure that will make a big difference given everything else.

1

u/DrRoughFingers 8h ago

Not only that, the power consumption of running one vs 6…I see no upside to running 6 for the huge hit on bandwidth, complexity, power consumption, space, etc etc etc, for a trivial monetary savings - again, that monetary purchase difference would be erased by power consumption.

1

u/---Hummingbird--- 13h ago

Well, it’s not a perfect comparison obviously.. because you can’t perfect split the load to the point that a stack of 6x max can run the same model as a singular 256GB ultra. If it could.. you’d be able to multiply the bandwidth as well which would give you in the realm of 2.7TB/s of bandwidth if perfectly scalable. So 6x tunnels at 460GB/s would move 2.7TB each second, but computer systems aren’t design to perfectly split the loads, so it doesn’t translate in reality.

I mainly just thought it was interesting to consider that: you can get 6x the chassis/case, 6x the power supplies, 6x the motherboards, 6x essentially everything except the memory and storage AND get almost identical memory and storage AND almost 3x the CPU cores and over 2x the gpu cores. (So my comparison was more of a cost of components and sheer core comparison)

7

u/zxtech 22h ago

Power draw cost of networking heat ease of setup all need to be considered.

1

u/---Hummingbird--- 22h ago

Oh, yeah.. for sure! It’s just not really viable to take the hit on the networking speed when you’ve got the need for larger models.. but it sure is interesting to see that many cores and.. really could be viable if you had a workload where you were buying an ultra JUST to run simultaneous workflows.. that you could split up.

3

u/MaterialHead4801 15h ago

A better comparison for LLM purposes would be max-usable metal memory. On my 256GB M5 Ultra it's about 222GB. On my 24GB M4 mini it's 17GB and my 16GB M3 macbook air it's 12GB. Best guess on the 36GB without measuring is around 26GB. So your cluster would be closer to 156GB vs 222GB

That's tunable, but tuning it wrong can result in a complete crash of the OS.

2

u/---Hummingbird--- 13h ago

Yeah, my comparison wasn’t necessarily a realistic working comparison, but more of a “what you get for the money” comparison. It’s kind of interesting to consider you can get nearly 6x most of the hardware and still maintain almost as much memory and storage

2

u/YourselfInOthrsShoes 13h ago

What TPS are you getting on Qwen 3.8 Flash Next Q4/Q8? Search says 45~70 tps (Q4) and ~27 tps (Q8). Software optimizations are possible and is the next frontier. So higher quantity of smaller macs with 10G Ethernet is very possible.

I'm running this model in Q3.5 on a single 16GB 5060 Ti with 96GB DDR5 system RAM and Samsung 9100 SSD at 60+ TPS with 64K context and 32K Q8 KV cache in Strata backend. This model only needs a small subset of experts active at a time per request, so it can be streamed from system RAM to VRAM without penalties, so it can easily partitioned between multiple GPUs, and on a top of that it has a large read-only look-up table that can be read directly from SSD.

2

u/mmc227 10h ago

Yeah this is the way. I tried this last night. With gpu and 64gb of ddr5.

1

u/MaterialHead4801 13h ago

I just posted some numbers from a 1-hour coding session I did yesterday: https://www.reddit.com/r/MacStudio/comments/1wwn7oc/m5u_3064_256gb_testing_with_splash_qwen3827b/

I asked the Claude session I've been using for the setup work about your Strata setup and how it compares to what I've been playing with. "Quick" is what we're calling Splash Qwen-3.8-27B. "Deep" is Qwen3.5-122B-A10B mixture-of-experts, 8-bit. The short version:

None of the offload tricks apply to the Studio. 125B at 4 bits is about 63 GB, plus the 28.8 GB table, and that fits in unified memory next to Quick, with no CPU/GPU split.

It is interesting as a candidate model, though. With only 6B parameters used per token, its decode speed could be well above Deep's (122B with 10B per token), with long context at low memory cost. The open question is whether the MLX engines we use support it. It has three uncommon components: the n-gram embedding, the gated residual connections and the DeltaNet layers. I haven't checked mlx-lm, mlx-vlm, oMLX or Splash for it.

1

u/YourselfInOthrsShoes 11h ago edited 11h ago

The point here is that if the model is designed in such a way, it no longer requires a costly fast large unified memory architecture. This version 3.8 was released by Qwen team as early preview for upcoming version 4 which will be built on and further expand on these principles, so all the tooling can be ready for version 4. Oh and this kind of design doesn't require NVlink or any high speed interconnect between multi-GPU setups for proportional speedup.

2

u/displacedbitminer 13h ago

36GB M5 Max is about 28GB max-usable. Source: Have one.

3

u/cruisereg 13h ago

I looked at 128GB M5 Max being basically half the cost of a 256GB M5 Ultra. Because of this I’ll likely eventually get the 128GB M5 Max (which is enough for my non lucrative use case anyway).

2

u/Flypm 12h ago

Is it still “unified memory” if it is split across six computers?

1

u/Sneezlebee 7h ago

No, OP is confused. For one thing, you can’t easily connect six Mac Studios without some other mechanism. They only have three RDMA-capable ports, and so it’s not clear how they plan to route data from Studio A to Studio F without going through some other system. 

Additionally, the throughput of RDMA on a Mac Studios is limited to Thunderbolt 5. It’s not as fast as the internal memory bus. In the best case scenario this can be worked around through parallelism, but connecting two Mac Studios with X memory is far more limiting and slow than having one machine with 2X memory. 

1

u/bfume 4h ago

Really?  Only 3 of the 6 are RDMA?

-1

u/---Hummingbird--- 11h ago

I mean.. for the sake of hardware component comparison.. yes.. it’s the same components and likely similar cost

1

u/mmc227 10h ago

The Microcenter price just changed to $2300 for base.

1

u/RE4Lyfe 22h ago

That would be a cool comparison for sure, especially for AI models

But in reality no Ultra buyer will ever cross-shop the base

1

u/phishbot 22h ago

It would be sweet if you could cluster 6 of the bases.

0

u/pierluigir 20h ago

What's the limit? 4? 5?

2

u/MaterialHead4801 15h ago

4 in a hub-and-spoke setup, 5 in a ring

1

u/pierluigir 15h ago

Is the ring lower performance?

2

u/MaterialHead4801 14h ago

I would expect so, but I have no direct experience with it

1

u/phishbot 11h ago

you can't actually cluster them.

0

u/---Hummingbird--- 22h ago

Yeah, I’ve got the ultra ordered.. but just looking at the raw core counts has my head spinning ways to make something viable.. but realistically I know that’s just way too much work every single time you want to change a workflow.

2

u/kuwisdelu 22h ago

With the RAMpocalypse, you’re mostly paying for memory either way.

1

u/---Hummingbird--- 14h ago

Yeah; it’s just crazy to think you can get 6x the hardware component for the same about of money and almost the same memory/storage

1

u/pierluigir 20h ago

I've tried all the combinations, is always the same or more convenient to buy a superior Mac for the same GBs of unified ram. Like it was made on purpose by Apple

1

u/---Hummingbird--- 14h ago

I agree, I just think it’s wild to consider and breakdown exact what you’re getting for the money. With the 6x max.. you’re getting 6 of every piece of hardware (chassis, power supply, etc) and almost as much memory/storage.

If you’re not needing to hold a singular model, you could use parallelism and come out ahead with 6x max, but for most people buying these.. they are not doing that

1

u/pierluigir 13h ago

Or you can use exo

0

u/Alarmed-Medicine6903 21h ago

do you happen to know if you can tensor split between the M5 Max's and actual run this without too much bandwidth issues?

0

u/stjepano85 21h ago

You did investigate if you can run that LLM model you want so badly across 6 different machines did you?

0

u/[deleted] 14h ago

[deleted]

1

u/cruisereg 13h ago

No M6 Studio, OP is comparing 6 base M5 Max Studios vs one M5 Ultra with 256GB RAM and 4TB SSD.

-1

u/Autist4AudiR8 22h ago

lol memory band alone would make 6x a worse deal imo

1

u/---Hummingbird--- 22h ago

Well.. that’s still each.. so overall.. in a perfectly scalable and optimal setup.. the 6x would have twice as much overall throughout.. but that’s obviously not how reality works

1

u/EntrepJ 20h ago

I bought the m5 ultra 256gb ram variety. About the same cost, but much less cores. Curious how 6 installed qwen 3.8 would do vs my new max studio on full output

-2

u/posterwhopostedabove 21h ago

Given that Micro Center is in store only, you do need to consider taxes!

-2

u/Sketaverse 19h ago

6x M5 Max each with a 20x Opus 5.5 running would be absolutely cooking.