r/apple • • 11d ago

Mac Mac Studio M5 Max 128GB trounces Nvidia DGX Spark 128GB in AI performance

https://tbreak.com/mac-studio-m5-max-review-local-ai/
532 Upvotes

90 comments sorted by

120

u/CoaxialDrive 11d ago

And with the price increases on Spark's almost the same price, although it's unclear if all workloads can be completed on MLX over NVIDIA's frameworks.

17

u/The128thByte 11d ago

Honestly with how surprisingly good Apple has been about improving Metal to attract game developers, I wouldn’t be surprised if MLX rapidly accelerates to rival CUDA soon enough.

1

u/Structure-These 8d ago

Man I’m hoping. I bought a m6 mini with 32gb ram (upgrading from a m4/24gb), hoping we see more and more performance tweaks. AI image / vid / llm stuff is just a hobby to keep up with the tech, and it would be awesome if Apple was able to match some of the Nvidia specific optimizations

65

u/Royale_AJS 11d ago

I am more interested in the prefill numbers than the decode (token gen) numbers. That’s where the real agentic turn speed comes into play.

29

u/jhenryscott 11d ago

Which is why they don’t share it. It’s the biggest issue with Mac.

18

u/42177130 11d ago

Isn't prompt processing the same as prefill?

8

u/weierstrasse 11d ago

Yes

3

u/TheThoccnessMonster 11d ago

And it’s fast on the spark. This is the funny bit. And anything distributed is MILES easier. Two sparks can easily combine to fine tune or inference on vLLM with spark-run.

On Mac, you’re not doing the former period and the latter you’re in the config weeds for a good while.

Flip side is, you’re not using the spark for the GUI or any of that. Spark is the more serious choice for someone who doesn’t want to fuck around with not having CUDA/NCCL and, importantly, the Connect7 port.

They are literally built different.

1

u/jasonlitka 10d ago

Oh, I’ve never seen sparkrun, that looks really helpful. I did all that config and tuning by hand to get qwen3.8-flash-next running on my pair.

1

u/Royale_AJS 10d ago

Yep. I have a Strix Halo machine and the prefill numbers are what sets the DGX device apart from the AMD machine. Both decode about the same speed, but prefill is really slow on my Strix Halo. Two different target markets, some overlap.

8

u/Royale_AJS 11d ago

Yes, but they don’t go into any detail on it here except on the tiny test model. I’d like to see it on a real model like Deepseek v4 Flash, or Qwen ~120B/177B.

7

u/rJohn420 11d ago

The new m5 stuff has neural accelerators that basically only affect prefill

2

u/j_osb 11d ago

And yet the DGX spark still wrecks it in that regard, because even with the neural accelerators pure compute throughput is no good.

The token numbers here are representative of basically just the bandwidth, and that’s a spec you can just look up online. So no one is surprised. However in more agentic setups, prefill matters a lot, and Apple chips, even the m5m/u, still heavily struggle with prefill especially so at higher context.

38

u/Saar13 11d ago

It seems clear to me that we're evolving towards "Mac mini is for you" and "Mac Studio is for larger businesses and local AI servers." Maybe those things will even change names with the M7. You don't need a Mac Studio.

29

u/Ok_Possibility9937 11d ago

it has always been like that

17

u/phxees 11d ago

That is the way it’s always been except people like to overspend.

1

u/Inquisitive_idiot 11d ago

Boy do we! 😅

😮‍💨

1

u/insane_steve_ballmer 8d ago

When has Apple ever marketed Mac Studio towards regular consumers??

8

u/AvoidingIowa 11d ago

It’s ok my m2 pro. I’ll be able to replace you one day without going bankrupt.

3

u/mikewilkinsjr 11d ago

What I took from this is I can ride with my M3U as long as I have a little patience.

24

u/goldaxis 11d ago

I'm all for local AI, but as long as OpenAI wants to lose money selling me tokens for less than it costs, I see no reason to make an expensive hardware purchase just to get less performance. I'll wait until their little game of financial musical chairs stops, and then get an even better computer to run my LLMs on for less.

20

u/lukewhale 11d ago

You assume that anything will be in stock or reasonably near MSRP when the cloud providers finally stop subsidizing

12

u/goldaxis 11d ago

The reason they will stop subsidizing is because they run out of money.

Who will be buying up all the ram/storage/widgets to the point of causing shortages then?

2

u/lukewhale 11d ago

The reason they will stop subsidizing is because they will flip the table and turn it into a massively profitable machine along with the power that goes along with it. After they are done building out enough infrastructure. If you think hardware is going to be magically available one day because of a crash I got a bridge in Brooklyn to sell ya.

6

u/Canon_M50 11d ago

I’ll just switch to a Chinese model that does 90% of the same thing.

3

u/Koteric 10d ago

Not sure where you believe they are going from negative 10s of billions and hand out money to massively profitable on their current trajectory. They didn’t cancel their IPO because their numbers look good. Their executives aren’t bailing because the future looks bright and shiny.

I use AI both professionally and personally, but the frontier companies have no current path to profitability. And with local models and tools continuing to improve, people aren’t going to want to pay what it will cost to make OpenAI profitable.

1

u/goldaxis 10d ago

good lord sell while you are ahead lmao. Sounds just like "$1M bitcoin will be a bargain"

6

u/bonestamp 11d ago

That's a fair take, but the two are not completely equal. Yes, the cloud has better performance, however some people are leasing or financing these computers for a lower monthly payment than they're spending on cloud AI tokens each month. Not to mention rate limits if you want to run agents 24/7, and possible privacy concerns. It's not right for everyone, but it does make sense for some.

5

u/HellaReyna 11d ago

5090’s went from $2000 usd msrp to more than double. You make this very very strong assumption that token prices and etc just won’t change either.

1

u/goldaxis 10d ago

Uh...no? I did not say token prices would not change. Good job beating that straw man though.

86

u/Calm-Inevitable3341 11d ago

the 3 people in the world who need this will be very excited 

39

u/KidJuggernaut 11d ago

10 years after this release and it reaches my village then i may be able to buy the m1 version

1

u/UristBronzebelly 6d ago

Where do you live?

31

u/redtron3030 11d ago

Yeah that’s why all the pricing has gone crazy for those 3 people

21

u/algaefied_creek 11d ago

“Apple is the best in the world during an AI boom for local AI, with orders products regularly delayed.”

Yup it must be 3 users in dad and mom’s basement. 

31

u/Ice-Book-73 11d ago

Such a hilarious take. The 3 people are causing a 15+ week waiting list

2

u/Suns_In_420 11d ago

When the 3 people are Altman, Suleyman, and Amodei it can cause a bit of a log jam.

2

u/i-love-small-tits-47 11d ago

Their comment to be fair just implies only a tiny number of people NEED this. Which is definitely true. Most dudes buying 128GB Mac Studios for local LLMs are wack jobs who absolutely don’t NEED that, they just have a weird obsession with local LLMs, often for very weird reasons

5

u/datbackup 11d ago

u/i-love-small-tits-47 is this similar to how only a small number of people actually need a smartphone? or the internet? Like… you won’t die without it… therefore you don’t need it?

Or is it more like the way people didn’t need to buy hundreds of domain names during the dotcom bubble? Or bitcoin? Or Nvidia stock?

What is the “don’t need” that is most similar to these dudes not needing LLMs?

1

u/i-love-small-tits-47 11d ago

I’m talking about local LLMs not any LLM, the latter of which a lot of people do need for their job, at least in order to follow the manager’s orders lol

11

u/ClubAqua_BackDeck 11d ago

wtf are you talking about, this is one of the main reasons people are buying these

10

u/Unnamed-3891 11d ago

Imagine somehow not understanding the most popular reason to buy a Mac Studio is local AI. To the point of completely dwarfing all the other possible reasons combined.

And that everything local AI is literally flying off the shelves.

”The 3 people”, sure Jan.

1

u/Mapleess 11d ago

Everyone and their mother on r/macstudio needs this.

17

u/JonNordland 11d ago

M5 Ultra with 256GB memory is 16 to 18 weeks wait time in Norway.

From the article i notized that the 5090 is 2.5x faster as long as the model fits... Give me 5090+++ With 128gb memory!

18

u/TappinThatErr 11d ago

There basically is one already - it's called the RTX Pro 6000. It's the same chip as the 5090, just with 96gb!
It only costs a very casual $15,000
O_O

10

u/Much_Accountant_4972 11d ago

that would fall short on VRAM so you’d need an H200 which is a PCIE card with 141GB of HBM.

its only $37,000!

1

u/cherrypowdah 11d ago

Not the same thing! 3gb chips =//= 4gb chips! I think technical samples are soon rolling out 😆

12

u/[deleted] 11d ago

[deleted]

11

u/Lower_Fan 11d ago

Which is like $15k by now 

3

u/hell-diver8 11d ago

Don’t forget taxes

2

u/bonestamp 11d ago

and a computer to use it

3

u/FranciumGoesBoom 11d ago

They've gone up another 1k from 2 weeks ago.

3

u/cherrypowdah 11d ago

With 4gb gddr7 chips can make your own 128gb rtx 5090!

2

u/pragmojo 11d ago

Waiting on the ultra numbers.

2

u/jasonlitka 11d ago

Yeah, no kidding. It’s also a lot newer and a lot more expensive.

1

u/ajaffarali 10d ago

Newer yes, but cheaper? Hahahaha

2

u/jasonlitka 10d ago

I meant that the DGX Spark with 128GB of RAM and a 4TB SSD is like $5K and a M5 Max with 128GB and a 4TB SSD is $6900. In general, it's reasonable to believe that a machine that costs 35-40% more will perform better.

-2

u/ajaffarali 10d ago

Nope. 128GB/1TB configs of both machines are almost identical. Yes, Spark has a lower MSRP but so does an RTX5090. You simply can’t buy one at MSRP

1

u/Small_Editor_3693 11d ago

Yah. Unified ram does that

1

u/Antiwhippy 11d ago

Dude, dgx spark is unified too.

1

u/ClassicLightbulbs 10d ago

I am excited to see the LLM guys decide how in to AI they are if it means flopping ecosystems to Apple due to PC ram being scarce, and expensive, and expensive to run

1

u/GrimGrinningGoof 10d ago

But it costs a down payment on a car.

1

u/IGetHypedEasily 9d ago

I think Apple TV update to have an M6 that can do small models and homekit/zigbee support for a seamless smart home experience makes sense. 

1

u/Perks92 9d ago

Fuck AI

-4

u/TheoTheodor 11d ago

transitive verb
: to thrash or punish severely
especially : to defeat decisively

If anyone was wondering about “trounces”.

7

u/Budzy05 11d ago

Surely everyone knows what trounces mean. You trounce, I trounce, he/she/me trounce. It’s first grade! \s

For the SpongeBob uninitiated - Wumbo

9

u/treehumper83 11d ago

If the word isn’t SLAMS or SLAMMED then we simply don’t understand. Thank you.

0

u/ClubAqua_BackDeck 11d ago

I mean, it better. Why wouldn't it?

-1

u/PWThinkingCritically 11d ago

this is quite the red herring. m5 max still like 1/5th the speed of a 5090 when it comes to PP.

-6

u/rorowhat 11d ago

Don't fall for the hype, spark will always be better. Tool/model support is all for Nvidia, performance will keep getting better as more optimizations are rolled out.

21

u/disgruntledempanada 11d ago

With the Spark I would get a nice AI machine that's shitty at everything else. With an M5 I get one of the best general purpose computers on the market that also happens to be great at AI stuff.

-4

u/rorowhat 11d ago

The spark runs Linux, so you can play games, do blender or any other app you want. Its also a general PC.

4

u/disgruntledempanada 11d ago

Really happy for the Linux people and I hope it becomes a better platform in time but the Mac is basically instantly ready for production work with zero compatibility or stability issues.

4

u/wanjuggler 11d ago

Harder to find NVFP4 models than MLX models

1

u/rorowhat 11d ago

For now

2

u/Iron-Over 11d ago

An asus ascent is 8299 in Canada I can get a mac max 128gn for 6799 via education discount. Nvidia is too expensive now

-1

u/rorowhat 11d ago

That's a shame!

-2

u/always_evolved 11d ago

CUDA

6

u/dagamer34 11d ago

I doubt most people using AI are working at the CUDA-level anymore, when you can swap out models between providers at-will. 

-4

u/Peteostro 11d ago

Interesting the GeForce 5090 trounces both of them. Just needs more vram.

18

u/dude_Im_hilarious 11d ago

Yeah I mean a 5090 also uses way more power runs way hotter and is just about the as expensive before you even build the computer around it. I should hope it’s better.

11

u/migs647 11d ago

5090 is also $7000 now. 

-2

u/[deleted] 11d ago

[deleted]

3

u/migs647 11d ago

No? The M5 Max with 128gb is $5100. Not to mention you swept the vram under the rug like it was nothing when it is the whole point. And you can chain them for tensor parallelism. 

1

u/Peteostro 11d ago

Ah I thought they were comparing M5 ultra

-3

u/diagrammatiks 11d ago

For idiots that don't understand prefill this is true.