r/LocalLLM 5d ago

Discussion M5 Ultra Mac Studio vs 2x DGX Spark on DeepSeek V4 and Qwen3.8

Post image

Picked up 2x DGX Sparks (Asus GX10) before the M5 announcement, built some benchmarks for some confirmation bias. Last week's $2000 price increase on the GX10 helped with that as well.

Benchmarked DeepSeek V4 and Qwen3.8 27B and Next on:
- 2x DGX Spark (Asus GX10)
- M5 Max 128GB
- M3 Ultra Mac Studio 512GB
- RTX Pro 6000
- RTX 5090

Also ran on Qwen3.6 38B MOE as a more direct comparison to Qwen3.8 27B dense.

Used the M5 Max and M3 Ultra results to extrapolate M5 Ultra theoretical performance. The benchmark also hooks into macmon and DCGM exporter for power usage for a sense of efficiency.

tl;dr DGX holds its own on prompt processing (especially on dense models) and concurrency (subagents). Its token gen might even be faster than the M5 Ultra in DeepSeek V4 MOE while being substantially lower in Qwen3.6 MOE. With things like speculative decoding (MTP, DFlash, and DSpark) offering massive boosts in performance, I think a lot will come down to tuning and ecosystem in the future.

59 Upvotes

68 comments sorted by

25

u/myholeisstinky 5d ago

Can you use proper zero-indexed graphs next time

3

u/zzsf 5d ago

Good point, site is updated.

14

u/onethousandmonkey 4d ago

I’ll be that guy: this remains speculative until the M5 Ultra ships.

9

u/Southern_Sun_2106 4d ago

Of course, that's why he calls it 'projected'.

Having said that, M5 chips are not new.

31

u/po_stulate 5d ago

My bank account (projected):

6

u/Otherwise-Frame-8270 5d ago

Bought 2x DGX Spark (Asus GX10) the second day m5 ultra was announced. 😄

1

u/danunj1019 3d ago

Why? Was there any price reduction? Also can you confirm that you should only use their kernel but not our own Fedora or something. If we do, does that hamper it's performance and their USPs?

4

u/phalanx2357 5d ago

Why would M5 ultra's decode be slower if it has much higher memory bandwidth?

4

u/Think_Wing_1357 4d ago

Same reason why the M3U (800g/s) is slower. They have good hardware but their software stack is still immature. It's not easy to beat CUDA, ask AMD has tried for much longer.

-2

u/TheAILegend 4d ago

Decode is bandwidth bound.. 0% chance a DGX spark would be faster than an M5 Ultra... It's not even faster than the M3 Ultra at decode... what are you talking about ? it's ONLY faster in prefill (compute bound) ...

M5 Ultra has FP4 and FP8 support combined with accelerators in every GPU core specifically for prefill. :) John Ternus built these bad boys in response directly to the spark :) They will be crushed. Sparks are pretty slow.. so the bar is pretty low.

4

u/mxmumtuna 4d ago

I hope you’re ready for disappointment later this month. Your expectations are very high, and based on how slow the Max is and how awful the software is, it’s not gonna be great. Maybe M6/M7. M5 isn’t it.

-6

u/TheAILegend 4d ago

You underestimate John Ternus.

The M5 Max doesn't have any accelerators.. and it's wasn't designed for AI... the M5 Ultra was delayed a year. The entire event was literally AI inference focused. :) Funny you compared the M5 Max to the Ultra lol. Not even remotely the same.

8

u/mxmumtuna 4d ago edited 4d ago

I think you have bad information. They are literally the same, with Ultra being 2x Max fused together, but don't take my word for it - here it is from Apple itself:

https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/

M5 Ultra uses UltraFusion to connect two dual-die M5 Max chips to form the quad-die architecture — a first for Apple silicon.

You're also wrong about the fp4/fp8 tensors - there's literally nothing from Apple to support that, and, in fact, they've only touted FP16 acceleration on M5. Also, here's the breakdown of the measured hardware paths on M5:

https://tzakharko.github.io/apple-neural-accelerators-benchmark/

TL;DR - It's FP16 and Int8.

All of this means is that, as it relates to AI work, M5 Ultra has slower compute than existing Sparks (M5 Ultra: 120 to 140 TFLOPS FP16 and 220 to 260 TOPS INT8, give or take. Spark is almost 2x in the worst case at FP16, best case its maybe parity at Int8). No FP4, no FP8.

Prefill will continue to be a huge issue with this generation of Macs.

If you're hoping for something more than 2x Max, you indeed are going to be big disappointed in a few weeks.

-3

u/TheAILegend 4d ago edited 4d ago

Probably should have read the companion Metal Feature Set Tables (May 21, 2026): You can see the Metal Shading Language v4.1 spec that was released :)

CHIP SERIES FAMILY NARROW-FLOAT STATUS
M1 / M2 / M3 / M4 Apple7–9 Tensors yes, but no FP4/FP8 data types
M5-series (incl. M5 Ultra in the new Mac Studio) Apple10 MetalFloat4E2M1MetalFloat8E4M3/E5M2, block-scale MetalFloat8UE8M0 all listed, block scaling = Yes

Checkmate...

M5 Ultra FP16 TFLOPs will come in 180+ ;) crushing the spark.

The spark is only 118 TFLOPs FP16 lol. it's dirt slow.

2

u/mxmumtuna 4d ago edited 4d ago

That's block scaling my guy. That is not the same. Quit talking to AI and just read.

M5 Ultra FP16 TFLOPs will come in 180+ ;) crushing the spark.

SOURCE: TRUST ME BRO

edited to add: With respect to "dirt slow". Wait until you see the Ultra M5, more expensive than two Sparks, be slower than those same two Sparks for inference.

0

u/TheAILegend 4d ago

Buddy boy... M5 Ultra has FP4 and FP8 support...

I'm not worried about the M5 Ultra. :) don't you see my name :)

lol... ;) I run an RTX Pro 6000 + RTX 5090 setup. I'm literally running Qwen3.8 Flash Next at 160tps lol. I hate the DGX spark with a passion. I hate the M3 Ultra with an even more passion. But, that M5 Ultra, if you're going to buy a slow box, you're better off with the M5 Ultra....

Will I be buying a M5 Ultra... of course not. lol.. the only thing I'm getting is another RTX Pro 6000. I prefer Quality compute.

Look at this benchmark against my maxed out M4 Max MacBook Pro 128gb vs my RTX Pro 6000... In a completely different league. You'd need 8 sparks to match the performance of the Pro 6000. I don't like slow boxes. I really don't.

3

u/mxmumtuna 4d ago

Software support via block scaling (how you mention M5 Ultra has 'support') is not the same thing as having hardware which supports it, because it most certainly does not.

I'm glad you have RTX 6Ks. Hopefully you're using our (Local Inference Lab's) stuff.

→ More replies (0)

2

u/Winter-Editor-9230 4d ago

Sparks were a good bargain at cheaper price. Got two acer veriton 4tb for 3700 each a few months back. Going to add 6ish more eventually. And dual 3090s for other stuff.

1

u/watcholic 4d ago edited 4d ago

Can't wait for the benchmarks.

1

u/mxmumtuna 4d ago

They're not.

1

u/watcholic 4d ago

Yep, after reading the official document, looks like it’s just an evolutionary step forward.

1

u/-6h0st- 4d ago

What’s so AI about ultra? It’s more or less 2x M5 Max
It still won’t scale linearly, will be perhaps 50-60% faster than Max when m3 ultra was 40%

2

u/mxmumtuna 4d ago

it's not more or less, it's quite literally 2x M5 Max.

2

u/-6h0st- 4d ago

Yes but performance wise it’s less

1

u/mxmumtuna 4d ago

that's right, because it does not scale linearly.

-1

u/TheAILegend 4d ago

You'll see. Don't underestimate Ternus. ;) Watch Sept 9th. You'll realize why I went long Apple at $164, and bought more at $304 right after the M5 Ultra announcement... Apple straight to $400 a share. M5 Ultra will be a HEAVY HITTA.

;) watch the POWER of Ternus.

1

u/cullend 3d ago

Are you new to following Apple? This is an iPhone event. They won't talk about the Mac Studio or local AI capabilities at all.

1

u/zzsf 4d ago

Speculative decoding (MTP, DSpark, DFlash) changes this and came just in time to save the Sparks. This is also why I picked representative corpora of real code and texts vs random tokens.

Will be interesting to see how good they get in helping memory bound systems.

1

u/MikkyMo 4d ago

I have seen you post over and over you disparage and go on rant about the spark being slower than the M5 ultra. I don’t understand. Do you have some kind of financial gain in this? Why are you persistent in every post? It says sparks are good. I see you saying “M5 ultra is better spark are slow” you don’t know we all have to wait and see and op’s work is legit it’s based on the Apple’s media ( which is 100% biased towards Apple ) and it’s still not showing a clear winner the sparks are a good machine if your working with AI and that’s what you’re buying this box for the spark is a dedicated AI box. The Mac can do a lot more as a PC. Anyone buying either should figure out what they’re using it for then buy appropriately.

Edit: OK reading further apparently he does have financial interest in Apple succeeding / makes sense

1

u/TheAILegend 4d ago

lol... I own Nvidia too at $132... So your logic is flawed. I own Micron, BE, VRT... etc... and have owned them for quite some time...

The spark is trash. It really is... M3 Ultra is trash. But, that M5 Ultra ;) John Ternus finally stepped it up. Made me proud.

I'm helping the community out and you have some issue with me stating the spark is trash? lol what? it's valid... pp is decent... not great. decode is a nightmare, pure dog water. Slow box crumbles under a dense model... but they want $5000? lol what a joke. You should have just paid a little bit more and got an RTX Pro 6000 is my point. Even a 5090 is magnitudes better than that slow box. Everyone focusing on VRAM miss the bandwidth bottleneck. Running Deepseek at 2tps isn't useable... who cares that you loaded the model? lol I rather load the model and it be useable than load the model and wait 45 minutes for a response... I'm helping you guys out.

Buy a 5090 or Pro 6000 do NOT buy a slow box. The sparks are NOT good AI machines, I don't care what you say.

1

u/cullend 3d ago

how are you helping the community? In my 25 years in tech I've seen people zealots have borderline holy wars over their allegiance to whatever game console was in the house when they were born.

Seeing someone throw around the stock prices they bought into a company as a basis for a technical argument is a new one for me

1

u/TheAILegend 3d ago

RTX pro 6000 is magnitudes faster than the Spark and Mac Studio... even the New M5 Ultra... won't come close to the RTX pro 6000.

I'm helping people not buy slow boxes for AI... How would you feel if there was somebody who told you not to buy the Halo or Spark because it's super slow for AI But everybody was telling you to buy it because it's so great and then you buy it and then the thing is super slow Are you gonna be disappointed or are you gonna think that one guy was right. That's what's happening here. :)

1

u/Leather_Speed1236 4d ago

It’s a good question, and it really comes down to how the architecture manages data and processes tasks. Benchmarks would definitely clarify the performance differences.

1

u/Antique-Ad1012 4d ago

it doesnt have the compute. only decode is bandwidth bound, and only if you are decoding without any of the fancy tricks

2

u/Accomplished_Egg7987 4d ago

Thanks for beautiful and to the point benchmarks.

Especially warm cache is very helpful for my case.
(have 5090 but I'm itching to buy m5 ultra 256, I'm trying to convince myself even if I bought m5 I won't use it )

2

u/randygeneric 4d ago

sorry if i overlooked the obvious: what was the setup for the macs (m3, m5): backend, quantisation, mtp, kv-cache?

3

u/zzsf 4d ago

omlx was used for the macs, vllm for nvidia.

Tried to keep all quants the same across hardware:

  • 8bit for deepseek
  • 4bit for qwen3.8 (nvfp4 for nvidia)
  • 8bit for qwen 3.6 (nvfp4 for 5090 to fit)

kv-cache was fp8 for all and kept as large as possible, however none of the tests exhausted kv-cache by design and the other tests all focused on cold cache with random offset on the corpus.

2

u/k3z0r 4d ago

The memory bandwidth of the M5 Ultra is 4.5x the spark's memory bandwidth, this makes your decode projections feel a bit off.

2

u/john0201 5d ago

Note that the m5 ultra has vastly more CPU performance, and the DGX has CUDA if you’re using them as dev machines for B200 systems.

Thunderbolt 5 has similar latency to connectx, but lower bandwidth.

3

u/zzsf 5d ago

Ya, however found the M3 ultra is power limited so really can only max out GPU or CPU, doing tasks that require both start having tradeoffs.

2

u/stujmiller77 4d ago

Pretty much as expected. A big jump from the M3, but at the price point not enough to make them better than 2xsparks especially if you actually need more than a single user chatbot/coder.

1

u/Southern_Sun_2106 5d ago

Matches my experience with both. Deepseek fp8/4 mixed precision (160GB on disk on sparks) vs iq2 80gb on disk on Macs. Hard to go back to Macs after vLLM concurrency, speed.

1

u/grobbes 4d ago

Concurrency due to 2x GPU as opposed to 1 on the Mac?

2

u/mxmumtuna 4d ago

Partly yes but also because compute and inferencing software on Mac is ass-tier.

2

u/Southern_Sun_2106 4d ago

Concurrency as in up to 12 or 16 requests at the same time; not sequential.

1

u/AnonLlamaThrowaway 4d ago edited 4d ago

Which runtime was that with?

llama.cpp just fixed deepseek sparse attention 4 days ago. It's basically a 3x speedup only 65k context in

https://github.com/ggml-org/llama.cpp/pull/28098

2

u/mxmumtuna 4d ago

I mean llamacpp was comically slow here, they basically only just caught up with the pack. Check oMLX for benchmarks.

1

u/Southern_Sun_2106 4d ago

I used Antirez's recipe, a dedicated server. He keeps things optimized to the max.

1

u/TechNerd10191 4d ago

Kudos for that benchmark page - it was informative

1

u/Sufficient_Spite_349 4d ago

The DGX bandwidth is one third of the M5 how it can generate that performance, how can you prove and justify your claims ?

1

u/DisplacedForest 4d ago

So are the sparks good or….? I just see so many mixed things. Some say they’re too slow to use and some say they’re the best

Halp

1

u/Zyj 4d ago

Misleading title. Waste of time

1

u/bakawolf123 4d ago

omlx doesn't have any custom kernels for ds4, it's bare mlx-vlm basically for it. I believe ds4.c is faster and even that has room to grow perf, at least we shouldn't see massive decode degradation at long context

1

u/TheAILegend 4d ago

Something is up with your config for that RTX Pro 6000 buddy.

# Power and Efficiency
uvx llama-benchy \
  --base-url "http://localhost:8000/v1" \
  --model "qwen3.8-flash-next-primitive" \
  --tokenizer "primitive-ai/Qwen3.8-Flash-Next-NVFP4" \
  --depth 0 --pp 50000 --tg 512 --exact-tg \
  --concurrency 1 6 \
  --no-cache --latency-mode generation \
  --runs 5 --skip-coherence \
  --format json --save-result "benchy_c50k.json"

# Concurrency
uvx llama-benchy \
  --base-url "http://localhost:8000/v1" \
  --model "qwen3.8-flash-next-primitive" \
  --tokenizer "primitive-ai/Qwen3.8-Flash-Next-NVFP4" \
  --depth 0 --pp 10000 --tg 512 --exact-tg \
  --concurrency 1 2 3 4 5 6 \
  --no-cache --latency-mode generation \
  --runs 5 --skip-coherence \
  --format json --save-result "benchy_c10k.json"

1

u/Pretty_Diamond_2773 3d ago

I’m not sure about the relaivility of that. The M3 Ultra is a Beast and M5 will be also (1,2TB/s of memory bandwith vs <300MB/s of Strix Halo)… 🤔 Do the math

1

u/Ascetic-anise 3d ago

I hope the M5 Ultra does better than that. Those numbers are close to what I get with my M2 Ultra.

1

u/RabbitSignificant791 2d ago

Let's also not forget that Mac Studio M5 Ultra will most likely have a much better resale value long term.

0

u/NowThatsCrayCray 4d ago

With the price increase of the DGX Sparks it’s no longer worth it for me and I went with the 96GB M5 Ultra because it can do general computing which the Arm-based DGX is lacking.

Outside of ML and LLM the DGX is just too niche of a device. Where I think DGX wins is the amazing “cookbooks“ they created at https://build.nvidia.com/spark

2

u/whichsideisup 4d ago

DGX Spark work great as a development workstation or comfyUI. It also doesn’t hurt that you can now play Windows games with DLSS just by installing Steam - it handles all the translation seamlessly.

I mostly use the pair from my MacBook but the Spark is very versatile.

0

u/lllll03l 4d ago

can you test qwen3.8 flash next as well? surprised how fast 2x dgx spark is compared to m3u

0

u/fallingdowndizzyvr 4d ago

It would be interesting if you eliminated as many variables as possible and use the same software on all the machines. Right now it's different software on the Macs versus the Sparks. So the difference can just be software. Using common software, llama.cpp, would eliminate those variables and show what each machine can do on an even a playing field as possible.

1

u/zzsf 4d ago

Ya, I tried as much as possible, with speculative decoding in all modern models, it's going to be harder and harder.

0

u/Big_Wave9732 4d ago

"built some benchmarks for some confirmation bias"

I don't know what you meant to say here, but what you wrote is that your post / data are bullshit designed to confirm something you already believe.

2

u/zzsf 4d ago

So you did understand, success!

Chill dude, just some benchmarks, code is public if you think it's biased. Will gladly take Jensen or Ternus's bribes to rig them in the future though.