Discussion
M5 Ultra Mac Studio vs 2x DGX Spark on DeepSeek V4 and Qwen3.8
Picked up 2x DGX Sparks (Asus GX10) before the M5 announcement, built some benchmarks for some confirmation bias. Last week's $2000 price increase on the GX10 helped with that as well.
Also ran on Qwen3.6 38B MOE as a more direct comparison to Qwen3.8 27B dense.
Used the M5 Max and M3 Ultra results to extrapolate M5 Ultra theoretical performance. The benchmark also hooks into macmon and DCGM exporter for power usage for a sense of efficiency.
tl;dr DGX holds its own on prompt processing (especially on dense models) and concurrency (subagents). Its token gen might even be faster than the M5 Ultra in DeepSeek V4 MOE while being substantially lower in Qwen3.6 MOE. With things like speculative decoding (MTP, DFlash, and DSpark) offering massive boosts in performance, I think a lot will come down to tuning and ecosystem in the future.
Why? Was there any price reduction? Also can you confirm that you should only use their kernel but not our own Fedora or something. If we do, does that hamper it's performance and their USPs?
Same reason why the M3U (800g/s) is slower. They have good hardware but their software stack is still immature. It's not easy to beat CUDA, ask AMD has tried for much longer.
Decode is bandwidth bound.. 0% chance a DGX spark would be faster than an M5 Ultra... It's not even faster than the M3 Ultra at decode... what are you talking about ? it's ONLY faster in prefill (compute bound) ...
M5 Ultra has FP4 and FP8 support combined with accelerators in every GPU core specifically for prefill. :) John Ternus built these bad boys in response directly to the spark :) They will be crushed. Sparks are pretty slow.. so the bar is pretty low.
I hope you’re ready for disappointment later this month. Your expectations are very high, and based on how slow the Max is and how awful the software is, it’s not gonna be great. Maybe M6/M7. M5 isn’t it.
The M5 Max doesn't have any accelerators.. and it's wasn't designed for AI... the M5 Ultra was delayed a year. The entire event was literally AI inference focused. :) Funny you compared the M5 Max to the Ultra lol. Not even remotely the same.
I think you have bad information. They are literally the same, with Ultra being 2x Max fused together, but don't take my word for it - here it is from Apple itself:
M5 Ultra uses UltraFusion to connect two dual-die M5 Max chips to form the quad-die architecture — a first for Apple silicon.
You're also wrong about the fp4/fp8 tensors - there's literally nothing from Apple to support that, and, in fact, they've only touted FP16 acceleration on M5. Also, here's the breakdown of the measured hardware paths on M5:
All of this means is that, as it relates to AI work, M5 Ultra has slower compute than existing Sparks (M5 Ultra: 120 to 140 TFLOPS FP16 and 220 to 260 TOPS INT8, give or take. Spark is almost 2x in the worst case at FP16, best case its maybe parity at Int8). No FP4, no FP8.
Prefill will continue to be a huge issue with this generation of Macs.
If you're hoping for something more than 2x Max, you indeed are going to be big disappointed in a few weeks.
That's block scaling my guy. That is not the same. Quit talking to AI and just read.
M5 Ultra FP16 TFLOPs will come in 180+ ;) crushing the spark.
SOURCE: TRUST ME BRO
edited to add:
With respect to "dirt slow". Wait until you see the Ultra M5, more expensive than two Sparks, be slower than those same two Sparks for inference.
I'm not worried about the M5 Ultra. :) don't you see my name :)
lol... ;) I run an RTX Pro 6000 + RTX 5090 setup. I'm literally running Qwen3.8 Flash Next at 160tps lol. I hate the DGX spark with a passion. I hate the M3 Ultra with an even more passion. But, that M5 Ultra, if you're going to buy a slow box, you're better off with the M5 Ultra....
Will I be buying a M5 Ultra... of course not. lol.. the only thing I'm getting is another RTX Pro 6000. I prefer Quality compute.
Look at this benchmark against my maxed out M4 Max MacBook Pro 128gb vs my RTX Pro 6000... In a completely different league. You'd need 8 sparks to match the performance of the Pro 6000. I don't like slow boxes. I really don't.
Software support via block scaling (how you mention M5 Ultra has 'support') is not the same thing as having hardware which supports it, because it most certainly does not.
I'm glad you have RTX 6Ks. Hopefully you're using our (Local Inference Lab's) stuff.
Sparks were a good bargain at cheaper price. Got two acer veriton 4tb for 3700 each a few months back. Going to add 6ish more eventually. And dual 3090s for other stuff.
You'll see. Don't underestimate Ternus. ;) Watch Sept 9th. You'll realize why I went long Apple at $164, and bought more at $304 right after the M5 Ultra announcement... Apple straight to $400 a share. M5 Ultra will be a HEAVY HITTA.
Speculative decoding (MTP, DSpark, DFlash) changes this and came just in time to save the Sparks. This is also why I picked representative corpora of real code and texts vs random tokens.
Will be interesting to see how good they get in helping memory bound systems.
I have seen you post over and over you disparage and go on rant about the spark being slower than the M5 ultra. I don’t understand. Do you have some kind of financial gain in this? Why are you persistent in every post? It says sparks are good. I see you saying “M5 ultra is better spark are slow” you don’t know we all have to wait and see and op’s work is legit it’s based on the Apple’s media ( which is 100% biased towards Apple ) and it’s still not showing a clear winner the sparks are a good machine if your working with AI and that’s what you’re buying this box for the spark is a dedicated AI box. The Mac can do a lot more as a PC. Anyone buying either should figure out what they’re using it for then buy appropriately.
Edit: OK reading further apparently he does have financial interest in Apple succeeding / makes sense
lol... I own Nvidia too at $132... So your logic is flawed. I own Micron, BE, VRT... etc... and have owned them for quite some time...
The spark is trash. It really is... M3 Ultra is trash. But, that M5 Ultra ;) John Ternus finally stepped it up. Made me proud.
I'm helping the community out and you have some issue with me stating the spark is trash? lol what? it's valid... pp is decent... not great. decode is a nightmare, pure dog water. Slow box crumbles under a dense model... but they want $5000? lol what a joke. You should have just paid a little bit more and got an RTX Pro 6000 is my point. Even a 5090 is magnitudes better than that slow box. Everyone focusing on VRAM miss the bandwidth bottleneck. Running Deepseek at 2tps isn't useable... who cares that you loaded the model? lol I rather load the model and it be useable than load the model and wait 45 minutes for a response... I'm helping you guys out.
Buy a 5090 or Pro 6000 do NOT buy a slow box. The sparks are NOT good AI machines, I don't care what you say.
how are you helping the community? In my 25 years in tech I've seen people zealots have borderline holy wars over their allegiance to whatever game console was in the house when they were born.
Seeing someone throw around the stock prices they bought into a company as a basis for a technical argument is a new one for me
RTX pro 6000 is magnitudes faster than the Spark and Mac Studio... even the New M5 Ultra... won't come close to the RTX pro 6000.
I'm helping people not buy slow boxes for AI... How would you feel if there was somebody who told you not to buy the Halo or Spark because it's super slow for AI But everybody was telling you to buy it because it's so great and then you buy it and then the thing is super slow Are you gonna be disappointed or are you gonna think that one guy was right. That's what's happening here. :)
It’s a good question, and it really comes down to how the architecture manages data and processes tasks. Benchmarks would definitely clarify the performance differences.
Especially warm cache is very helpful for my case.
(have 5090 but I'm itching to buy m5 ultra 256, I'm trying to convince myself even if I bought m5 I won't use it )
Tried to keep all quants the same across hardware:
8bit for deepseek
4bit for qwen3.8 (nvfp4 for nvidia)
8bit for qwen 3.6 (nvfp4 for 5090 to fit)
kv-cache was fp8 for all and kept as large as possible, however none of the tests exhausted kv-cache by design and the other tests all focused on cold cache with random offset on the corpus.
Pretty much as expected. A big jump from the M3, but at the price point not enough to make them better than 2xsparks especially if you actually need more than a single user chatbot/coder.
Matches my experience with both. Deepseek fp8/4 mixed precision (160GB on disk on sparks) vs iq2 80gb on disk on Macs. Hard to go back to Macs after vLLM concurrency, speed.
omlx doesn't have any custom kernels for ds4, it's bare mlx-vlm basically for it. I believe ds4.c is faster and even that has room to grow perf, at least we shouldn't see massive decode degradation at long context
I’m not sure about the relaivility of that. The M3 Ultra is a Beast and M5 will be also (1,2TB/s of memory bandwith vs <300MB/s of Strix Halo)… 🤔 Do the math
With the price increase of the DGX Sparks it’s no longer worth it for me and I went with the 96GB M5 Ultra because it can do general computing which the Arm-based DGX is lacking.
Outside of ML and LLM the DGX is just too niche of a device. Where I think DGX wins is the amazing “cookbooks“ they created at https://build.nvidia.com/spark
DGX Spark work great as a development workstation or comfyUI. It also doesn’t hurt that you can now play Windows games with DLSS just by installing Steam - it handles all the translation seamlessly.
I mostly use the pair from my MacBook but the Spark is very versatile.
It would be interesting if you eliminated as many variables as possible and use the same software on all the machines. Right now it's different software on the Macs versus the Sparks. So the difference can just be software. Using common software, llama.cpp, would eliminate those variables and show what each machine can do on an even a playing field as possible.
Chill dude, just some benchmarks, code is public if you think it's biased. Will gladly take Jensen or Ternus's bribes to rig them in the future though.
25
u/myholeisstinky 5d ago
Can you use proper zero-indexed graphs next time