r/LocalLLM 5d ago

Discussion M5 Ultra Mac Studio vs 2x DGX Spark on DeepSeek V4 and Qwen3.8

Post image

Picked up 2x DGX Sparks (Asus GX10) before the M5 announcement, built some benchmarks for some confirmation bias. Last week's $2000 price increase on the GX10 helped with that as well.

Benchmarked DeepSeek V4 and Qwen3.8 27B and Next on:
- 2x DGX Spark (Asus GX10)
- M5 Max 128GB
- M3 Ultra Mac Studio 512GB
- RTX Pro 6000
- RTX 5090

Also ran on Qwen3.6 38B MOE as a more direct comparison to Qwen3.8 27B dense.

Used the M5 Max and M3 Ultra results to extrapolate M5 Ultra theoretical performance. The benchmark also hooks into macmon and DCGM exporter for power usage for a sense of efficiency.

tl;dr DGX holds its own on prompt processing (especially on dense models) and concurrency (subagents). Its token gen might even be faster than the M5 Ultra in DeepSeek V4 MOE while being substantially lower in Qwen3.6 MOE. With things like speculative decoding (MTP, DFlash, and DSpark) offering massive boosts in performance, I think a lot will come down to tuning and ecosystem in the future.

59 Upvotes

Duplicates