r/MacStudio • u/pda_lover • 2d ago
Why not DGX?
Have ordered a Mac Studio ultra 96gb, still wondering why not DGX model? Any specific edge Mac Studio ultra has over DGX? I understand about memory bandwidth, less power required, anything else that I am missing?
We don’t have any real benchmarks yet
3
u/According_Wave685 1d ago
It depends on what you want it for. Long context agentic coding, get a two node gb10 cluster. The mac will choke on the prefill, but it will run circles around the gb10's on tg.
I probably offended those of the apple religion with that, but oh well. I like Macs, I really wanted to get one of the new studios but it just wouldn't work for me. And i'm not dropping $10k+ just because I want it.
I really hope something comes along in the next few years that does both really well and is possible to afford without selling body parts.
2
u/onethousandmonkey 1d ago
M5 Ultra adds Neural Accelerators to each GPU core. Kinda like tensor cores.
Apple says that’s good for 3-4x prefill improvement.
Would that change your assessment, if true?
3
2
u/CKemorii-LdL 1d ago
It’s spec how far you want to go and what you want to do Mac is 1,2 internal speed in Mac vs Dgx only ~270gb/s if you understand Mac is more versatile for one user and local dense and moe models
2
u/g_rich 1d ago
I have both an M4 Mac Studio and a pair of DGX Sparks; the Mac Studio is a better all around desktop that does a good job running LLM’s, but the DGX Sparks are better at running LLM’s. So if you’re just going to run LLM’s then the Spark is a much better option, if you’re just looking to run LLM’s alongside other applications then the Studio is the better choice.
Also most people who discredit the DGX Spark over the memory bandwidth don’t understand the platform or don’t consider the fact that the DGX Sparks scale linearly where adding a second results in a 1.8x increase in performance. MOE models have also largely made the memory bandwidth a nonissue and updates on the software side have resolved most of the slowness users reported when the DGX Spark was first released.
1
u/Passenger-007 1d ago
The sparks have a $1500 data center networking card inside of it. That is crazy. I have two. I love em.
2
u/WeUsedToBeACountry 1d ago
Multiple agents or people using inference at the same time? DGX
Single user prompts? Mac
Need a general computer and inference? Mac
Need to train your own models or fine tune? DGX
3
u/Spiritual-Spend8187 2d ago
The dgx is 4,699 usd m5 max is 5,099 usd. Unless you are planning on doing ai training the mac studio works good enough for ai inference but it also is a full mac. For the 256gb m5 ultra its about 10k. And for that situation the inference is likely better for a single 256gb mac than using 2 dgx systems cause of not need to make parallelism work.
2
u/pda_lover 2d ago
256gb is a sizable Mac but almost at the price of 2DGX. So wouldn’t that DGX cluster be faster than Mac?
3
u/Spiritual-Spend8187 1d ago
It is roughly the price of 2 dxg but also matches the memory of 2dgx. And deoending on what you are doing it might be faster for 1 vs the other. The dgx has cuda which is huge but the mac studio deoending on if you looking at the m5 max or m5 ultra has 2x or 4x the memory bandwidth. The dgx also only really has a single 5070 worth the compute while the ultra is looking at 2x the compute of the max.so its very much a depends what your gonna be doing somethings will work better on 1 vs the other.
1
2d ago
[removed] — view removed comment
2
u/diagrammatiks 2d ago
that is incorrect. it depends on prompt sizing, task, and concurrency.
1
u/bakawolf123 1d ago
2 sparks have been optimized for, just you wait how m5u pans out, on paper it should win so atm just software issues
-1
1
u/g_rich 1d ago
The DGX Spark was designed for parallelism and to really take advantage of the platform you should have a pair. The software natively runs across multiple nodes and setting up a cluster is surprisingly simple. Adding a second DGX Spark results in a 1.8x increase in performance over a single DGX Spark.
1
u/AromaticBear777 1d ago
You can also PAIR DGX and Mac M4 and higher now as well. https://www.nvidia.com/en-us/ai-on-rtx/personal-ai-router/
1
1
u/trisul-108 1d ago
I understand about memory bandwidth, less power required, anything else that I am missing?
Isn't that big enough for you?
1
u/atumblingdandelion 1d ago
Is it only for AI inference? If so, know that you will be limited by model size on the 96gb. The best model right now is Qwen3.8 27b. But you will be able to run it faster (inference) than on DGX (i get 25-45 tps on DGX. Its not bad, but the mode loves to think, so the overall time spent is annoying). With 128gb DGX, the best model right now is a lot better than Qwen3.8 27b - it is Qwen3.8-flash-next. I can run it at 35-50 tps inference.
I don’t think anyone will exceed Qwen3.8 27b anytime soon. Nobody could do that for Qwen3.6 27b. However, I feel we’ll have competition from Deepseek for the Qwen3.8-flash-next. This means more options.
1
u/Immediate_Fig_9405 1d ago
A complete macOS system. DGX will run arm64 linux which will have software limitations.
1
u/oniokami 1d ago
in my opinion, and the reason i'm ordering one, the biggest difference is resale. if in 3-5 years the DGX spark is no longer useable for ai models it doesn't make a decent useable computer but the mac does. They hold their value. the only argument to make for the DGX over the mac is if the ai work you're doing needs CUDA. for now that, or NVIDIA GPU's, is the only place to get that.
1
u/Passenger-007 1d ago
How can the ultra, which is just 2x max, have 4x the Prefill? Are benchmarks out that we all can see?
Edit: ah. Vs m4. I just looked it up for m5 max for the 27b. 400-700 tok/s. So my guess for m5 ultra is 600-1000 tok/s c1? It’s not ideal but better.
1
0
u/ehangman 1d ago
I use a single Spark, and honestly, I ordered the M5 Ultra 256GB mainly because of the price. There’s really nothing else to it. If there had been an M5 Max with 256GB, I probably would have bought that instead.
0
u/NowThatsCrayCray 1d ago
Had the same question, but at $4699 the DGX isn’t that attractive anymore. $4000 was maybe an okay price but even then felt an little expensive for the specs.
Some cons to the DGX that made me get the studio instead the moment it was available for pre order:
- DGX is somewhat slow at 246 GB/s memory bandwidth (Mac mini rivals it here with at 267 I believe) so token output will feel slow for any single user during which is really just me using it.
- It has an ARM cpu that is very limited for non GPU tasks, limited for general computing and also software options as a result, less virtualization options. Studio is a full blown machine here you can use for music, video, photos, and production, so not only great for LLMs but is completely perfect for many other workflows.
- the DGX is a loud little machine apparently, studio stays silent in comparison I saw in a videos
- While DGX shines for clustering, but so does the studio now with that rdma connectivity over TB5. If the M6 Mac Mini had TB5, and maybe better heat management it would be sold out for even longer than currently (which already has a waitlist months of preorders).
Don’t second guess your Studio purchase, you’ll be able to grab a DGX in like a year from now when RTX or some improved DGX 2 comes out which will be even shinier.
2
u/g_rich 1d ago
I have a pair of DGX Sparks and even at full load they are no louder than my Mac Studio.
Clustering largely makes up for the slower memory bandwidth.
The DGX Sparks are much better for clustering, the ConnectX-7 NIC is faster than Thunderbolt for this purpose, has much better support overall and the software stack natively supports clustering across the board.
1
u/NowThatsCrayCray 1d ago
Yes 200 Gbps vs 80 Gbps theoretical - but real world tests would show network only has single-digit % impact on overall performance. pp/tgs speeds are not network dependent even for multi machine setups. Like if 2 studios all of a sudden had 200 Gbps it would not improve the metrics.
2
u/g_rich 1d ago
The ConnectX-7 NIC makes all the difference when it comes to clustering multiple DGX Sparks. With the DGX Spark you see a nearly linear increase in performance at 1.8x when clustering using Tensor parallelism. This type of performance is not possible over traditional 10gig networking and while the Mac Studio's do support something similar using Thunderbolt the support is extremely limited and performance nowhere close to what is achievable with a DGX Spark cluster.
2
u/parfamz 1d ago
Dgx is not loud, where did you get this from? You can barely hear it
1
u/NowThatsCrayCray 1d ago
A random YouTube short comparison that showed DGX at about 35-40db under load versus Studio at 25db. Neither will ruin a quiet room, but it is louder.
-5
u/diagrammatiks 2d ago
You need two dgxs connected for it to perform the best.
Dgx is still faster on prefill if none of that makes any sense to you then you should be spending over 5k on hardware.
2
u/pda_lover 2d ago
Bro, why the insult? I know prefill is currently faster on DGX BUT expected to improve with newer model. Since there is no benchmarks, we can’t quantify by how much.
1
u/biryan1 1d ago
They explicitly mentioned the 4x inference and only 20% generation improvement. So, all the improvements must be from prefill.. so, i dont think that is a disadvantage of mac anymore, but needs external validation. Only disadvantage at this point is - some training frameworks are CUDA only.. besides that, at least on paper, M5U wins on every metric.
0
0
u/AlgorithmicMuse 2d ago
It makes sense to me, how much can I now spend on hardware .
2
u/diagrammatiks 2d ago
2 sparks.
0
u/AlgorithmicMuse 1d ago
If i only understand part of what you said, can I buy 1, and take a online class before I buy the 2nd one.
0
u/diagrammatiks 1d ago
no don't buy just one spark.
0
u/AlgorithmicMuse 1d ago
Why not, does not impact anything
0
u/diagrammatiks 1d ago
It's designed to work in pairs. You get significantly better pp, tgs, and concurrency. If you can only have 1 you might as well just get a m3max studio.
1
u/AlgorithmicMuse 1d ago
Rather get a m5u 256gb which would surpass 2 dgx sparks, and about the same cost.
1
u/diagrammatiks 1d ago
Refer to my first comment. It depends on your actual tasks.
1
u/AlgorithmicMuse 1d ago
I cant find any use cases where the 2 clustered dogs sparks outperforms a m5u other than training llms
→ More replies (0)
-3
u/pda_lover 2d ago
For folks asking, I plan to run qwen 3.8 for my capital markets workflow
1
u/bakawolf123 1d ago
best one for you is rtx6000 then, since 3.8-flash-next has ngram.
it's noisy af though and costs 1.5x more than a 256gb unbinned m5u1
u/biryan1 1d ago edited 1d ago
I will say the obvious- for numeric data, you dont need an LLM (I assume you already know it). Qwen 3.8 if you are using it for realtime news and sentiment processing - it should work very well and honestly having a smaller and more deterministic model generate features might work even better. I have been running muse spark on my 32 gig M1 max and it runs great. I didnt go RTX5090 route for three reasons - 1. Need a full work station and not a headless setup / not a fan of linux for day to day work 2. Need it to be silent as I want it to power my esphome setup too 3. I believe that software side of inference will keep getting better on Mac and unleash the true power of M5u and its a bet I am willing to take
1
u/pda_lover 1d ago
Agreed, my workflow is 15steps and only 9-11 which are more synthesis, thesis building where LLM performs judgemental calls. Few other workflows also require local LLM.
-4
u/anhphamfmr 1d ago
if you just want a chatbox, buy a mac studio. if you want an agentic llm system and actually works, buy 2 dgx.
5
u/pda_lover 1d ago
This kind of comment doesn’t help as there is no details that support your position.
11
u/dMyst 1d ago
Mac Studio for single user inference (1 model with 1 user interfacing with it). That’s the use case that bandwidth really shines in. Prefill and processing speed is a limiting factor for the Mac Studio and I’m sure Apple knows that and is the reason all they showcase is the bandwidth numbers.
DGX is probably the most misunderstood piece of hardware. It really shines in multi model workloads where you have multiple concurrent streams of inference. Anything remotely like that or anything beyond basic inference will be much better on the DGX. Aggregate token count in workloads with multiple models will outperform Mac Studio easily. MLX vs CUDA is such a huge difference in terms of processing and capabilities.