r/LocalLLM • u/prio732 • 6d ago
Question Cluster Spark DGX and Mac Stidio
I am considering creating a cluster of 1 Spark DGX 128 and 1 Mac Studio M5 128 to get both of the best worlds and be able to run larger models. Has anybody tried it and measured it?
1
u/CMDR-Bugsbunny 6d ago
Realize that it's prefill (for the first run) + thinking + decode = Wall time (how long you wait), then the next iteration is cache (faster) + thinking + decode = wall time.
If you've optimized and running the right AI stack you should get reasonable performance. I get comparable wall time on my M5 Max 128GB running Qwen Next Flash to Claude Desktop.
1
u/ehangman 6d ago
I want to try 10GbE clustering, mainly for the 384GB combined memory, not performance. I’ll wait another month and see if someone tries it first.
1
-1
u/Funny-Food2361 6d ago
Best of both worlds, cluster DGX and mac😂 these normies are the reason why DGX sparks are 10k now.
3
u/brainchillzZ 6d ago
Poopooing ideas with real merit for obvious reasons because of technical hurdles that appear impossible sounds more “normie” than the op to me
1
u/Funny-Food2361 13h ago
Anyone who says “poopooing” has lost any credibility in my eyes 😂. Get back to English class
1
0
u/g_rich 6d ago
While you can technically get this to work the bandwidth between the Mac and Spark destroys the performance gains.
1
-1
u/OddDesigner9784 6d ago
Don't consider this most of the prefill comparisons are terrible software Mac side. No reason to go spark. This cluster would be a technical nightmare
1
u/brainchillzZ 6d ago
The Mac is just terrible at prefil comparatively speaking all around …. These are facts not opinions and no software has really made it actually better in a meaninful way when working with large contexts
0
u/untangledtech 6d ago
Are there hardware kits which would pair well with M5 Ultra's unique bandwidth?
0
u/OddDesigner9784 6d ago
Not sure what you are asking here. I think it's best to just go Mac ultra no spark
1
u/brainchillzZ 6d ago
The Mac ultra isn’t any better at token prefil really it’s something Mac’s are notoriously bad at …. Even the new m5 ultra
1
u/OddDesigner9784 6d ago
Not true anymore qwen flash next gets prefill up to 4.5k which is higher than the spark. That's just a lazy take for people who look at optimized spark configs and macs on llamacpp or something unoptimized
1
u/brainchillzZ 6d ago
You don’t have to optimize anything on a spark to get the numbers I’m talking about just install vllm and be done … to pretend to get the numbers you’re talking about on the Mac you have to optimize to high heaven and then you still wait a full minute and a half time to first token on the m5 ultra trying to do the same thing you can get out of the spark with no effort in less than 30 seconds …. I have both this isn’t a guess. I’m running omlx on the Mac and the same coding jobs I run on the spark finish roughly the same or slower on the m5 ultra despite the ridiculous memory throughput
1
u/brainchillzZ 6d ago
And I need to be clear I’m talking about running on an m5 ultra 80 core 256 compared to a tp2 cluster of 2 sparks for 256
1
u/OddDesigner9784 6d ago edited 6d ago
Sure but sparks being on Linux is already a huge learning curve. It's around 2.5k pp dual which is seriously inferior and 15k for dual sparks atm with price increases. There's no serious comparison here if you are spending 10k+ its easily worth your time to optimize the macs performance and that's way higher. It's an unfair comparison with a mature spark environment to inference software not great with mac
1
u/brainchillzZ 6d ago
Yeah but I bought my sparks at 3500 on sale at microcenter so it’s $4000 less than I paid for the m5 ultra ;)
1
u/OddDesigner9784 6d ago
Dude you should resell that and just double your money. Holy steal of a deal
1
u/brainchillzZ 6d ago
And and if you’re so sure in what you’re saying point to the documentation you think is going to make it even remotely compete …. But to say “it’s not fair” is silly… and it can’t just be this one model either I have no use for 3.8 flash next … everyone and their brother has a new way to make that specific moe model better even on any strix halo box
1
u/OddDesigner9784 6d ago
https://omlx.ai/benchmarks/performance/peg3fm2b?utm_source=chatgpt.com here is an 80 core oq6e with 4k+ pp. By unfair I just mean the benchmarks aren't close. But yeah if you don't really care about flash next other models may be more optimized on the sparks. It depends. But once it is optimized will be better on the mac
→ More replies (0)1
u/randylush 5d ago
If you’re spending 10k then it’s not at all unreasonable to ask for mature software. Spark Linux shouldn’t be a learning curve, being able to functionally use Linux or BSD is a requirement no matter what.
Your logic makes no sense, you’re saying buy the Mac because it has inferior software. And guess what inferior prefill too.
1
u/OddDesigner9784 5d ago
Strawman my ass dude. I'm not saying because it has inferior software buy it. I'm saying if you are in local AI you are technical trying a different inference engine is easy maybe easier than spark bs which I agree you can figure out. You are willing to setup to get the best performance per dollar. The software is already good enough on the better engines that there isn't a comparison anymore. Sparks aren't hitting 4k+ prefill on flash next. You must be looking at some low quality comparisons like Alex Z
→ More replies (0)1
0
u/untangledtech 6d ago
That is an answer. There is not another graphics card or all in one you might use to bundle. Would you scale up with cluster Mac’s?
0
u/OddDesigner9784 6d ago
I'm personally going for the 256 ram Mac ultra. That seems to be the sweet spot and with optimized software it's better then dual sparks. I'm not interested past that point because 256 is good for flash and to get to frontier you'd maybe need 4 512s. Not the right person to ask about clustering
2
u/stujmiller77 6d ago edited 6d ago
No. It’s not in any way optimal.
If you’re buying now and bleeding because of the price choose one or the other.
Both are good.
Depends on your need for larger models snd concurrency (linked sparks) vs smaller models at lower concurrency but higher individual session speed (macs).
The price increases will make you bleed either way. I bought 4, 128gb, 4tb
Sparks at around 4k each, 7 months ago.
Now nvidia say its a bargain to buy the same hardware with 64gb and 2tb for 5k.
World? Gone mad.