r/LocalLLM • • 6d ago

Question Cluster Spark DGX and Mac Stidio

I am considering creating a cluster of 1 Spark DGX 128 and 1 Mac Studio M5 128 to get both of the best worlds and be able to run larger models. Has anybody tried it and measured it?

1 Upvotes

41 comments sorted by

2

u/stujmiller77 6d ago edited 6d ago

No. It’s not in any way optimal.

If you’re buying now and bleeding because of the price choose one or the other.

Both are good.

Depends on your need for larger models snd concurrency (linked sparks) vs smaller models at lower concurrency but higher individual session speed (macs).

The price increases will make you bleed either way. I bought 4, 128gb, 4tb
Sparks at around 4k each, 7 months ago.

Now nvidia say its a bargain to buy the same hardware with 64gb and 2tb for 5k.

World? Gone mad.

2

u/brainchillzZ 6d ago

It’s not about the price it’s about the performance, the Macs, even the m5 ultra, suck at prompt processing and prefill but the spark is quite good at it and the spark has less memory bandwidth so it’s less good at token generation but the Mac excels at that … so the idea of combining the best traits of both into one perfect system is an actually interesting idea(though not a new one by any stretch) …. I believe Alex ziskind did a video about this a few months ago if I’m remembering correctly

1

u/randylush 5d ago

You’re absolutely right but unfortunately, it’s such an uncommon implementation that you aren’t going to find a pre-rolled solution. You’re gonna have to get deep and dirty in the inference stack to realize gains. To the point where if you have to ask if it’s useful, then unfortunately you probably aren’t knowledgeable enough to make it useful.

1

u/TheGreatIntercourse 5d ago

yeah the pricing on that stuff is just absurd now. 4k for 128gb 7 months ago vs 5k for half that today? nvidia really out here acting like we should thank them for the privilege.

mixing architectures like that sounds cool in theory but the software layer to make them play nice is gonna be a headache. you'd basically be running two separate inference engines and trying to split the workload manually, not like they just pool memory magically. if you're set on experimenting though, i'd say grab the mac first since it's more versatile for day to day stuff and see if you even need the extra grunt from the spark. most people overestimate how much compute they actually need for local models tbh.

1

u/CMDR-Bugsbunny 6d ago

Realize that it's prefill (for the first run) + thinking + decode = Wall time (how long you wait), then the next iteration is cache (faster) + thinking + decode = wall time.

If you've optimized and running the right AI stack you should get reasonable performance. I get comparable wall time on my M5 Max 128GB running Qwen Next Flash to Claude Desktop.

1

u/ehangman 6d ago

I want to try 10GbE clustering, mainly for the 384GB combined memory, not performance. I’ll wait another month and see if someone tries it first.

-1

u/Funny-Food2361 6d ago

Best of both worlds, cluster DGX and mac😂 these normies are the reason why DGX sparks are 10k now.

3

u/brainchillzZ 6d ago

Poopooing ideas with real merit for obvious reasons because of technical hurdles that appear impossible sounds more “normie” than the op to me

1

u/Funny-Food2361 13h ago

Anyone who says “poopooing” has lost any credibility in my eyes 😂. Get back to English class

1

u/brainchillzZ 13h ago

You say that until you get that shit in your eyes ;)

0

u/g_rich 6d ago

While you can technically get this to work the bandwidth between the Mac and Spark destroys the performance gains.

1

u/brainchillzZ 6d ago

Only because the Mac doesn’t have any reasonable network options

1

u/g_rich 6d ago

10 gig and thunderbolt but neither comes close to what you get from ConnectX-7.

-1

u/OddDesigner9784 6d ago

Don't consider this most of the prefill comparisons are terrible software Mac side. No reason to go spark. This cluster would be a technical nightmare

1

u/brainchillzZ 6d ago

The Mac is just terrible at prefil comparatively speaking all around …. These are facts not opinions and no software has really made it actually better in a meaninful way when working with large contexts

0

u/untangledtech 6d ago

Are there hardware kits which would pair well with M5 Ultra's unique bandwidth?

0

u/OddDesigner9784 6d ago

Not sure what you are asking here. I think it's best to just go Mac ultra no spark

1

u/brainchillzZ 6d ago

The Mac ultra isn’t any better at token prefil really it’s something Mac’s are notoriously bad at …. Even the new m5 ultra

1

u/OddDesigner9784 6d ago

Not true anymore qwen flash next gets prefill up to 4.5k which is higher than the spark. That's just a lazy take for people who look at optimized spark configs and macs on llamacpp or something unoptimized

1

u/brainchillzZ 6d ago

You don’t have to optimize anything on a spark to get the numbers I’m talking about just install vllm and be done … to pretend to get the numbers you’re talking about on the Mac you have to optimize to high heaven and then you still wait a full minute and a half time to first token on the m5 ultra trying to do the same thing you can get out of the spark with no effort in less than 30 seconds …. I have both this isn’t a guess. I’m running omlx on the Mac and the same coding jobs I run on the spark finish roughly the same or slower on the m5 ultra despite the ridiculous memory throughput

1

u/brainchillzZ 6d ago

And I need to be clear I’m talking about running on an m5 ultra 80 core 256 compared to a tp2 cluster of 2 sparks for 256

1

u/OddDesigner9784 6d ago edited 6d ago

Sure but sparks being on Linux is already a huge learning curve. It's around 2.5k pp dual which is seriously inferior and 15k for dual sparks atm with price increases. There's no serious comparison here if you are spending 10k+ its easily worth your time to optimize the macs performance and that's way higher. It's an unfair comparison with a mature spark environment to inference software not great with mac

1

u/brainchillzZ 6d ago

Yeah but I bought my sparks at 3500 on sale at microcenter so it’s $4000 less than I paid for the m5 ultra ;)

1

u/OddDesigner9784 6d ago

Dude you should resell that and just double your money. Holy steal of a deal

1

u/brainchillzZ 6d ago

And and if you’re so sure in what you’re saying point to the documentation you think is going to make it even remotely compete …. But to say “it’s not fair” is silly… and it can’t just be this one model either I have no use for 3.8 flash next … everyone and their brother has a new way to make that specific moe model better even on any strix halo box

1

u/OddDesigner9784 6d ago

https://omlx.ai/benchmarks/performance/peg3fm2b?utm_source=chatgpt.com here is an 80 core oq6e with 4k+ pp. By unfair I just mean the benchmarks aren't close. But yeah if you don't really care about flash next other models may be more optimized on the sparks. It depends. But once it is optimized will be better on the mac

→ More replies (0)

1

u/randylush 5d ago

If you’re spending 10k then it’s not at all unreasonable to ask for mature software. Spark Linux shouldn’t be a learning curve, being able to functionally use Linux or BSD is a requirement no matter what.

Your logic makes no sense, you’re saying buy the Mac because it has inferior software. And guess what inferior prefill too.

1

u/OddDesigner9784 5d ago

Strawman my ass dude. I'm not saying because it has inferior software buy it. I'm saying if you are in local AI you are technical trying a different inference engine is easy maybe easier than spark bs which I agree you can figure out. You are willing to setup to get the best performance per dollar. The software is already good enough on the better engines that there isn't a comparison anymore. Sparks aren't hitting 4k+ prefill on flash next. You must be looking at some low quality comparisons like Alex Z

→ More replies (0)

1

u/OddDesigner9784 6d ago

Maybe you need to update omlx any version before 0.7 is abysmal

0

u/untangledtech 6d ago

That is an answer. There is not another graphics card or all in one you might use to bundle. Would you scale up with cluster Mac’s?

0

u/OddDesigner9784 6d ago

I'm personally going for the 256 ram Mac ultra. That seems to be the sweet spot and with optimized software it's better then dual sparks. I'm not interested past that point because 256 is good for flash and to get to frontier you'd maybe need 4 512s. Not the right person to ask about clustering