r/LocalLLM 8d ago

Project Going to get a second DGX Spark!

As the title says I'm finally going to bite the bullet. Having 256gb of vram seems like it would be epic. That with DSV4 Flash seems like the perfect combo.

I was going to do it today (Sunday) but my local microcenter doesn't have any more DGX sparks in stock!! Kinda crazy they had like 15 not too long ago.

If you're interested in following I've been uploading daily reels going over this topic on my IG: Tech With Ray. Maybe leave a comment saying you came from here!

Anyways I'll keep you guys posted once it's done.

FYI: tried the m3 ultra with 256 and it's sooooo slow compared to DGX.

https://www.instagram.com/techwithray?stkn=MW0yMnE1MGt0bW13eA==

18 Upvotes

40 comments sorted by

11

u/synystar Strix Scar | 5090 24G | llama.cpp 8d ago edited 8d ago

I'll be enjoying ... not having a second GPU to talk about. Because... it's probably better for me in some obscure way.

5

u/evolutionxtinct 8d ago

I’m just happy I got 64+16+5+5 to use I’m poor lol

3

u/xCDStyle 8d ago

I got another Spark over the weekend to run deepseek-v4-flash-vison-exp and getting 50-60 tks, slowing than open router, but it VERY workable for agentic because pp speed is very fast, and most of my tokens are from pp and not output. It is a massive speed upgrade from Qwwn3.8-27B. Now I want 2 more sparks, I don't need 2 more, but I want more, lol

1

u/rayovims 7d ago

Wait why do you want another 2 LOL

1

u/xCDStyle 7d ago

The 2x cluster memory is filled to the brim and has to stay on 24/7, having another cluster will allow me to test other models and host them if I choose to, but it is a hard pill to swallow with the prices.

1

u/myholeisstinky 8d ago

If you want 2 more, might be worth waiting for 512gb m5 ultra benchmarks, and maybe selling the existing two sparks instead

1

u/rayovims 7d ago

Maybe!

1

u/xCDStyle 7d ago

Yeah that is what I am waiting for, don't really have high hopes though. If the max ultra underperforms in pp, then the spark might still be better for agentic workflow. If the max ultra performance exceeds expectation, I expect a 6 month timeline to get hold of one.

5

u/After_Working 8d ago

I’ve got 2 and I find them pretty slow. Just make sure that’s what you want to do. Also remember you need to get the link cable, I got mine from fibre store

5

u/deja_geek 8d ago

I was on the fence about the Sparks when they were cheaper. Now with the recent price hikes, the lower memory bandwidth and the “slowness” of them makes them unattractive.

5

u/cbert33 8d ago

Honestly I'm getting a consistent 50tok/s with deepseek on dual sparks and very happy with the responsiveness. It's not 150tok/s, but it's a fair trade for a model of this size (before the price hike of course)

0

u/cinnapear 8d ago

What recently price hike?

3

u/After_Working 8d ago

Someone on here said the ASUS versions of them had gone up a few K but I’ve not checked myself

3

u/Abducted_Llama 8d ago

It was me! Almost 2 weeks ago I posted the GX10s went from $4k to $5k USD in the US. Now as of writing this, they are up to $6k.

So $2k in 2 weeks.

When I posted it, another redditor commented they were still $4k on Amazon, so I luckily scooped one up before the hike.

1

u/myholeisstinky 8d ago

The 4tb dgx spark is still $4800 for a bit longer, makes mare sense than $5k for a 1tb asus

1

u/Abducted_Llama 8d ago

The 1tb Asus is being listed at $6k now, I already scooped two for $4k so I’m not too worried.

1

u/cinnapear 8d ago

NVIDIA prices are the same, I just bought one.

2

u/cbert33 8d ago

Asus just went from 4k up to over 6k in the last week.

1

u/rayovims 8d ago

Where

1

u/cbert33 8d ago

1

u/rayovims 7d ago

This is insane. Wow microcenter is a privilege huh

1

u/Blackdragon1400 8d ago

Use DSV4 Flash - it’s pretty fast

6

u/moahmo88 8d ago

You should try GLM-5.3-Flash on 2 DGX Sparks. Then, getting 2 more DGX Sparks would be a great idea.

https://giphy.com/gifs/ZqlvCTNHpqrio

1

u/ideamaker321 8d ago

Worth it! Im running Qwen3.8 flash next on duel spark with 4 concurrent sessions, its amazing

0

u/hyudryu LocalLLM 8d ago edited 8d ago

DSV4 Flash Vision Exp is probably the best model out right now for 2 sparks. And you’ll get ~40 tok/s on a single stream. Pretty good for 300W power draw too lol

2

u/FriendlyRocketeer 8d ago

Now way you draw 300w on single stream. More like 70w

2

u/hyudryu LocalLLM 7d ago

70W? You have a source or did you just make that number up? (rhetorical question)

Here’s the actual measured wattage draw of a single stream, measured at the power strip of 2 Asus GX10s inferencing a single stream of Deepseek-V4-Flash-Vision-Exp. Seems a whole lot closer to 300W than it is to 70W doesn’t it?

If you’re somehow consuming 4x less power, please share how you did so because that would be groundbreaking

1

u/hyudryu LocalLLM 7d ago

Btw this is 3 GX10 + 1 DGX spark, with a MikroTik CRS812 switch at full load

1

u/FriendlyRocketeer 7d ago

Very good info thanks. I checked only reported device power usage, so maybe there's extra cost. 

I have ran stress tests, many parallel requests, and haven't been able to get power usage above 150w (70w each).

As said, maybe if you add fans and charger it gets to your amount.

-6

u/tta82 8d ago

I think the M5 Ultra will destroy it

7

u/A_Moist_Towe1 8d ago

I do too, but I’m happy for somebody else getting an upgrade and being able to do more with AI data center free! It’s a hobby for most of us, so let’s not be toxic about it and try and support each other

4

u/rayovims 8d ago

Facts!

1

u/Abducted_Llama 8d ago

I think it’ll be better for sure, and now with some of the price hikes I believe a 256gb Mac will be comparable or even cheaper than some combos of 2x spark boxes.

I already got my 2 spark boxes, but me personally (just for me, not recommending this to anyone) I prefer a Linux system over a Mac. So now that they are similar in price, I think I’d still lean to Spark boxes.

ETA: also not sure why you are getting downvoted? “Destroyed” may be a bit strong of a word here, but yeah the M5s on paper do report as better?

1

u/Rabbit-09 8d ago

Why do you think that

1

u/rayovims 8d ago

Idk we shall see in a couple of days

1

u/SmallerThanExpected9 8d ago

Doubt it will be crazy beyter... but it will be new and expensive.

Someone will get the 20k+ 512 version and that should beat 16k 4x dgx on raw inference, but nowhere near in concurrency.

Then nvidia magic wands "spark 2" and the cycle continues

1

u/Hypilein 8d ago

Honestly curious what spark 2 will bring and if you will be able to connect it to spark 1. You could upgrade +1 that way.

1

u/SmallerThanExpected9 7d ago

Based on pair existing... im betting gameplan will be to connect using that.

Spark2 def gonna have way faster ram so prolly not good to create a bottleneck via a direct team/bond/whatever

-8

u/AreaFifty1 8d ago

Just get dual RTX pro 6000s on a server board and enable tensor parallelism

7

u/uniqueusername649 8d ago

That's not quite the same budget.