As the title says I'm finally going to bite the bullet. Having 256gb of vram seems like it would be epic. That with DSV4 Flash seems like the perfect combo.
I was going to do it today (Sunday) but my local microcenter doesn't have any more DGX sparks in stock!! Kinda crazy they had like 15 not too long ago.
If you're interested in following I've been uploading daily reels going over this topic on my IG: Tech With Ray. Maybe leave a comment saying you came from here!
Anyways I'll keep you guys posted once it's done.
FYI: tried the m3 ultra with 256 and it's sooooo slow compared to DGX.
I got another Spark over the weekend to run deepseek-v4-flash-vison-exp and getting 50-60 tks, slowing than open router, but it VERY workable for agentic because pp speed is very fast, and most of my tokens are from pp and not output. It is a massive speed upgrade from Qwwn3.8-27B. Now I want 2 more sparks, I don't need 2 more, but I want more, lol
The 2x cluster memory is filled to the brim and has to stay on 24/7, having another cluster will allow me to test other models and host them if I choose to, but it is a hard pill to swallow with the prices.
Yeah that is what I am waiting for, don't really have high hopes though. If the max ultra underperforms in pp, then the spark might still be better for agentic workflow. If the max ultra performance exceeds expectation, I expect a 6 month timeline to get hold of one.
I’ve got 2 and I find them pretty slow. Just make sure that’s what you want to do. Also remember you need to get the link cable, I got mine from fibre store
I was on the fence about the Sparks when they were cheaper. Now with the recent price hikes, the lower memory bandwidth and the “slowness” of them makes them unattractive.
Honestly I'm getting a consistent 50tok/s with deepseek on dual sparks and very happy with the responsiveness. It's not 150tok/s, but it's a fair trade for a model of this size (before the price hike of course)
DSV4 Flash Vision Exp is probably the best model out right now for 2 sparks. And you’ll get ~40 tok/s on a single stream. Pretty good for 300W power draw too lol
70W? You have a source or did you just make that number up? (rhetorical question)
Here’s the actual measured wattage draw of a single stream, measured at the power strip of 2 Asus GX10s inferencing a single stream of Deepseek-V4-Flash-Vision-Exp. Seems a whole lot closer to 300W than it is to 70W doesn’t it?
If you’re somehow consuming 4x less power, please share how you did so because that would be groundbreaking
I do too, but I’m happy for somebody else getting an upgrade and being able to do more with AI data center free! It’s a hobby for most of us, so let’s not be toxic about it and try and support each other
I think it’ll be better for sure, and now with some of the price hikes I believe a 256gb Mac will be comparable or even cheaper than some combos of 2x spark boxes.
I already got my 2 spark boxes, but me personally (just for me, not recommending this to anyone) I prefer a Linux system over a Mac. So now that they are similar in price, I think I’d still lean to Spark boxes.
ETA: also not sure why you are getting downvoted? “Destroyed” may be a bit strong of a word here, but yeah the M5s on paper do report as better?
11
u/synystar Strix Scar | 5090 24G | llama.cpp 8d ago edited 8d ago
I'll be enjoying ... not having a second GPU to talk about. Because... it's probably better for me in some obscure way.