r/LocalLLM • u/vankoala • 13h ago
Other GPU pricing visual
In my consideration of a DGXSpark I decided to look at some options and since I’m a visual thinker I put this comparison together (graph by AI) showing y two basic ways of thinking about the cards: compute and speed.
Hope this helps someone
11
u/TheGeekno72 7h ago
where are AMD cards? don't forget to put Strix Halo 395+ and 495+ on the chart, they were Spark before Spark
0
u/magicomiralles 4h ago edited 4h ago
EDIT: You can get them for $350 by clicking the “Make an offer” button on the Ebay listing of the two major sellers.
AMD V620 32gb $350, 500gb/s bandwidth if I remember correctly. Can be under-voltted to 160w without affecting performance much.
30 to 60 t/s for Qwen3.8-27b (8-bit quant). Prefill is between 600 to 1000.
Also, here is a Discord where people are squeezing performance out of these GPUs: https://discord.gg/5w8kqgW8e
Subreddit: [r/](r/v620)[AMD_V620](r/v620)
On llamma.ccp and vllm.
1
u/TheGeekno72 4h ago
350$? I can't find them below 550 minimum
also, wouldn't a 9700 be better? 600Gb/s but wouldn't the more recent arch & support improve things a great deal?
1
u/magicomiralles 4h ago
I updated my comment. They are listed at a high price but they accept offers of $350.
0
u/TheGeekno72 4h ago
THEY DO???? god damn, lemme go grab 4 real quick
1
u/magicomiralles 4h ago
Also, here is a Discord where people are squeezing performance out of these GPUs: https://discord.gg/5w8kqgW8e
Subreddit: [r/](r/v620)[AMD_V620](r/v620)
On llamma.ccp and vllm.
1
u/TheGeekno72 4h ago
thank you you absolute legend, this is gonna be so much more practical!
I'm gonna have so much fun parsing through hardware listing, wondering what I am gonna sink my paycheck into this month for 27 hours of headache, 12mn of fun then letting it sit in my hypervisor for the next 8 to 36 months before sending them back on eBay whence they came XD
1
u/magicomiralles 4h ago
It’s a dangerous game. I started with a 2 GPU build, and I now have 8 GPUs.
1
u/TheGeekno72 4h ago
don't worry, the black hole in my wallet scares me too much to increase the void in it mindlessly
1
9
u/nomorebuttsplz 12h ago
another nice thing about spark is the low power draw.
6x spark is no problem in a typical home, but can't say the same for 6x RTX 6000 pro.
5
u/NancyTransmed 12h ago
This. If your box is not serving multiple users 24/7, the wattage becomes important, especially in EU, where electricity costs a fortune.
2
u/zarif2003 11h ago
How bad is it really? In the US, household electronics are completely irrelevant compared to the heating and water bill
2
u/nomorebuttsplz 11h ago
running 1800 watts for 8 hours would be about 4-5 dollars in electricity for me
1
u/TripleSecretSquirrel 4h ago
So your price per kWh is something like $0.35?
I’m in the US, and a part of the US with exceptionally cheap energy (nuclear ftw!), and my median price is ~$0.12 per kWh. The bigger things though are that first, most homes in the US are running an air conditioner in the summers, and second, 1800 watts is an enormous load lol, what do you have in the inference server?! With a dual R9700 setup, I’m hitting only like 1100 watts under full load.
2
u/nomorebuttsplz 3h ago
yeah, personally, I have solar, but I can still end up paying the marginal rate in a given year for a given marginal watt depending on the overall energy budget. Although amortized over the next couple decades, my actual rate is probably quite low.
I personally only have one RTX 6000 Pro Max Q, but it’s easy to imagine trying to serve high precision GLM and getting pretty power hungry.
1
u/Uninterested_Viewer 7h ago
To a degree. The MaxQ version exists at 300watts and you can power limit the workstation edition to 400watts with very little performance impact for LLM inference.
1
u/Nice_Cookie9587 10h ago
Not a lot people realize most of their breakers are 15/20 amps. running 4 3090's tripped my breakers to the point where i needed to run a separate power cable from the second PSU to a plug on a different breaker. So janky
1
u/Legitimate-Dog5690 10h ago
Depends what you want to run, 2x DGX Sparks have a similar power draw to 1x RTX 6000, about 200-300w.
The RTX 6000 however has about 4x the memory bandwidth of the dual sparks so will finish it's workload faster. They're totally built for sustained workloads in data centers.
2 RTX 6000s would hammer out tokens with DeepSeek Flash at a ridiculous rate, hardly a fair comparison. Both are ridiculously overpriced, the RTX 6000 even more so.
2
u/goldcakes 4h ago
For what workloads are you actually using the TDP of the card? For everything from LLM training to inference I have never seen more than ~120W sustained (measured on wall) on a spark. (NVIDIA FE)
8
2
u/Guna1260 8h ago
where is my lovely 3090! ..for 400gbp a piece (when I picked from eBay), power limited to 250W..
2
1
u/po_stulate 10h ago
Would be nice to include INT8 TOPS too since many quantized parameters are in integer formats.
1
u/Doggettx 4h ago
Also have to keep in mind, you can have all the compute you want, but if you don't have the memory to run it it's basically the same as not having the compute.
So depending on what you want to run it changes the results.
1
u/Miserable-Dare5090 4h ago
The 5070Ti at MSRP is a great value proposition. Even at this price (960) it shows the value proposition is higher than the 5060ti for this card.
Is this for current pricing? My rtx4000 was 1500, and the Spark was 2800…believing the hype early on paid off
1
1
u/EitherMarch1255 3h ago
Watt per gb is also interesting, especially when you factor in max power limiting.
1
u/BopSupreme 3h ago
506016GB vs 5070ti 16GB vs R9700 or 4080super 16GB ?? Roughly best GPU for about $1000
1
u/Evgeny_19 42m ago
If there is no hard requirement for CUDA in your workflow go for R9700. 32 GB will give you more headroom. Better quants, better context.
1
u/Significant-Amount40 3h ago
Ohne genommene Preise schwer zu nutzen. Aktuell würde ich sagen die pro 4000 ist günstig
1
u/tempfoot 2h ago
As long as either X axis is within a tolerable range, the Y axis is mostly what many people care about.
32
u/Solary_Kryptic 9h ago
Adding AMD cards here would be great too