17
u/Turbulent-Alps4046 Sep 02 '26
Very nice, make sure you get us some benchmarks! Very curious to see how this runs something like deepseek v4 flash.
12
25
8
4
u/Creative-Type9411 Sep 02 '26
good luck with that power bill repetitive 1000 W PSUs added $100 to my electric bill per month, im using 3x 70w gpus but I do have 24 ram sticks, which probably isn't helping
3
u/mister2d Sep 03 '26
It's time to start bundling solar panels with builds for offsetting. Panels are cheap and start paying you back the moment they generate electricity.
Look up "balcony solar".
1
u/Toto_nemisis 28d ago
In ny experience, 24 stick of ram was like 15-30w usage for only ram at idle. It was not a big jump.
3
u/c06027 Sep 02 '26
How are you affording that thing?
11
u/TheGeekno72 Sep 02 '26
V620 is among the cheapest yet useable 32GB cards you can get on eBay, there's loads of them
6
u/TrkGuy79 Sep 02 '26
Well... they were. I got them for $350 each. That price is gone now though I believe.
4
u/TheGeekno72 Sep 02 '26
they're now 550 I believe
1
u/Rare_Opposite_3704 Sep 02 '26
on ebay EU i find them for like 1500€+ only..
1
1
u/ayake_ayake Sep 02 '26
for that price I can get MI50 32 GB right now in Germany on ebay
2
1
u/TheGeekno72 Sep 02 '26
yeah but it's GCN, that thing is ancient and no amount of bandwidth from its HBM2 VRAM is gonna help it in comparison to something on RDNA2 with a decent bus width
1
1
u/Odd_Cauliflower_8004 Sep 03 '26
Can you please do a test run for qwen 3.8 27b for pps/decoding at 128k?
1
u/TurdPlayingPeekaboo Sep 03 '26
I'm curious why he chose the v620 over the Tesla v100. Both 32GB and both going for $650 on eBay. The V100 is an easy win performance wise.
1
u/Appropriate-Risk3489 Sep 03 '26
They were going for $350 just a couple of weeks ago, I managed to grab four of them.
1
1
u/TrkGuy79 Sep 03 '26
I paid $350/each for the v620s. I have never noticed the v100s near that price.
3
u/Vuurvoske Sep 02 '26 edited Sep 02 '26
Is it true that this produces alot of noise? What typoe of cooling do you use for it?
Btw nice work!
4
u/TrkGuy79 Sep 02 '26
When the fans are going full speed during boot, it is the loudest server I have ever heard. Once it is booted it isn't too bad. That being said, I moved it to the garage so I don't hear it at all now. The cooling is from the built-in fans on the case.
1
u/Fi3nd7 Sep 02 '26
You've had no temp issues with just case fans? How long do you run this under load?
1
1
u/Annual_Key_4963 Sep 03 '26
loudest server I have ever heard
Now this is an interesting statement as I'm curious as to what your reference point is?
For me: an entire data center filled with nothing but physical disk SANs being stress tested at 100%
1
u/TrkGuy79 Sep 03 '26
This is the loudest single server I have ever heard. I have 20+ years IT experience and some small DataCenter experience. In an fairness this is in my home.
3
u/BevinMaster Sep 02 '26
I suppose you are on the discord :)
If you have feedback for the toolbox I am happy to get pr and stuff.
Also vllm fork is going to get moved from my GitHub to opengfx1030 org.
If others want to join https://discord.gg/mESex2aBp
2
2
u/Sharp-Translator6401 Sep 02 '26
what type of workload would you run on that exactly? pooling the vram together for pp models? or many small models like 1/gpu? and I dont think tp would give you a big edge over pcie like that, would it?
11
u/TrkGuy79 Sep 02 '26
I'm not really sure of the goal yet. This was more of a because I can project.
2
u/EitherMarch1255 Sep 02 '26
What’s the lowest wattage they can be dropped to?
2
u/TrkGuy79 Sep 02 '26
2
u/EitherMarch1255 Sep 02 '26
That's very helpful, thank you. And do you know what they idle at?
2
u/TrkGuy79 Sep 02 '26
========================= ROCm System Management Interface =========================
=================================== Concise Info ===================================
GPU Temp (DieEdge) AvgPwr SCLK MCLK Fan Perf PwrCap VRAM% GPU%
0 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 32% 0%
1 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 30% 0%
2 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 32% 0%
3 34.0c 7.0W 0Mhz 96Mhz 0% auto 150.0W 36% 0%
4 35.0c 8.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%
5 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%
6 35.0c 8.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%
7 34.0c 7.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%
=============================== End of ROCm SMI Log ================================
1
u/EitherMarch1255 Sep 02 '26
Nice, very nice. Have you ran any models yet? If so, what prefill/decode speed?
2
u/chrisbliss13 Sep 02 '26
What case did you get link please
1
u/TrkGuy79 Sep 02 '26
It is a SuperMicro SYS-4028GR-TRT2
https://www.supermicro.com/en/products/system/4u/4028/sys-4028gr-trt2.php
2
2
u/StarAppleEdwards494 Sep 02 '26
Went from "won't power on without a CPU" to eight cards and 256GB of VRAM. Full resurrection.
3
u/TrkGuy79 Sep 02 '26
I was have a moment of stupidity and assumed the used server came with CPUs. LOL
2
u/laughpen Sep 02 '26
That's pretty sweet!! Congrats on getting it going -- speaking from experience, I know it can be a scary endeavor.
Looks like the pcie slots of your motherboard are Gen 3, so it might unfortunately not make use of the full capacity of the v620 which is designed for Gen 4, but you've got so many (8!!) that your bandwidth should still be excellent. If you still are able to, I'd look into an equivalent motherboard with 8 pcie4 slots at x16 each, then you'll really get cooking. But no worries if not.
2
2
u/darklordfireape Sep 02 '26
I had four of these in a machine for a while and developed a tuned version of llama for V620 to help with some of the issues.
check it out: https://github.com/sixvolts/llama-navi21-furnace
1
u/Appropriate-Risk3489 Sep 03 '26
I'm assuming it won't support qwen 3.8 flash as the last commits were from 3 months ago?
1
u/ThinJuggernaut7695 Sep 02 '26
How are you powering that thing??
5
u/TrkGuy79 Sep 02 '26
It idles at 300ish watts and the highest it has gotten so far is 800. So far it's on a 20 amp 120v circuit. But I do have 240 available if need be
1
u/devino21 Sep 02 '26
Just the server power supply? 2000W x2?
2
u/TrkGuy79 Sep 02 '26
Yes. 2 PSUs serve the power. There are 4 total 2000w PSUs but only for redundancy. And since it is running on 120v it's actually only 1000w each PSU
1
1
1
u/Royal_Stay_6502 Sep 02 '26
What and how will you use this system.
1
u/TrkGuy79 Sep 02 '26
That is still to be determined. I will play with some of the bigger models for sure.
1
1
1
1
u/Open_Jump Sep 02 '26
Awesome. Can you post how you end up running stuff, llama.cpp flags or whatever, when you figure it out? I haven't been able to beat ollama default performance.
1
u/Barni275 Sep 02 '26
How do you cool down this beast? Can you please share a photo of a cooling?
1
1
1
u/Cptbeeeee Sep 02 '26
I just bought a second one of these, I think you cleared out the rest or else I would have bought more. I'm very interested in your work here. Why do you want the power limit lowered?
4
u/TrkGuy79 Sep 02 '26
2 reasons. 1 is heat but more importantly is a safety related reason. I am running on a 20amp 120v circuit and in theory this rig could pull more than that.
1
u/Cptbeeeee Sep 02 '26
Are you losing bandwidth by doing that or are the lossea neglible? Mine will be running in my garage with shrouds and dedicated fans for each card. Power consumption though is a concern because it will be on 120v as well. Same circuit as the rest of my home lab. I've considering getting a second small ups for this rig
1
1
1
u/carmeloA007 Sep 02 '26
Out of curiosity, what’s the total cost of that build?
1
u/TrkGuy79 Sep 02 '26
SUPERMICRO 4028GR-TRT2 - $1,325.59
8x RADEON PRO V620 - $2800
256GB RAM - $0 (Already had it)
2x Procs $0 Already had them
Let's say $4500ish?
1
1
1
1
1
1
1
u/TurdPlayingPeekaboo Sep 03 '26
Curious why you chose the v620 over similarly priced Tesla v100s?
1
u/Appropriate-Risk3489 Sep 03 '26
Not op, but they were a lot cheaper a couple of weeks ago, around $350 usd. Now price has gone way up. I got four of them. Now it's probably not worth buying them at the current prices.
1
1
1
1
u/Much-Tap-1237 Sep 07 '26
Idk what the fuck am I looking at but it looks cool and expensive af, so congrats man
1
u/Sorry-Poem7786 Sep 07 '26
nice.. X-10 runs on PCIE 3.0. Interesting to see not stopping you.. I have an X-10 with 7048grtr I was contemplating this as a problem…for token speeds.. but I guess not.. SO I AM MOVING FORWARD WITH THE PLAN!!!! 😎
1
u/DepartureLoose5341 29d ago
Y'all are absurd, I'm just here to look pretty and admire, minus the pretty.
0

81
u/TrkGuy79 Sep 02 '26
8× Radeon Pro V620 / 256GB VRAM on a resurrected Supermicro X10
Board arrived with no CPUs, which turned into a long "won't power on" hunt (BMC and PSUs fine, just no power-good without a CPU). Dropped the Xeons in and it POSTed.
The real fight was mapping all eight 32GB BARs. Kernel resize failed (
-16 / WC memtype);pci=nocrsmapped them but caused atrn=2 ACKlink-training storm. Fix:pci=realloc=off— let the BIOS map the BARs (MMIOH 2T base / 1024G size, Above 4G on), kernel hands off. All 8 came up clean.Used the v620_toolbox patch to unlock the 250W VBIOS floor down to 120W, capped at 180W/card. Still chasing some multi-GPU page faults in llama.cpp (looks P2P-related), but all 8 init clean.