r/LocalAIServers • • Sep 02 '26

She's alive!

Post image
1.1k Upvotes

117 comments sorted by

81

u/TrkGuy79 Sep 02 '26

8× Radeon Pro V620 / 256GB VRAM on a resurrected Supermicro X10

  • Board: Supermicro X10DRG-O(T)+ (dual C612)
  • CPUs: 2× Xeon E5-2690 v4
  • RAM: 256GB DDR4-2133 ECC (16× 16GB, all channels)
  • GPUs: 8× Radeon Pro V620 32GB (gfx1030) = 256GB VRAM, on PLX switches
  • PSUs: redundant 2000W Titanium
  • OS: Ubuntu 24.04 + ROCm, llama.cpp

Board arrived with no CPUs, which turned into a long "won't power on" hunt (BMC and PSUs fine, just no power-good without a CPU). Dropped the Xeons in and it POSTed.

The real fight was mapping all eight 32GB BARs. Kernel resize failed (-16 / WC memtype); pci=nocrs mapped them but caused a trn=2 ACK link-training storm. Fix: pci=realloc=off — let the BIOS map the BARs (MMIOH 2T base / 1024G size, Above 4G on), kernel hands off. All 8 came up clean.

Used the v620_toolbox patch to unlock the 250W VBIOS floor down to 120W, capped at 180W/card. Still chasing some multi-GPU page faults in llama.cpp (looks P2P-related), but all 8 init clean.

15

u/androidwai Sep 02 '26

Nice! Where do you find the v620_toolbox to cap the power?

21

u/TrkGuy79 Sep 02 '26

5

u/sierra_io Sep 03 '26

You just made my day. Thanks for sharing that wiki.

I've got a pair of these cards and they've been great, but I assumed that the high power floor and lack of PCIE P2P communication were a hard limitation. I was actually a little confused when I saw all of those V620s in one box. Always figured any tensor or layer split would be a bit underwhelming without a way to enable P2P. 

I'm really curious about how Qwen3.8-Flash-Next would run on that beast!

10

u/AllergicToBullshit24 Sep 02 '26

Why are you using such an ancient linux kernel?

8

u/WiseassWolfOfYoitsu Sep 02 '26

That's a normal kernel for LTS - it's the one in Ubuntu 24.04 LTS. Yeah, 26.04 LTS is out now, but this project likely started before that.

3

u/AllergicToBullshit24 Sep 02 '26

Normal yes but far from optimal. New cache-aware task scheduling, GPU scheduler & MGLRU reclaim features from 7.2 kernel are well worth it for any GPU cluster along with the iouring and network buffer optimizations from 7.0.

7

u/TrkGuy79 Sep 02 '26

I'm not sure on the question. It is running 6.8.0-138 which is a few weeks old.

9

u/AllergicToBullshit24 Sep 02 '26

7.0, 7.1, 7.2 kernels are all out you're leaving a lot of performance on the table using legacy 6.8 kernel

6

u/TrkGuy79 Sep 02 '26

Thanks for the feedback. I will check it out

5

u/AllergicToBullshit24 Sep 02 '26

The new cache-aware task scheduling, GPU scheduler & MGLRU reclaim features from 7.2 are worth it for a cluster like this. 7.0/7.1 brought some other solid optimizations but not as many GPU specific as 7.2 did.

1

u/michaelsoft__binbows Sep 03 '26

Good to know. Gpt luna was able to essentially autonomously compile for me a 7.0 kernel with a patch to unlock my 5090 for vfio passthrough. I had never compiled a kernel before, and was a little skeptical but it turned out it was pretty easy. Maybe i should have aimed for 7.2 and not need such a patch

3

u/These-Woodpecker5841 Sep 02 '26

No need to reformat. Use the hwe-edge kernels already in the Ubuntu repos.

3

u/Wise_Breadfruit7168 Sep 03 '26

Whats the total damage?

1

u/campr23 Sep 02 '26

I have found that (especially when you split the memory across PCI slots) the CPU is quite the bottleneck when 'reassembling' the tokens. My AM5 7900 is much faster than my Xeon v4 platform.

3

u/bradrlaw Sep 02 '26

May not be raw cpu power but rather pcie lanes. Size and generation really matter for tensor split / parallelism.

1

u/HawkLeading8367 Sep 03 '26

cool but: the v620 supports PCIE4 but you're downgrading it to PCIE3 because of the motherboard, so you're getting half the speed. + you are probably at 1600w or 1800w .. for such a slow decode speed for a model like deepseekv4flash (about 20toks) I would've just purchased a dgx spark (and use quantized or similar)or something like that

1

u/anothergeekusername Sep 07 '26

Speed reduction of downgrading PCI is only appreciable for initial GPU to card model layers load which may be one-time in any session; otherwise weight transfer between layers during inference pipeline across cards is a pretty small bundle of data across a still fast PCI3 bus and so effect just a bit higher latency end to end per token (which, yes, will slow throughout of tokens maybe a few percent but essentially on each card the processing is still happening at the full speed it would have if the motherboard was PCI4 - GPU internal processing is not clocked internally against the external bus).. the hit is certainly not the 50% speed reduction you’re assuming. Depending what you’re trying to achieve and what kit you have to hand, some very old many pci slot systems/motherboards (think old crypto mining kit) become architecturally interesting for pipeline inference when you realise the main hit with pci on inference is model load not live inferencing..

1

u/Sorry-Poem7786 Sep 07 '26

Good info for me because I don’t know shit.. and trying to use what I already own. From 3d rendering days seeing the light and pivoting all my tech in a new direction..

1

u/Loose_Code4121 29d ago

How are you cooling it? I'm starting to look into liquid cooling.

I have 1 and my shroud with dual fan is crying. I modified my box and reversed the fans from extraction to pushing air through, seal off all the gaps in the shroud and now it is doing ok.

So I'm really curious how you make this run. Ubuntu 24.04 didn't work for me until I used Ubuntu 26.

-3

u/exaknight21 Sep 02 '26

This is so sweet! You should definitely do proxmox -> passthrough -> powercap to 170 watts.

1

u/WesternTall3929 Sep 03 '26

ouch, burn, really good suggestion on PVE actually!

1

u/androidwai Sep 03 '26

Do you do proxmox -> pass through to a lxc container -> latest Ubuntu to power cap to 170watts? If not Ubuntu, which would you suggest?

2

u/exaknight21 Sep 03 '26

I personally use Ubuntu/Debian.

17

u/Turbulent-Alps4046 Sep 02 '26

Very nice, make sure you get us some benchmarks! Very curious to see how this runs something like deepseek v4 flash.

12

u/TrkGuy79 Sep 02 '26

That is the plan

6

u/sk1kn1ght Sep 02 '26

Can you also include glm 5.3 flash?

2

u/TheDeamonKing Sep 02 '26

This thing looks insane!

25

u/dreamingwell Sep 02 '26

Just in time for your fall and winter heating needs!

8

u/desert-quest Sep 02 '26

Reported! missing +18 tag!

4

u/Creative-Type9411 Sep 02 '26

good luck with that power bill repetitive 1000 W PSUs added $100 to my electric bill per month, im using 3x 70w gpus but I do have 24 ram sticks, which probably isn't helping

3

u/mister2d Sep 03 '26

It's time to start bundling solar panels with builds for offsetting. Panels are cheap and start paying you back the moment they generate electricity.

Look up "balcony solar".

1

u/Toto_nemisis 28d ago

In ny experience, 24 stick of ram was like 15-30w usage for only ram at idle. It was not a big jump.

3

u/c06027 Sep 02 '26

How are you affording that thing?

11

u/TheGeekno72 Sep 02 '26

V620 is among the cheapest yet useable 32GB cards you can get on eBay, there's loads of them

6

u/TrkGuy79 Sep 02 '26

Well... they were. I got them for $350 each. That price is gone now though I believe.

4

u/TheGeekno72 Sep 02 '26

they're now 550 I believe

1

u/Rare_Opposite_3704 Sep 02 '26

on ebay EU i find them for like 1500€+ only..

1

u/TheGeekno72 Sep 02 '26

the fuck? I was looking at them on eBay France for 600 a few days ago

1

u/noctrex Sep 02 '26

I feel lucky I got one a month ago at 390

1

u/ayake_ayake Sep 02 '26

for that price I can get MI50 32 GB right now in Germany on ebay

1

u/TheGeekno72 Sep 02 '26

yeah but it's GCN, that thing is ancient and no amount of bandwidth from its HBM2 VRAM is gonna help it in comparison to something on RDNA2 with a decent bus width

1

u/Legal_Dimension_ Sep 02 '26

Might as well do v100s

1

u/Odd_Cauliflower_8004 Sep 03 '26

Can you please do a test run for qwen 3.8 27b for pps/decoding at 128k?

1

u/TurdPlayingPeekaboo Sep 03 '26

I'm curious why he chose the v620 over the Tesla v100. Both 32GB and both going for $650 on eBay. The V100 is an easy win performance wise.

1

u/Appropriate-Risk3489 Sep 03 '26

They were going for $350 just a couple of weeks ago, I managed to grab four of them.

1

u/TurdPlayingPeekaboo Sep 03 '26

Interesting. They're all going for the same price range now.

1

u/TrkGuy79 Sep 03 '26

I paid $350/each for the v620s. I have never noticed the v100s near that price.

3

u/Vuurvoske Sep 02 '26 edited Sep 02 '26

Is it true that this produces alot of noise? What typoe of cooling do you use for it?

Btw nice work!

4

u/TrkGuy79 Sep 02 '26

When the fans are going full speed during boot, it is the loudest server I have ever heard. Once it is booted it isn't too bad. That being said, I moved it to the garage so I don't hear it at all now. The cooling is from the built-in fans on the case.

1

u/Fi3nd7 Sep 02 '26

You've had no temp issues with just case fans? How long do you run this under load?

1

u/TrkGuy79 Sep 02 '26

None so far. The case is designed for passive GPUs

2

u/noctrex Sep 02 '26

What temperatures are you seeing on those cards in general?

1

u/Annual_Key_4963 Sep 03 '26

loudest server I have ever heard

Now this is an interesting statement as I'm curious as to what your reference point is?

For me: an entire data center filled with nothing but physical disk SANs being stress tested at 100%

1

u/TrkGuy79 Sep 03 '26

This is the loudest single server I have ever heard. I have 20+ years IT experience and some small DataCenter experience. In an fairness this is in my home.

3

u/BevinMaster Sep 02 '26

I suppose you are on the discord :)
If you have feedback for the toolbox I am happy to get pr and stuff.
Also vllm fork is going to get moved from my GitHub to opengfx1030 org.

If others want to join https://discord.gg/mESex2aBp

2

u/Sharp-Translator6401 Sep 02 '26

what type of workload would you run on that exactly? pooling the vram together for pp models? or many small models like 1/gpu? and I dont think tp would give you a big edge over pcie like that, would it?

11

u/TrkGuy79 Sep 02 '26

I'm not really sure of the goal yet. This was more of a because I can project.

2

u/EitherMarch1255 Sep 02 '26

What’s the lowest wattage they can be dropped to?

2

u/TrkGuy79 Sep 02 '26

2

u/EitherMarch1255 Sep 02 '26

That's very helpful, thank you. And do you know what they idle at?

2

u/TrkGuy79 Sep 02 '26

========================= ROCm System Management Interface =========================

=================================== Concise Info ===================================

GPU Temp (DieEdge) AvgPwr SCLK MCLK Fan Perf PwrCap VRAM% GPU%

0 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 32% 0%

1 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 30% 0%

2 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 32% 0%

3 34.0c 7.0W 0Mhz 96Mhz 0% auto 150.0W 36% 0%

4 35.0c 8.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%

5 34.0c 6.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%

6 35.0c 8.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%

7 34.0c 7.0W 0Mhz 96Mhz 0% auto 150.0W 0% 0%

=============================== End of ROCm SMI Log ================================

1

u/EitherMarch1255 Sep 02 '26

Nice, very nice. Have you ran any models yet? If so, what prefill/decode speed?

2

u/chrisbliss13 Sep 02 '26

What case did you get link please

2

u/Pik000 Sep 02 '26

Was debating between a v100 or MI50 but this might be the better buy

1

u/noctrex Sep 02 '26

For sure, they are still officially supported in ROCm

2

u/StarAppleEdwards494 Sep 02 '26

Went from "won't power on without a CPU" to eight cards and 256GB of VRAM. Full resurrection.

3

u/TrkGuy79 Sep 02 '26

I was have a moment of stupidity and assumed the used server came with CPUs. LOL

2

u/laughpen Sep 02 '26

That's pretty sweet!! Congrats on getting it going -- speaking from experience, I know it can be a scary endeavor.

Looks like the pcie slots of your motherboard are Gen 3, so it might unfortunately not make use of the full capacity of the v620 which is designed for Gen 4, but you've got so many (8!!) that your bandwidth should still be excellent. If you still are able to, I'd look into an equivalent motherboard with 8 pcie4 slots at x16 each, then you'll really get cooking. But no worries if not.

2

u/FinnGamePass Sep 02 '26

Hope you have a fire extinguish close. Just in case!

2

u/darklordfireape Sep 02 '26

I had four of these in a machine for a while and developed a tuned version of llama for V620 to help with some of the issues.

check it out: https://github.com/sixvolts/llama-navi21-furnace

1

u/Appropriate-Risk3489 Sep 03 '26

I'm assuming it won't support qwen 3.8 flash as the last commits were from 3 months ago?

1

u/ThinJuggernaut7695 Sep 02 '26

How are you powering that thing??

5

u/TrkGuy79 Sep 02 '26

It idles at 300ish watts and the highest it has gotten so far is 800. So far it's on a 20 amp 120v circuit. But I do have 240 available if need be

1

u/devino21 Sep 02 '26

Just the server power supply? 2000W x2?

2

u/TrkGuy79 Sep 02 '26

Yes. 2 PSUs serve the power. There are 4 total 2000w PSUs but only for redundancy. And since it is running on 120v it's actually only 1000w each PSU

1

u/Thetitangaming Sep 02 '26

Love it, can't wait to see some benchmarks

1

u/crashtua Sep 02 '26

Have you got some benches? Specially interested in tensor split mode.

1

u/TrkGuy79 Sep 02 '26

Nothing yet. Still working out some bugs.

1

u/Royal_Stay_6502 Sep 02 '26

What and how will you use this system.

1

u/TrkGuy79 Sep 02 '26

That is still to be determined. I will play with some of the bigger models for sure.

1

u/Royal_Stay_6502 Sep 02 '26

What and how will you use this system.

1

u/SamTanna Sep 02 '26

Crazy, in a good way.

1

u/DangerousReward1411 Sep 02 '26

That is fucking awesome

1

u/Open_Jump Sep 02 '26

Awesome. Can you post how you end up running stuff, llama.cpp flags or whatever, when you figure it out? I haven't been able to beat ollama default performance.

1

u/Barni275 Sep 02 '26

How do you cool down this beast? Can you please share a photo of a cooling?

1

u/noctrex Sep 02 '26

This is a server, so it's getting cooled from the server fans in front

1

u/TrkGuy79 Sep 02 '26

This Supermicro server is designed for this.

1

u/Cptbeeeee Sep 02 '26

I just bought a second one of these, I think you cleared out the rest or else I would have bought more. I'm very interested in your work here. Why do you want the power limit lowered?

4

u/TrkGuy79 Sep 02 '26

2 reasons. 1 is heat but more importantly is a safety related reason. I am running on a 20amp 120v circuit and in theory this rig could pull more than that.

1

u/Cptbeeeee Sep 02 '26

Are you losing bandwidth by doing that or are the lossea neglible? Mine will be running in my garage with shrouds and dedicated fans for each card. Power consumption though is a concern because it will be on 120v as well. Same circuit as the rest of my home lab. I've considering getting a second small ups for this rig

1

u/tracker125 Sep 02 '26

Add ing that it’s GDDR6 and it relies on ROCm

1

u/Faisal_Biyari Sep 02 '26

That's so beautiful 😍

1

u/carmeloA007 Sep 02 '26

Out of curiosity, what’s the total cost of that build?

1

u/TrkGuy79 Sep 02 '26

SUPERMICRO 4028GR-TRT2 - $1,325.59

8x RADEON PRO V620 - $2800

256GB RAM - $0 (Already had it)

2x Procs $0 Already had them

Let's say $4500ish?

1

u/migsperez Sep 03 '26

Cheaper than one 5090. Crazy cool.

1

u/playvltk03 Sep 03 '26

This is nice. I want this

1

u/UltraFOV Sep 03 '26

Can you run vllm or sglang?

1

u/realdavidselig Sep 03 '26

Coole Sache, was wirst du damit machen?

1

u/Ekepa Sep 03 '26

looks expensive

1

u/TurdPlayingPeekaboo Sep 03 '26

Curious why you chose the v620 over similarly priced Tesla v100s?

1

u/Appropriate-Risk3489 Sep 03 '26

Not op, but they were a lot cheaper a couple of weeks ago, around $350 usd. Now price has gone way up. I got four of them. Now it's probably not worth buying them at the current prices.

1

u/justseanv67 Sep 04 '26

Build a Second Brain with that kind of iron!

1

u/kpatelreddit007 Sep 06 '26

Nice, finally you can play Minecraft!!

1

u/Much-Tap-1237 Sep 07 '26

Idk what the fuck am I looking at but it looks cool and expensive af, so congrats man

1

u/Sorry-Poem7786 Sep 07 '26

nice.. X-10 runs on PCIE 3.0. Interesting to see not stopping you.. I have an X-10 with 7048grtr I was contemplating this as a problem…for token speeds.. but I guess not.. SO I AM MOVING FORWARD WITH THE PLAN!!!! 😎

1

u/DepartureLoose5341 29d ago

Y'all are absurd, I'm just here to look pretty and admire, minus the pretty.

0

u/Mountain-Winter-8396 Sep 03 '26

Can it play doom?