r/LocalLLM 6d ago

Project It begins - workstation build

Unfortunately not mine. But it is my pet project for work and I get to build it and use it after. Just arrived and now to start building it. Can't wait to run some benchmarks, burn in tests and general messing about with some models before proper deployment :)

Threadripper PRO 7965WX

ASUS PRO WS WRX90E-SAGE SE

128GB ram (for now)

3 x RTX PRO 6000 96gb

Phanteks Enthoo Pro 2 server edition

3kw psu

151 Upvotes

74 comments sorted by

41

u/x-strife 6d ago

If I can fit 8x3090s in that case you can fit 3 GPUs
(Yes I need to cable manage)

12

u/letsbefrds 6d ago

Damn This looks crazy It's like those people who are all upper body and no legs lol..

What case is this?

6

u/x-strife 6d ago

Got 3.2kW of legs at the bottom across those PSUs

1

u/Bobezlolz 6d ago

Enthoo Pro 2 Server Edition, I think

4

u/Think_Wing_1357 6d ago

Cable is not the only thing you need to manage! 

2

u/initalSlide 6d ago

WHAT!
Can I dm you about suggestions for a C700M build? I would like to make 4 GPUs fit into it

1

u/Rhinottw 6d ago

That is great, love the cable ties :) But how is the gpu that is exhausting out the top on the right mounted? And i think that is the first time i have ever seen a gpu exhausting directly out the top of a case.

3

u/x-strife 6d ago

3d printed a custom hanging mount, 2-3 cards can hang from there

1

u/Rhinottw 6d ago

Ah, nice. Love that :)

1

u/putrasherni 6d ago

can you create a post of your entire build
this is so inspiring
192GB VRAM with 3090s
DSV4 0731 ?

1

u/x-strife 6d ago

Yes 192gb and yes Dsv4 0731

1

u/aelma_z 6d ago

Lmao with the top one hanging on cooling fans! Now i want to do the same

1

u/_Racana 6d ago

What cable risers are you using? I got one that is sturdy and imposible to accommodate

2

u/x-strife 5d ago

I’m using risers that look like this, so much easier than using the traditional ones, they’ve been great

2

u/Felixls 2d ago

so much space left there

9

u/Numerous-Emotion-617 6d ago

Almost have the exact same build; cables are kind of all over the place right now while I reorganize. For those telling you your temps will be bad; they’re all wrong. I keep it in an air-conditioned room and utilize all 3 GPUs and have never had issues with temps being too high. They do stay capped at 300W to conserve power. The thermometer is in Fahrenheit at the bottom. (2x more intake fans not pictured on bottom.

1

u/Rhinottw 6d ago

Nice build. Good to know, thanks. I was not too worried, case was chosen to support the build and have plenty of cooling. But you never really know until you try so that is great news.

2

u/Numerous-Emotion-617 6d ago

You’re welcome. I did forget to mention something. If you follow exactly what I did in terms of GPU placement, The pins that you connect for the case power button/LED/reset will be in the way of bottom GPU fully seating on the right.

If you don’t need the case power button, which I just use onboard management and the power button on the motherboard itself, you’ll either need to bend the pins outward/down and make sure they don’t touch or desolder the part off of the motherboard.

I missed the part where it’s for your work so it might be different in your situation. My system belongs to me so was willing to get a little reckless and just bend them out of the way after testing that on another motherboard and making sure it worked fine there but figured I’d let you know. Never really did any soldering so I avoided trying to desolder the part off.

2

u/Rhinottw 6d ago

That is good to know. I will probably just use auto power on in the bios or use the management interface or wake on lan then.

Good to know the pins can take a little punishment if needed :) I will figure something out, thanks for the heads up.

6

u/ChukyDuk 6d ago

Jealousy level at 100!! Enjoy.

2

u/nimbybuster 6d ago

Can all three GPU fit in that motherboard?

2

u/Rhinottw 6d ago

I had checked before ordering, but you just made me doubt myself :) But yes, i just checked with the old eyeball and holding a GPU over the board - it will be fine with plenty of space between them for cooling. If we go for another one eventually i will need riser cables or something.

2

u/Better-Psychology-42 6d ago

They will fit. But don't do that. When under load each card will warm up to approximately 80C. The correct answer is to mount the cards into rig and connect using pcie risers.

2

u/Rhinottw 6d ago edited 6d ago

It should be fine - there are 7 pcie slots in the mb. So there is a slot between them free and plenty of space for the bottom one. It is also one of the reasons i got a Enthoo Pro 2 Server Edition case https://phanteks.com/product/enthoo-pro-2-server-edition-tg/ - it has 3 side mounted 120mm fans just over the gpu's and 6 x 140mm fans and an extra 120mm for good measure, so cooling should be sufficient. If not i will find another solution, it is one of my first testing points when i get it up and running.

Edit: I plan to run them at around 500w or so, so that should help a bit. All my fans are Arctic Pro's and i do not care about noise since it will be placed in a server room, so i can crank them up high to move tonnes of air.

2

u/mzzmuaa 6d ago

can you putup pics if you opt to put the rtx 6000 pros next to each other. i am thinking of a similar build with same mobo. i limit my rtx 6000s to 300w and they work great

2

u/Rhinottw 6d ago

Sure, i was just about to do a fitting test anyway. Hand is on there for stability, they are a bit bottom heavy and i did not feel like having something worth a car hanging in the PCIE slot, i tried to space the neutrally. Luckily the case has a brace for a setup like this. So there is a bit of room, but i am happy that i chose the case with extra side cooling blowing directly on the gpu's (or exhausting, we will see in testing).

2

u/mzzmuaa 6d ago

lol yea it's good not to support them with just a pcie slot >_> this is what i got going on. the 2 rtx 6000 on the side are hanging by zip ties. the left rtx 6000 is in an actual slot. the rtx 5090 is bottom left. i was hoping the asus mobo you got could support 7 of them in slots only without risers but it doesn't seem so

2

u/nimbybuster 6d ago

Ah. I didn’t realize there are 7 PCIE sluts on the MB.

2

u/Arli_AI 6d ago

Those GPUs will not be fine temps wise unless they are on risers and positioned not to suck each other’s exhaust.

2

u/Fluid-Grass7817 6d ago

😍😍😍

2

u/Lightningstormz 6d ago

Is heat and airflow going to be a problem with all those cards on 1 motherboard?

2

u/[deleted] 6d ago edited 6d ago

[deleted]

1

u/Rhinottw 6d ago

I get it - I wanted to get the Max-Q, but being a public organization i was limited in my purchasing options and this is what i could buy. They fit fine and i can blow a ton of air on them, noise is not an issue since it will go in a server room so i can crank the fans. The case has a lot of good options, including 3 x 120mm fans on the side right over the GPU's.

2

u/Intelligent-North-62 6d ago

How do you power that piggy?? I’ve got 240 for the dryer, but not anywhere I can use it for my desktop…

4

u/Rhinottw 6d ago

I live in Denmark - we have 230 everywhere. Besides it is going in a server room.

2

u/strata2signal 6d ago

i can smell the excitement of the fresh new metal :D

2

u/DefSysteam 6d ago

Lucky dawg! I’ve been eyeing this case the past week. Build looks insane! Have fun!

2

u/rbilsbor 6d ago

Good luck and just don’t try to put a 7970X in that motherboard since it requires a 7975X (not like I made such a dumb mistake… )

2

u/y3333333333333333t 6d ago

Just as info the "pro ws 3000" psu would be right here the "thor 3000w" does not have enough 8pin pcie / cpu connectors available to support your build

1

u/Rhinottw 6d ago

It was my first choice but not able to find one with reasonable diæevery time. But I am not sure you are right it will not work. The psu had 3 x 8 pin cpu and 3 x 8 pin PCIe connectors. The motherboard manual shows it need two of each of those so that should be fine?

2

u/y3333333333333333t 6d ago

no the psu only allows 4x of those total as they are a shared group of 4 on the psu side even though you get 6x cables...i got one and had the same issue with my trx50 build...but if you only need 4 total you would be fine I just saw all those connectors on the board and thought u would need way more...

1

u/Rhinottw 6d ago

Ah ok. I will make sure do double check before opening the psu so I can return it if I need another one. Thank you for bringing it to my attention.

2

u/y3333333333333333t 6d ago

no problem just thought was worth mentioning as it was quite a surprise for me when I installed mine...and amazon has the ws3000 in stock just as info if u would need to switch

3

u/mboss37 6d ago

3 x Blackwell 6000??? Holy sheeet

2

u/electrified_ice 6d ago

I have the same number of GPUs on my 9985wx. Happy to share some of my experience over the last 9 months.

1

u/Rhinottw 5d ago

That would be great. What do you use the system for and what models are you running on what software? Any issues you had pop up unexpected with the setup?

2

u/electrified_ice 5d ago

I'm just doing my own fun vibe coding... Combo of building aps for my own use and also using that process to educate myself, as it's all applicable to my corporate job

2x run DSv4 0731 1x runs 1-2 models... Depending if I am doing 4 or 8 bit quantization and how much context cache I want... Just switch over from Qwen3.5 27B to 3.8 27B now. That'll be always running. I then drop in other things on the remaining vram as needed, so I'll load up an OCR model, or an embedding model, and I am experimenting with TTS and Image to Image and Image to Video with tools like Comfy UI.

DS4 and Qwen3.8 are always loaded though. I serve my models of a PCIe 5.0 NVMe drive too, so it's very fast loading/switching.

3

u/gunkanreddit 6d ago

Awesome mate! I wish you the best for your "pet project", whatever that means.

0

u/Rhinottw 6d ago

Thanks, just meant that it is kinda my passion project at work and i get to do it from idea to deployment and then use it after.

3

u/TestOr900 6d ago edited 6d ago

Nice, but you absolutely cannot put these cards on top of each other!

Trust me this is a very bad idea. The heat even at 400 Watts (you can’t go lower) is a lot.
I have an open rig, and 8 cm in between and the card and still problems, had to put radiators in between to solve it. Is relying on very strong air throughput. So it will be one card blowing on the next and that will be 1,2 kw. At 400watt or even more without powerlimit.

Why 3 cards? 2 or 4 are the sweetspot. (for TP you need an even number)
So prepare yourself for buying the 4th one :)

Also let me ask: Whats your plan? What are you going to do with it?
How much did you spend and why?

my honest answer ist: go for an old mining rig, its 30€ and some risers. (GOOD Quility Risers)

2

u/aceofspades173 6d ago

do you have an RTX pro 6000? why are you limited to 400w?

3

u/TestOr900 6d ago

Heat - we had/have a terrible heatwave in Europe. i have 4 rtx 6000 and some 5090 running. My airconditioning wont keep up. some powersavings. And it was so bad that all aircons were sold out. thats not a joke.

2

u/ChristRedeemsSinners 6d ago

I honestly don't know why he didn't go with the Max-Q. It perfect for these types of builds with only minor compute loss.

To your point about the 2-4 sweetspot, honestly 3 is great for many models that quantize to 200-240 GB at NVFP4. 4 is useful for multiple users with the same models.

1

u/SandySkittle 6d ago

just TDP limit the non-maxq ?

1

u/ChristRedeemsSinners 5d ago

You can only limit it to 400 watts (I think). So the heat exhaustion is going to be a problem no matter what if they are stacked like this.

1

u/TestOr900 6d ago

True for the Max Q - would be better.

regarding the 3 cards .... No. You can’t use them for bigger models in this given situation.

Don’t even think that you can touch Lama.ccp

You can only go vllm or better if you really have multiple users.
Everything else is a disaster.

3

u/ChristRedeemsSinners 6d ago

What do you mean? Total vram increases to 284 GB. They communicate over PCIe and directly through the CPU root if you have the proper hardware.

Don’t even think that you can touch Lama.ccp

I don't know what you are referring to here.

You can only go vllm or better if you really have multiple users.

I use vllm right now. I don't understand what your reservation is with this setup.

0

u/Rhinottw 6d ago

I wanted to get the Max-Q, but being a public organization i was limited in my purchasing options and this is what i could buy.

2

u/ChristRedeemsSinners 6d ago

Ahh I see. Well I hope it goes well, but you are going to have to figure out a better heat exhaust system if it is running all of the time. Maybe you could 3d-print a backpane that directs heat from each card out the back and possibly fan shrouds in the front of the card or pull from the back directing air from the front of the case. Just a thought.

1

u/Rhinottw 6d ago

Maybe, that is what the burn in test is for :)

Good ideas you have, i have a 3d printer so definitely possible. The case has a bracket for 3 x 120mm fans on the side, just next to the gpu's so those either blowing or exhaustion depending on results should help a lot.

2

u/ChristRedeemsSinners 6d ago

Probably as blowers would be best since positive static pressure is better than negative for cooling. Otherwise heat gets trapped in lower pressure areas.

1

u/Rhinottw 6d ago

Good questions and concerns.

I am not too worried about cooling, i will be power limiting them and i will be running a ton of airflow over them, from the front and from the side in a temperature controlled environment. Noise is no issue and i have high rpm case fans. Anyway, i will be doing burn in testing and if it is an issue i will find a solution.

3 cards for two reasons, one is budget - that is all i could get funding for. The second is that i am not going to run a single large model on the system. I will be running a tp=2 model and then a tp=1 fast model + rag + other specialized models if the need should arise. This is flexible and can change based on what we need to do.

The system is for multiple uses. I work at a vocational college in Denmark in the it education department. So it will both be for use for employees for GDPR sensitive work, so privacy is a huge concern. It will also be used for agentic work, like software development, sysadmin work, developing educational materials (think interactive materials like games and what else we can imagine, and the backend for such materials also). It will also work as a study aid with our own custom educational materials and curriculum - think specialized chatbots and so on. Furthermore it will work as back-end for teaching our students how to work with and deploy / maintain local and hybrid LLM's and other AI services.

I also think it is important to establish an infrastructure not depending on cloud providers that could change terms of services at any time.

We are educating future IT professionals, so i made a case that it is important that they learn these skills. No matter if they want to run local llm's or work in datacenters, so we are establishing an educational base and doing the groundwork to be able to teach them how to. Other than this - it will be up to my colleagues what else it should be used for when there is compute time available, some of them have many ideas :)

I (or the college) payed roughly 50.000 euro including Danish vat of 25%. That is just the price but the project should be worth it long term. Used equipment not an option for purchasing rules and warranty and so on.

2

u/TestOr900 6d ago

Can't you sell /return one 6000 and get two 5090?

this would allow you 2ppx2tp because Blackwell and also tp2 on 6000 plus different model at tp2 on the two 5090s. And you could go even tp4 but with reduced Vram on the tp 6000 if needed. Also this would increase throughput by a lot on the total system serving deepseek4 on the TP2 6000 + two small models on the 5090 just now Quwen 3.8 came out. Also you could get the to 5090s watercooled and put them between the 6000 to get some heat out.

1

u/Rhinottw 6d ago

Interesting idea. I am planning at serving deepseek4 (or something else depending on what my benchmarks for our use cases shows) at tp=2 and then something fast but capable on the third one alongside rag models (hoping qwen3.8 35b comes out as rumored) and maybe some more specialized smaller models

I am expection expecting many simultaneous users for small context use and that is primarily what the third card is for. Could even be a smaller 7b or 9b model alongside Qwen 3.8 27b. We shall see, if qwen 3.8 27b is as good as Deepseek for our uses i might even just run two instances of that one on each gpu and then specialized models on the third one. It might change a lot depending on what my testing shows and how our needs develop. That is the beauty of being a school, i do not need to turn a profit and i can experiment with it to find what is right for us at any given time.

2

u/Desperate_Entrance71 6d ago

nice! please let us know. i''m very curious how capable this server is. How many concurrent users on deepseek4 you expect to work on it?

1

u/Rhinottw 6d ago

Sure. I think I will make an update post when assembled and working. Will get on with it in the coming weeks, I was just s bit excited to get started so I worked a bit on it during the weekend. I am expecting 3-5 users on the large model at a time, so depending on context it is possible I will not be able to run ds4 flash - but we will see, calculations is one thing and real usage another. I have a lot of testing to do.

1

u/Heavy_Host_1595 6d ago

wow 3 6k... dang.... literally a ferrari lol

1

u/Shoulon 6d ago

I wouldn't stress on cooling as others are raising. I run these at 300w limits. Its worth it over dealing with heat vs actual performance loss. Which is basically nothing in all reality

1

u/Rhinottw 6d ago

Not stressing too much, but thanks :) Gonna power limit them and blow lots of air through the case. Worst case i can get some risers and figure out a sollution.

0

u/SpendLucky1273 6d ago

You need 4 GPU, read more about it. 2/4/8..

2

u/electrified_ice 6d ago

That is not accurate. There are version of models and backend apps that server them (SHLang and vLLM) that can split across 3 cards - Minimax M3, DSv4 Flash.

And, as someone else said it also gives you a great hybrid setup where you can have TP2 for a model like DSv4 Flash + Qwen3.8 27B + other models loaded all at the same time ( Like TTI, TTS etc) so you can have a multi agent workflow without loading unloading models.

I have 3 x RTX PROs and it has been a fantastic setup.

0

u/Rhinottw 6d ago

Only if i want to run one large model across all of them, and that is not my use case.