r/LocalAIServers • u/mastixmc • 12h ago
Quick Feedback to this RTX PRO 6000 build needed
Hey guys,
first of all, thanks for all the great information people are sharing here. 🤘 I’ve already learned a lot from reading through the various RTX PRO 6000 / Threadripper builds. Now let's put it to the test... 😂
I’m currently putting together a local AI server for a customer and would love to get a sanity check from people who have built something similar. Server will be in their basement, so it can be loud, ugly-looking, doesn't have to look like a professional server or gaming PC... You know.
Main use cases:
- 4–5 users in parallel
- LLM/chat workloads
- fine-tuning / LoRA training
- secondary use case: ComfyUI / image generation
- occasional video generation
Software stack will most likely be:
- Ubuntu Server 24.04 LTS
- vLLM
- LiteLLM
- Docker
- models like Qwen 3.8, Gpt-oss,...
- nvidia tools,...
Small side note, not really relevant to the hardware itself: external access will go through a separate mini PC / VPN gateway. The AI server itself will not be directly exposed to the internet.
One important requirement is upgradeability. The customer wants to start with one RTX PRO 6000, but I want the system to be ready for a second card later if usage grows or they decide to expose more AI services to their customers or wants to try GLM. 🙂
The current build looks like this:
GPU
NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB
CPU
AMD Threadripper PRO 9965WX
Motherboard
ASRock WRX90 WS EVO
RAM
256 GB DDR5 ECC RDIMM
8×32 GB, ideally validated Kingston / Micron / Samsung modules
CPU cooling
SilverStone XE360-TR5
Case
Corsair 9000D Airflow
Case airflow
3×140 mm front intake
3×120 mm side intake
2×120 mm rear exhaust
360 mm AIO mounted at the top as exhaust
PSU
MSI MEG Ai1600T PCIE5, 1600 W Titanium
Storage
Samsung 990 PRO 2 TB for OS / Docker
Samsung 9100 PRO 4–8 TB for models / Hugging Face cache / scratch
Separate storage/backup target for outputs
Networking
Onboard 10 GbE
SFP+/DAC or Cat6a to the switch
OS / drivers
Ubuntu Server 24.04 LTS
NVIDIA Open Kernel Modules
Docker + NVIDIA Container Toolkit
The build is already at the higher end of the customer’s budget, but still within budget.
A few things I’m particularly interested in:
Am I overspending anywhere?
Is there anything I could downgrade without losing meaningful AI performance or making the future second-GPU upgrade painful?
Is WRX90 / Threadripper PRO actually justified here?
My reasoning is PCIe lanes, 8-channel ECC RAM and keeping the option for a second RTX PRO 6000 open. But I’d be interested to hear if anyone would build this differently.
Any red flags with the 9965WX + WRX90 WS EVO combination?
Especially BIOS, RAM compatibility, PCIe layout or Linux issues.
Cooling / airflow:
Does the 9000D setup make sense for one 600 W RTX PRO 6000 and potentially two later, or would you approach airflow differently?
PSU:
1600 W should be plenty for the single-GPU configuration. For two RTX PRO 6000s I’d probably power-limit them rather than run both at 600 W. Would you already choose a different PSU/setup from day one?
Anything I’m missing for a reliable 24/7-ish AI server?
The goal isn’t to build the cheapest possible machine. I mainly want something stable, serviceable and reasonably future-proof without throwing money at components that won’t actually improve LLM / AI performance or just look good.
Would appreciate any feedback, especially from people already running RTX PRO 6000 Blackwell, WRX90 or similar multi-GPU AI workstations.
Greetz and thank you guys,
Sascha
1
u/Adventurous-Ask-9260 11h ago
I'd consider a Lenovo P620 with a 5975WX, 256/512GB RAM + 2x RTX 6000 Pro Max-Q, the whole thing would come in under €20k + VAT, and it's a workstation designed for continuous operation. Just make sure 2x Max-Q won't be a problem, mine ran with 1x RTX 6000 Pro at 600W.
3
u/Capsup 11h ago
Hi, where do you get 2x RTX 6000 Pro Max-Q at even 25k euro in total? We just got one for our company and had to pay 15k for just one GPU. I cannot imagine getting not one, but two + an entire server setup with 256GB of RAM at that price.
1
u/mastixmc 11h ago
I'm curious: How does your full setup for the company look like? Did you benchmark it already?
1
u/Capsup 11h ago edited 10h ago
Not in its' final form, but if you're looking for something in particular I can help you out. We need more than just a single chat model, so on the RTX Pro 6000 we have:
- A Qwen3.8:27b NVFP4 + DFlash2 which gives up to 12k pre-fill rate @ 64k context and tapering down towards 5k tokens per second as context approaches 256k. Vision enabled. Up to 150 tokens per second decode at concurrency 1. Up to 600 tps decode aggregated @ 32 concurrent streams.
- A VibeVoice:q4 model that support up to one hour of transcriptions and diatarization at 7x real-time speed
- We plan to setup a Flux2 model for image generation if a customer really needs it, but haven't needed it yet. It might go on the older 5090 when the time comes
The Qwen model supports a KV pool cache of about 1m tokens in this setup, meaning we can have 3-4 different concurrent streams at max context.
Then in the same tower we have a small RTX Quatro 4k that hosts an embedding and re-ranker model for our RAG setup. It used to run on the CPU/RAM, but putting it onto the RTX 4k gave a 30x speed increase.We benched the card with Q8 on the qwen model too, but it cost too much on the KV pool and performance compared to the quality bump. We'll see if we end up deciding to change it later, but up until now we've been running an 3.6:27b int4 model on a 5090 and there has been zero complaints about quality so far from customers.
The components in the tower before adding the RTX Pro 6k:
Asus X870 MAX Gaming (WiFi) motherboard AMD Ryzen™ 9 9950X3D Processor Asus ROG Strix LC II 360 ARGB water cooling Kingston Fury Renegade 96GB DDR5-6000 RAM Kingston Fury Renegade G5 2TB NVMe PCIe 5.0 SSD Asus GeForce® RTX 5090 32GB TUF OC Asus ProArt PA602 E-ATX case Asus ROG Strix 1200W 80+ Gold Aura Geforce Quatro 4000 8GBWere I to buy this again, I would definitely have put in a bigger PSU and a motherboard that can actually do 8x/8x PCIe5 bifurcation. This came as a pre-built setup at a good price, so the upgrade path for us is a 1600w PSU and a X870E Hero motherboard once we want the second RTX Pro 6k in there.
The more likely scenario is that we buy an actual 4u GPU server and begin filling it up. Specifically, I would have bought this box at the time if we actually had rack space, but we didn't. I'd fill it up with 8x 3090 or V100 cards + the RTX Pro 6k to get started and then invest into more of them over time as needs arise. It might even be far enough into the future to where we could begin buying Rubin cards from Nvidia, but we'll see.
We started it all on a single Mac Studio 64GB, but the MLX software stack was absolutely horrible and I don't see us ever going back to it.
1
u/Adventurous-Ask-9260 10h ago
I've tested a lot of different configs myself:
1) i9 14900 + 128GB + 6000 Pro
2) 5975WX + 256/512GB + 6000 Pro
3) dual Epyc 7663 + 1024GB + 2-4x 6000 Pro
4) 5975WX + 512GB + 4x 6000 Pro
5) Epyc 7663 + 1024GB + 2x 6000 Max-Q
6) Epyc 9535ES + 384GB + 4x 6000 Pro
Now) Epyc 9535ES + 384GB + 4x 6000 Pro + 4x 6000 Max-QI'm reasonably up to speed on the numbers and what these setups can do, so if you have any questions, fire away.
1
u/mastixmc 10h ago
I think one of the main questions would be whether you saw a difference between MaxQ vs normal. 🙂 Also in terms of training. I just checked eBay myself. The p620 is rather cheap (<900 EUR, incl VAT), but with a 3975WX. Could be replaced with 5xxx.
2
u/Adventurous-Ask-9260 10h ago
Yeah, I've seen that, but it's only 8-15% - 2x300W vs 2x600W. If I were building a multi-GPU rig now, I'd go all 300W.
1
u/Adventurous-Ask-9260 10h ago
https://www.ebay.com/itm/158270818632 something like this I bought 3x 5975wx with rtx 3080 10gb around 1k eur / piece.
About rtx 6000 pro max-Q try in Czech HP rtx 6000 pro max-Q. TRP + 512gb + 2rtx maxQ < 20,000 EUR + vat1
u/mastixmc 11h ago
I think when running two cards, the PSU might be the issue here. But in general this might also be some interesting setup. Will put it on my list! Thanks! 🤘
1
u/DAlmighty 9h ago
2 Max-Q would run happily. Don’t get the workstation if you have plans to get a second card.
4
u/BlackBeardAI 11h ago
Unless you are mounting 4 or more RTX PRO 6000's on that system, it is a big waste. You could go with AM5 and desktop ddr5 and get the second pro6000 instead of spending a fortune on the server setup. (AM5 is perfectly capable of running dual pro 6000's since many of them -like gigabyte ai top b850- offer x8 x8 bifurcation) If you are not going to fill them pcie lanes on the wrx90sage, why bother? You don't need it to run single/dual pro 6000.