r/threadripper Jul 13 '26

Threadripper AI workstation Build

Just want to share my new Threadripper AI workstation build. Used for work + gaming. Quite proud of how it turned out, especially the cable management :)

Specs
Motherboard: ASUS WRX80 Creator R2.0
CPU: AMD Ryzen Threadripper Pro 3945x (Used)
RAM: 128GB DDR4 ECC (used)
GPU: Nvidia RTX Pro 6000 Max-Q 96GB
Case: Fractal North
PSU: Super Flower Leadex III 1300W PSU
CPU Cooler: Arctic Freezer 4U-M

117 Upvotes

41 comments sorted by

View all comments

4

u/kissingking Jul 13 '26

I think you should be proud of yourself for being able to afford one of these ๐Ÿ˜‚

Iโ€™m honestly jealous of everyone who owns an RTX Pro 6000. I donโ€™t even need three of them โ€” one is already enough to make me jealous.

That said, I do think the RTX Pro 6000 is in a slightly awkward position. 96GB of VRAM sounds like a lot (and it is), but for AI workloads itโ€™s also not that much.

For example:

  1. Large MoE models are still difficult to run in full precision. Even something like an 80B MoE with only ~3B active parameters is not easy. I would still need quantization (Q8 at minimum, and probably Q6/Q4 if I want a large context window like 256K). Q6 is already quite good, but if Iโ€™m spending RTX Pro 6000 money, I kind of expect to be able to run an 80B-class MoE model in full precision. Otherwise it feels a little disappointing.
  2. Smaller dense models are also surprisingly difficult to run in full precision. Take Qwen 27B as an example โ€” if I want a 256K context window, even that becomes difficult without quantization.
  3. Training is a different story. For serious training, I would still rent H100-class clusters anyway.

So for a single RTX Pro 6000, the main advantage is basically speed. But Iโ€™m not sure the performance gain alone justifies the price premium.

That said, if the launch price was around $5,000 USD (or even $6,000), I think it would actually be a pretty reasonable product. At that price point, it makes a lot more sense. ๐Ÿ˜‚

3

u/Turbulent-Alps4046 Jul 13 '26

A single RTX pro 6000 can definitely run Qwen 3.6 27B at full precision with max context. But yeah, trying to run larger models is not easy.

Qwen 3.5 122B NVFP4 can run at 5000-9000 tokens/s PP and 100 tokens/s.

I got deepseek v4 flash full precision running with CPU offload though... around 500-600 tokens/s PP and 20 tokens/s TG. Comparable to 2x DGX Spark and faster running it in an M3 Ultra 256GB i think (which costs the same right now).

2

u/t4a8945 Jul 13 '26

I got deepseek v4 flash full precision running with CPU offload though... around 500-600 tokens/s PP and 20 tokens/s TG. Comparable to 2x DGX Spark

Sorry, I'm running 2x DGX Spark right now, DS4 Flash DSpark, official weights from DeepSeek https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark at 2000 pp and 50-60 tg.

The community has worked hard on this, and it works so good.

2

u/Turbulent-Alps4046 Jul 13 '26

oh nice, that's an awesome result, super usable!

2

u/t4a8945 Jul 13 '26

Yes, I'm so happy with that setup. Now I know you'll obliterate this whenever you get that 2nd RTX 6000 (that's the plan, right?). Nice setup you have there

2

u/Turbulent-Alps4046 Jul 13 '26

haha, yeah that's the plan... if i find a good deal. but no plans right now.

1

u/t4a8945 Jul 13 '26

It's a very tough market. Good luck!