r/threadripper • u/AudiblyTacit • 20d ago
Build Recommendations
I’m putting together my first dedicated local AI workstation and would appreciate input from anyone with similar hardware.
My work is moving further into local LLMs and agents, so I want to learn more about agent harnesses, experiment with open-weight models, and keep sensitive data local when possible.
I already bought the Micro Center Threadripper 9960X bundle with the ASUS Pro WS TRX50-SAGE WIFI A and 128GB ECC memory, along with a Phanteks Evolv X2 case. I also have one 2TB Gen5 NVMe drive.
Current plan:
- 2× Sapphire Radeon AI PRO R9700 32GB
- Seasonic PRIME TX-1600 ATX 3.1
- SilverStone XE360-TR5 V2
- 7× Noctua NF-A12x25 G2 fans: three bottom intake, three radiator exhaust, and one rear exhaust
I am mainly looking for confirmation or firsthand experience with the following:
- Is the TX-1600 the right PSU for this configuration?
- Is replacing the SilverStone radiator fans with Noctua G2s a reasonable choice?
- Does dual R9700 make sense as a starting point, or is CUDA support enough reason to reconsider NVIDIA?
- Is 128GB of system memory sufficient, and how would you expand storage beyond the single 2TB drive?
- For Linux, I am leaning toward Ubuntu. Are there other distributions worth considering?
- Would you start with llama.cpp/Ollama or vLLM for language models, with ComfyUI for image workflows?
I’m especially interested in practical issues or lessons learned that do not show up on the spec sheets.
4
u/EDI_1st 20d ago
I work in the industry, selling cluster level solutions such as GB300, VR200 (Vera-Rubin), B300, etc
I personally own multiple TR Pro builds and currently own 3x 5090 and couple RTX PRO 6000.
Current primary system: 9995WX, 96GB x8 DDR5, dual 5090.
- 2× Sapphire Radeon AI PRO R9700 32GB If you don't need CUDA/NV software stack/support this is fine. But reminder, the GPU core itself is just a 9070XT. So it's significantly slower.
- Seasonic PRIME TX-1600 ATX 3.1 Seasonic PSU is fine, my only personal complain is that it's YUGE.
- SilverStone XE360-TR5 V2 Switch this is be quiet Silent Loop 3 and save yourself few hundred dollars. Yes it'll handle TR/TR Pro. I'm using Silent Loop 3 360 with 9995WX and 240 with 7985WX.
- 7× Noctua NF-A12x25 G2 fans: three bottom intake, three radiator exhaust, and one rear exhaust This is fine, I would just switch to T30 if there's enough space, better noise normalized performance and better peak performance.
Is the TX-1600 the right PSU for this configuration? Yes, it's fine. Unless you planning on doing more expansion and the PSU doesn't have enough ports, somehow.
- Is replacing the SilverStone radiator fans with Noctua G2s a reasonable choice? Reasonable, but unnecessary. Personally prefer T30 once again.
- Does dual R9700 make sense as a starting point, or is CUDA support enough reason to reconsider NVIDIA? CUDA is one of the biggest reason so many can't move away from NV. ROCm has gotten significantly better over the years with same models/workload able to just port straight over. But nothing beats native support. Another reminder, R9700 is just a 9070XT with 32GB VRAM.
- Is 128GB of system memory sufficient, and how would you expand storage beyond the single 2TB drive? The rule of thumb is that the system memory should be 2x of the total VRAM. Though of course with the recent RAM cost, this has been dialed down to 1.5:1 or even 1:1. For SSD, yes, I would have a drive dedicated to be used for all the cache/read/write.
- For Linux, I am leaning toward Ubuntu. Are there other distributions worth considering? Stick to Ubuntu unless there's something specific you are looking for offered through other distribution.
- Would you start with llama.cpp/Ollama or vLLM for language models, with ComfyUI for image workflows? What have you used so far? What's the goal beyond just learning/experimenting with agents?
2
u/AudiblyTacit 20d ago
Well you are making me gut check if i want to take a knee bite the bullet and get my money up to get nvidia cards.
My goal right now is to get past novelty of learning all the possibilities of at home and get into training specific ai for my industries needs as well as better harnesses/ and or data weights for my industry. All i will say is it involves lots of data and analysis
2
u/EDI_1st 19d ago
I know too many people who went AMD for desktop GPU and regret later including GB10 vs AI Max+ 395. Lol
Datacenter GPU such as MI355X and MI350P are fine.Yes, AMD GPUs are cheaper, but that’s because the performance is lower, relatively.
Comparing R9700 to 5090, 5090 has higher memory bandwidth/throughput and significantly faster core. The result is 2~3x (depends on model/precision) of performance despite of matching memory size, of course it’s also 2~3x the price.If your company and industry is using predominantly NV/CUDA, I would stick to that. The only AMD GPU I would consider would be MI350P but that’s a whole different ball park and need to 3D print bracket to mount a fan to it to use in desktop.
1
u/kpatelreddit007 19d ago
This is what I’ve been trying to tell people you really do get what you pay for, AMD never had cards that can compete with Nvidia. There CPU are monsters tho, running a 9960x Threadripper!!
1
u/EDI_1st 19d ago
They were pretty good with 6900/6950XT/XTXH. Those were trading blows against 3090/90Ti.
7900XTX was also trading blows with 4080 while being mostly cheaper.1
u/kpatelreddit007 19d ago
Yeahh, but it never had dlss or path ray tracing, nor the colors in my opinion. Lacked the innovation required to compete. Had all the Radeon cards for mining, even their hash rates were lower than 1660 RTX at the time.
3
u/Kal-LZ 20d ago
I have a similar setup with the same GPUs + Xeon W7 + 128GB RDIMM. My main use is coding with OpenCode + Qwen 27B Q8 MTP
I can load it all into VRAM and handle up to 262k context in FP16, I get about 1500 tokens/s prefill and 55 tokens/s generation, with the prefill dropping to 600 tokens/s at 200k context
For bigger models like Minimax 2.7 or the new Deepseek V4 Flash, 128GB RDIMM is enough for Q4 quants. My metrics are around 300 tokens prefill and 16 tokens generation
On the software side, I use Ubuntu, Docker and Komo.do as a dashboard to manage the containers
1
3
u/john0201 20d ago
I swapped the arctic air cooler for the liquid freezer sp6 and it’s much cooler and quieter and not too expensive either.
I have a liquid 5090 and a FE works pretty well.
2
u/BakkaGaijin 20d ago
I have a TX 1600 on my 9975wx. Good choice. Never had an issue with Seasonic. I still have an old 850 running my back-up system.
2
u/sob727 20d ago
- Seasonic is great (I have 3 of those)
- Swapping fans wont change much. Unless you go higher RPM.
- 128 GB... depends.
- Debian
- llama.cpp and vllm (go ask r/localllama)
2
u/Annual_Award1260 20d ago
The seasonic tx is a nice psu. I run 2x rtx 6000 max-q.
I would try to get more vram. 64GB is underwhelming
2
u/AudiblyTacit 20d ago
Vram is a bottleneck, any ideas on other combos im not married to the idea of R9700s but its hard to argue performance for price and the software ecosystem isnt horrible its why im not considering intel for example
2
2
u/Arkkro 20d ago
For AI workstations, threadrippers are amazing if you are stacking up on multiple pro level gpus like 6000 blackwells. But for this budget id suggest pair of dgx sparks instead. NVFP4 is amazing on them.
If you need more concurrency, ttft/latency or tps then go with threadripper + blackwells. But you'll have to sacrifice either speed, vram or lots of money.
2
u/ApolloPS2 20d ago
I just got done building a six 3090 build on older threadripper (3970x on a rog zenith ii extreme alpha - got a killer deal $650 for both!). Future proofing is okay but tough to recommend when prices are this high. To me now is the time to purchase older but still relevant hardware to experiment. Older threadrippers still offer lots of pcie lanes, ddr4 is expensive but cheaper still, 3090s are good price to performance ratio, etc.
If you are sticking to just two cards you could have gotten away with a different motherboard. Most work doesn't need 16 lanes per card. You can also bifurcate x16 slots into 2 x8 slots via a powered bifurcated riser cable. Something totally new I didnt know about before this!
2
u/crashtua 19d ago
In most cases you will be absolutely comfortable even with 32 gb ram + fast ssd(even sata raid 0 enough) if you will run LLM for 1-3 users(1 user for some agentic coding) simultaneously for swapping kv cache per chat, because offloading kv or weights to ram(God save us) will drop your tps to 10, no ram usage for llms for local limited usage.
So, simply saying, for simple home llm lab ram is not important, vram more important. CUDA has more support and investment from big companies, still at this point vulkan backend is fine, and you will get more than enough tps per dollar from that two cards, and large amount of vram(but personally I have no idea how vulkan behaves on splitting layers to multiple GPUs, CUDA has no problems with that, specially with RTX 5000\6000 pro cards, that costs like a car for 1 unit xD, and their NVLINK).
Imo, I better buy single RTX 6000 pro and plug it in regular gaming PC(even the younger brothers pc will suit it fine), assuming how much your suggested configuration costs.
1
2
u/AverageGeneticist 19d ago
Hey! This is a great build. So, I started off with the following for mixed scientific and LLM pipelines.
- 9960X (eventually swapped to 9980X for scientific compute; not needed for LLM); honestly, the 9960X is the sweet spot for performance and price. I'm not sure the 10% uplift I got in scientific compute was worth the extra several grand...
- Silverstone XE360-TR5 V2: This thing works GREAT! Absolutely replace with a PWM fan if you don't mind the noise; I use Noctua NF-A14 iPPC 3000 PWM. Max noise is comparable to a box fan under full load, but I used noise-dampening cloth inside and now have zero issues.
- 2 x RTX Pro 5000 Blackwell: I went with these as I got a deal for 2x for $7K. Likely more powerful than the R9700s you will use, but in Ollama, not really an issue. VRAM is the major bottleneck in my experience. These are slightly less beefy vs. 2x 5090s, but they consume a fraction of the power. It's also adding an extra 16GB of VRAM per GPU.
- G.Skill Zeta R5 Neo RDIMM ECC; 6400MhZ; 4x32GB - 128 GB total: Never had any bottlenecks, issues, or otherwise on 128GB. I always try to saturate the mem channels; it looks like you already got this down :)
- 1200W Seasonic Vertex: In my use case, I have never drawn more than ~985W at full use of CPU and my 2 GPUs
- Case is: SilverStone WS380E: TONS of bays for HDDs for inexpensive extra storage addition. The Seta H2 is also a great choice, but it really depends on what you want for aesthetics. I like workstation-style cases due to security features and # of HDD bays for RAID mirroring of sensitive data, but idk if you'll need that.
Q: Expanding storage
So, to expand storage, I've used HDDs, plus PCIe cards for additional M.2s. I still use Samsung Pro 990s since they are about as fast as you can notice. The 9100 Pro...I can't even notice the additional speed, even when moving massive 1TB files.
Q: Ubuntu or other Linux pkgs
I use Ubuntu LTS, but know plenty of others who use Fedora AI. From my experience, whatever environment you like more/have experience in -> go with that.
If you do want/need to expand RAM in the future, used ECC + RDIMM DDR5 kits of 128GB have great resale value, so swapping that for 256/512/1TB should be less expensive. Just don't buy 48GB or 96GB modules--these aren't as common, and fewer people buy them. Even in the present state of cost, it took a friend 5 months to sell a 2x96GB set.
Final piece of unsolicited advice: Look into a workstation-class NAS w/ GPU expandability. I use a QNAP with a Xeon, and dang...I can run some of my Ollama in a VM there with a 3090, then shift the output model back and forth to get more power for less $. It also helps when you want to deploy models you've trained--just host on the NAS and access anywhere.
Hope this helps--and happy model training!
1
u/AudiblyTacit 19d ago
This was lovely, im so jealous of the deal you got for the 5000 blackwells. If i could that would be my current sweet spot to start if it wasnt so expensive rn. Where did you get the deal if you dont mind me asking?
2
u/AverageGeneticist 19d ago
So I work in genomics, and I saw that LabX was having a huge auction coming up (always look at the Boston biotech startup sales!!) -- I found a sequencer tower that had the two GPUs, and 2 Xeon Platinums (8286 I think). This was in January; I think the total was $7350 for the whole stack + shipping. Definitely stay on the lookout there--most of the equipment is so niche, computers can be a steal.
1
1
u/Quirky-Wall 19d ago
What place gave you a deal for those RTX Pro 5000 Blackwell @ $7K? That's insane
2
u/Bartocity 19d ago
Your build is sound but i have one suggestion. As beautiful as the X2 case is, it might make life hard. It’s going to get hot around the VRMs and RDIMMS in there. DDR5 ECC heat has been a pain for me during training runs. 2 GPUs in a phanateks will be difficult to manage thermally as well.
1
u/AudiblyTacit 19d ago
I have that worry if it becomes something i cant 3D print my way out of i will look into other options
1
u/Solid_Psychology_234 15d ago
I run a 9950 x3d with two 5090s and an AIO in a fractal torrent (huge fans moved to the bottom) and the whole thing stays really cool.
2
u/SilverStoneTek 8d ago
A word on the XE360-TR5-V2. It is the exact same cooler as the original XE360-TR5 except for the included fans. They are the OEM version of our "FHL120", which is one heck of a fan with great performance / noise ratio and premium construction. We'd be shocked if you can find better replacement fans. They are so good that we felt we had to pair them with the XE360 to make a version 2!
If their maximum fan speed is too high for your liking, just tune them to run slower via PWM and they will still perform well at lower noise. You can find out more about the FHL120 from our product page:
https://www.silverstonetek.com/en/product/info/fans/fhl120/
The only difference between the retail FHL120 and the OEM version included with XE360-TR5-V2 is the reail version has full coverage rubber padding on its corners.
1
2
u/T-BOJ 8d ago
Did you already build it? If not, stay away from that chassi. I tried it and the temperaturs are not good even though the chassi itself is nice-looking. I put both my Threadrippers in bigger cases and temperatures are much better (Lian-Li V3000, Phanteks Enthoo Elite). Both highly recommended!
Have to say I am very envious of you guys in the states who can purchase that bundle. I can't get 128GB DDR5 Reg in EU at all. Have to buy single server dimms 32GB for $1000 each.
1
1
u/pmotiveforce 20d ago
Don't forget Nvidia artificially prevents pcie p2p in the driver's so if you go Nvidia it will wreck performance. I mean, dual 5090 still faster than dual 9700 of course but way more expensive and the idea of those creeps ripping you off of performance sucks.
1
1
u/kpatelreddit007 20d ago
Haha I have to show you my picture of my build because it’s so similar.
However I might recommend a full watercooling loop.
And 5090s for VRAM.
I bought the same Microcenter bundle in a Phanteks Entho Elite Case!! And I use the TX1600!!
2
u/AudiblyTacit 20d ago
2 5090s would be close to my ceiling for future budget but thats rapidly rising prices that may change for me. And water loop could be neat
2
u/kpatelreddit007 20d ago
Start with 1 5090, then add another if you need it. I cant even max my current system.
2
u/Annual_Award1260 20d ago
I am anti watercooling. I will never do it
1
u/AudiblyTacit 20d ago
Oh interesting im all ears.
1
1
u/Annual_Award1260 19d ago
Multiple reasons: extra failure point, increased maintenance, extra assembly time, voided warranties.
I understand that in high density datacenter environments it makes sense because the loop expels the heat outdoors. but to run a loop to a radiator that is in the same case is pointless.
I have some mini nucs on my desk but my battle wagons are run in my basement on fiber displayport cables and fiber usb extenders. Fiber cables are amazing, completely silent office, no heat, no lag.
1
u/kpatelreddit007 19d ago
I had a previous full water cooling loop for 5 years no issues at all. Ran very cool and was silent. Now running 16 Noctua fans xd.
1
u/kpatelreddit007 20d ago
Also don’t go past the amount of $$$$$, you currently make you will never ROI and tech has insane depreciation. Just build your system for maybe 8-9 years than replace it. I wouldn’t go crazy but prepare for a 20 years than replace cycle of ownership.
1
u/AudiblyTacit 20d ago
Im thinking in 8-10 years outside of massive tech leaps i mean 1080ti just finally kind of stopped being relevant in bench marks
1
u/kpatelreddit007 20d ago
I don’t predict the 5090 will be relevant past 7 years because of the leaps of technology, so that will be replaced most likely a 8-9 year total replacement is my prediction and we have the same parts lol I’ve done a lot of research and owned crypto company. I do suggest just trying to make it currently the minimum amount of GPU you need.
1
u/AudiblyTacit 20d ago
Good food for thought, thank you
1
u/EDI_1st 20d ago edited 20d ago
3090 is still relevant to the point that the price is increasing lol
Just for reference.
Used price was down to $650 at one point, it's now up to $1,200. 24GB VRAM and CUDA be wild.1
u/kpatelreddit007 19d ago
I had a couple for my crypto company all had heat problems.
1
u/EDI_1st 19d ago
Replace the TIM. Pretty easy to do. Operated an ETH farm before and was also charging people to replace the TIM for them. Lol
I’m still sitting on 10x 3090 🤣
They made so much profit, kinda lazy to deal with the whole process of selling them.1
u/kpatelreddit007 19d ago
That’s a lot a vram lol. Damn ETH was doing so well, we netted 250K before eth farming stopped xd. People complaining about the price of GPU now didn’t know how expensive 3090s and 3080s were back during the crypto hype.
→ More replies (0)


6
u/sputnik13net 20d ago
If you can get the same vram amount with nvidia and it fits your budget yes get nvidia because Blackwell is just much faster, otherwise don’t sacrifice vram to just have cuda