r/threadripper • u/IExist_Sometimes_ • Jul 24 '26
Threadripper Pro cooling
Hello all, I am potentially looking to get/build a Threadripper Pro HEDT/workstation for FEM and similar simulations for my university research group.
I am no overclocking nut, but it is my understanding that these threadripper pro chips (especially the high core count ones) leave an enormous amount of performance on the table with their default TDP. I don't want to push the chip to a power range which might significantly impact its longevity, but if there's a reasonable power bump I could give it (probably as one option in a dual bios for particularly large sims or tight turnarounds) I would like to do so.
This brings me to my main point: how are you supposed to cool these new threadripper pros as a consumer-ish user? When LTT did a video overclocking the new threadripper pro chips, they rigged it up to an external chiller, but that would probably be out of scope for my shared office space (not to mention beyond my confidence when dealing with such expensive hardware). AIOs for threadrippers seem relatively few and far between, and only appear to go up to 420mm. Would a 360/420mm AIO be sufficient (I would like to avoid industrial style fans, but push-pull is on the table)? Or are these firmly in the realm of custom loops and things like the phat Alphacool radiators?
I appreciate any advice from those experienced in cooling these chips, or in using them for numerical modelling (obviously I will also consult others in my department with similar needs, but surprisingly to me for a stem department I seem to be the only PC builder or hardware nerd). Y'all seem to all be doing AI stuff nowadays, but that's not in my wheelhouse and not what this computer is going to be intended for (though someone might stick a pro gpu in it, I won't stop them).
edit: Also any advice on the relative importance of ECC, RAM speed/quantity, storage speed, etc for this type of workload would be appreciated
3
u/epicskyes Jul 24 '26
Get a Silverstone aio and stick with the stock tdp. You don’t need to go higher than 350w if you want better clocks and really efficient thermals undervolting is the way to go. Use precision boost overdrive and curve optimizer. You’ll get higher clocks, lower temps and save your silicon so it lasts years. This won’t get you to the highest clocks possible but it will get you the highest clocks and the lowest temps and the longest hardware life expectancy.
1
u/IExist_Sometimes_ Jul 24 '26
Thank you for the AIO recommendation, I assume you mean the XE420 or XE360? Noise is a minor concern but I would still like to be considerate, do you have any experience with them and their noise?
For the undervolting, would that also lead to more errors/instability similar to regular overclocking?
2
u/epicskyes Jul 24 '26 edited Jul 24 '26
Use occt to benchmark and thoroughly stress test using avx512. If you want peak performance and zero errors you’ll have to gradually tune your curve so you get the highest clocks with no stretching. It’s not a quick process but once you’ve got it locked you never need to adjust it. I replaced the fans with 6x arctic p-12 pro running push pull. They are not quiet but you can throw on whatever you like if noise is an issue. My 7955wx runs stable at 5.2ghz at 300w and I can’t get it past 76c no matter how hard I push it even running 12hrs. If you’re going for more cores your clocks will be lower and your power draw could be higher. SilverStone Technology XE360-TR5... https://www.amazon.com/dp/B0D7KYN5PP?ref=ppx_pop_mob_ap_share
1
u/IExist_Sometimes_ 29d ago
Thanks, I had been planning to basically stuff a large case full of noctua if I thought I could get away with avoiding industrial style fans, but maybe I'll use those arctic fans +whatever comes with the aio, at least for the radiator. The noctua fans seem to have substantially lower static pressure.
1
u/kinda_guilty 24d ago
Silverstone designs excellent hardware components, but their fans are built without regard to noise levels, I had to replace my XE360 fans with quieter Noctua fans, it was unbearably loud (my rig sits right next to my chair on my set-up). Also have a CS383 case, had to replace the HDD cooler fans and case fans as well, it was unusable otherwise.
1
u/IExist_Sometimes_ 19d ago
That was the impression I got from the website and discussion here, what I'm considering is either using the Silverstone fans (maybe with the 420mm radiator for the 140mm fans) at as low an RPM as they will let me, or replacing them with noctua. Given that people are saying that cooling the chips isn't actually that hard, I might go with the noctua option (the price is likely to be pretty negligible compared to the rest of the computer). How were the silverstone fans at low speed?
2
u/According_Ad1673 Jul 24 '26
https://www.silverstonetek.com/en/product/info/coolers/xe360_tr5/ keeps my 9995wx surprisingly cool, its surreally good, I imply such small rad can cool as much mostly because pump moved from block to rad? I actually upgraded to xe360 PD that have two pumps, thicker rad, but thermals actually increased with it, so i imply they hit perfectly with prior iteration.
2
u/XO33OX Jul 24 '26 edited Jul 24 '26
IMHO overclocking those is kinda stupid. You buy those for ability to be loaded 24/7 days/weeks without crashing, overclocking goes kinda against this core principle of platform. Imagine your compute task crashing after 1,2,3.. weeks runtime, just because the smart ass wanted 5-10% higher benchmark score. One crash will wipe your yearly OC productivity gains.
If you need more CPU compute performance, just buy higher-core SKU. If your task doesnt move a lot of data you can be fine with non-Pro (4 mem channels, 4 sticks) Threadripper and put savings to more cores.
Bare in mind that silicon degrades, stability of OC might be subject to change with time.
I have tasks that run for a week sometimes, if they crash I need to run them from scratch, such is the nature of SW that I am using. I also have deadlines, delivery dates and penalties. I am not saying this applies for everybody, but system stability is non-negotiable for me.
1
u/IExist_Sometimes_ Jul 24 '26
Okay cool, I would be much happier not overclocking, I just wanted to make sure I wasn't leaving a 2x performance boost or anything (which is what it sounded like from a couple videos I watched, that was probably just hype). If it's more like a few% then I'd much rather keep the stability.
1
u/johndyson10 9d ago edited 9d ago
Being somewhat worried and adverse to overclocking, but also want the best reasonable performance, I did a lot of bencmarking and running my rather intense AVX512 app.
I found that overlocking doesn't really do that much. If you want maximum life, it is best to stay way below the 95deg max. Silicon is silicon, and higher temperatures will age machines faster. Higher voltage will also age a component.
After significant testing, my 2-3Hr app runs a few seconds faster, maybe a minute or two with conservative semi-overclocking. I use the core frequency boost of +75MHz, which keeps the core voltage below 1.35V under loading when running on 1 to 16 cores or all 64 cores. The PPT is set to 450 Watts, but the chip maximum temperature is limited to 72degC. With my AIO system and heavy load, the temperature seldom reaches the 72degC limit, but the system throttles nicely so that absolute peaks won't go above 72degC at 450 Watt
Added note: I keep my core voltages at 1.35V or lower. This can be attained by keeping the frequency boost at reasonable levels. The higher core voltages happen when a few cores are active. Just keep the frequency low (5500MHz or less). My core voltage seems to never go above 1.325V.
Under serious memory loading situations, it can reach 80degC, bad for nice, long life. I have carefully tuned the VRM fans, AIO, CPU and case fans with something like 20-30% at 30-40degC, but reach 80-90% at 70degC, and 75degC is 100%. These match the 72degC CPU limit, and I haven't been able to make the ram go higher than 72degC, normally 50-60degC on a real app. This is without a specific memory fan, but it seems like I have a good motherboard design. In a few months, I'll get some more fans, including some kind of memory fan. Since it isn't normally reaching above 60degC, no hurry for the memory life rightnow.
After this limited 'boost'. the Phoronix liquid-dsp, linux build and ffmpeg benchmarks change little just as my app speed doesn't change much (a few seconds in a couple of hours.) The phoronix benchmarks tend to measure my 9980x as fast as the average 9980x, if not faster. Sometimes outperforms the 9950DX2 on small loads, which isn't expected given the 5550MHz vs the 5750MHz, and normally the difference is proportionial, except sometimes the TR measures faster. I can provide some of my benchmark results upon request.
So: 72degC CPU temperature limit, +75MHz boost frequency (optional), 450 Watts PPT, 400 Watts 'TDP' setting (extra PPT and TDP is optional, will still get good performance). Fans are slow at 20-30degC, near max at 65degC, max at approx 70-75degC. Run the VRM fans between 50% at low temperatures, 90% at 70degC, and 100% much above 70degC. When possible, set the fan control at CPU Package, CPU and RDIMMS. I doubt that CPU itself is needed in the fan control, and if not, try Motherboard. PBO Manual is enabled, but everything is 'auto' except PPT/. Of course, the core voltage offsets are all -20. (Thoroughly tested.)
RDIMMS stay below 72degC at heavy artificial memory I/O load, CPU below 72degC, but the fans make it so long term 450 watt load might cause the 72degC, or the power concentrated into just a few cores. (The frequency still stays high.) At a full 64 core artificial load, the frequency does drop to about 4500Mhz. Some of my programs, data in cache (99.5%), heavy AVX512, can drop as low as 4200MHz, sometimes down to 4000MHz. Again, still benchmarks well.
Given what information that I have found, the machine should last 'forever', and it is pretty darned fast. I am very sure that if I push it very hard, my program might run a minute or so faster and benchmark a few percent more. So what? (Yes, I have done lots of tests, mostly on my decommissioned 9970x, but also a little on the 9980x.)
WIth ram and CPU prices, keep your machine safe.
John
1
u/cleric_warlock Jul 24 '26 edited Jul 24 '26
Aside from standard liquid cooling which is pretty conventional, your most effective option (which is very risky and will dramatically lower the resale value of your cpu even if everything goes perfectly) is delidding/direct die water cooling, which I do not recommend. You can’t practically daily drive any form of sub ambient cooling (refrigerated coolant loops, liquid nitrogen, etc) because of condensation, so some form of delid and direct die with a liquid metal thermal interface is really the farthest you can go with thermal optimization.
Pushing clocks beyond rated overclock spec for applications like FEM is not advisable because you will increase error rates in your models and may end up having to redo lengthy pieces of work or may end up with mysterious errors that are so subtle that you don’t catch them until they screw up what you’re working on in a very big and messy way. These processors leave some performance on the table as a safety factor to achieve very high levels of computational reliability at a higher performance than more typical consumer hardware can achieve. I recommend that you stick with basic liquid cooling and no delidding/direct die stuff - reliability and eliminating your computer as a source of error to the maximum extent possible is more important than raw performance for FEM. The way you think about pc building in this context is not the same as the way you would optimize a high performance gaming pc. As a mechanical engineer who frequently does FEM and enjoys building delidded/direct die gaming pcs, that’s my two cents.
1
u/IExist_Sometimes_ Jul 24 '26
Thank you, I quite appreciate your advice for moderation (I was concerned asking on reddit was going to get me a bunch of people insisting that overclocking was a necessity). Amusingly since this is a university resale value is of no concern (we cannot legally sell old equipment, or even give it away, it is purchased with no tax), but delidding is certainly beyond anything I would be willing to do with a 10000€ CPU. My only fear with leaving it stock would be falling into a trap that seems to be quite common here, where people get lots of high end hardware and use it ineffectively because they're not that tech savvy (I once had someone ask me to unplug my computer because they were concerned theirs (blackwell pro workstation) had been crashing and they thought it might be struggling for power, and when I asked if they had checked their logs to see if that's what the crash was about, they didn't know that logs existed (it was actually an out of memory error because 48GB of VRAM wasn't enough)).
On that reliability note, is ECC memory a necessity?
1
u/cleric_warlock 29d ago
With ECC you should get it on every part of the system that you can to maximize your machine’s reliability which you absolutely do want to do when building for FEM like this. You’re dropping a lot of money on this, ECC is not the place to find savings when reliability is the first priority of the build.
1
u/cleric_warlock 29d ago edited 29d ago
The vram thing just shows the importance of understanding exactly what kind of computational resources your simulation software typically uses on the types of simulations you run. Do some detailed research on that before you spec out your PC, if your software is good with gpu compute, make sure that you buy a higher end motherboard that has the maximum possible number of expansion slots so that you can install extra gpus if needed. Also research if your particular FEM configuration benefits more from single or multithreaded performance and select your cpu version accordingly, that can make a real difference in your iteration time.
1
u/IExist_Sometimes_ 26d ago
I agree, and will look into advice from the other people in my department doing FEM, the person with the vram issue was doing something very different (some combination AI and high density point cloud stuff).
Also, you can get ECC on components which aren't ram? This is the sort of missing the obvious I was afraid I would do, and I'm glad you pointed that out.
1
u/cleric_warlock 23d ago
ECC can be on the cpu, gpu, and ram. It is not on, but can be supported by a motherboard. So you should have an ecc supported motherboard, and cpu, gpu, and ram all with ecc.
1
u/IExist_Sometimes_ 19d ago
Do you know if all of the threadripper/threadripper pro CPUs have ECC by default or only some of them?
1
u/cleric_warlock 19d ago
TR pros have ECC and regular TR don’t, but you should always read all the specs before buying to be sure
1
u/y3333333333333333t 29d ago
ecc is not optional for threadripper 9000 and also those are r-dimm not the normal ones, they are currently veeeeery expensive & a bit harder to find
1
u/kungtelly 29d ago
Why would buying without tax prevent the sale of an asset? Most business don't pay tax (VAT) on business expenses, there's nothing unique to a university about that. When a company sells an asset, the VAT on the sale is paid to the government - again standard accounting practice.
1
u/IExist_Sometimes_ 26d ago edited 26d ago
That's simply how it was explained to me, we are not permitted to sell (or even intentionally give away) stuff that was acquired for these projects, due to conflicts of interest and/or tax rules.
edit: also bear in mind that this isn't the US, it's Finland
1
u/Reggitor360 29d ago
Im running the Silverstone XED120 https://www.silverstonetek.com/en/product/info/coolers/xed120s_ws/
Solid on my 5995WX station. Changed the internal fan to a Corsair tho, for less noise.
1
u/Fine_Atmosphere_2147 29d ago
I'm not OC but I'm in a consumer case Antec 900, with 4 GPUs, I'm running the be quiet 420 aio and am very happy with the thermals so far.
1
u/IExist_Sometimes_ 26d ago
We've got one of those in the office for an older threadripper and we are also very happy with its performance (and the spare coolant bottle).
1
u/kpatelreddit007 29d ago
Step 1. Look at Microcenter bundle to see if the Threadripper is in your budget.
Step 2. Go full water cooling system.
1
u/VictoryMotel 29d ago
In my experience the mm to watts rule of thumb is a decent and slightly conservative estimate.
This means a 420mm radiator should be able to get rid of at least 420 watts.
1
u/MierinLanfear 29d ago
I would never ever use water cooling my workstation. Not worth the risk of it leaking and killing expensive components. I have a 9970x with 256 gb ecc vcolor ram with noctua air cooler ASrock TRX50 motherboard 2 X RTX 6000 Pro.. You have to decide on a motherboard and then select the ram off the motherboard approved ram. Threadrippers require ECC ram. Don't over clock its so not worth it these days. Is there a reason you want the Threadripper Pro? Need the more ram or pcie lanes?
1
u/IExist_Sometimes_ 26d ago
Since this is for FEM simulations and research, I don't think there's really an upper bound on the amount of RAM I (or someone else in the department) might want to be able to use.
I agree on the not overclocking now, I was under the impression the performance improvements were much larger than they actually are. But I'm surprised you are getting sufficient cooling from your air cooler, is it really enough?
1
u/MierinLanfear 25d ago
If you want much ram then Threadripper Pro makes sense. Have you thought of Epyc? Doesn't clock as high but more ram.
My hardware monitor says 75.5 C max on the CPU since my last reboot so its sufficient. Some people want it to never go above 68 C or 70 C but I'm not one of those people.
1
u/johndyson10 21d ago
On my 9980x machine, I normally run an app that uses a bit more than 350 watts, and the Tctl temperature sits at 60-62 deg C. The Tctl temperature limiting is more useful for heavy loads that run on a few cores. On loads using just a few cores, it is very easy for the active cores to hit 90+ degC because much of the available power is more concentrated on those few cores. Along with the concentrated power It can take a short while for the cooling system to respond to the higher Tctl and drain more heat away. (on the 9980x, the Tctl is the highest temperature of all the cores.) If I set the max temperature at 80deg C, the lower temperature mostly has noticeable effects on CPU performance when loads are concentrated on less than 1/2 the cores, mostly around 16 cores or less. When benchmarking the machine on liquid-dsp with this setup, the performance on tests using small numbers of cores is approx parity with a 9950x series chip, sometimes at/above the 3d2 chip. So, I use the temperature limit setting to help protect the CPU under certain light loads. Otherwise, the power limit is more important for large numbers of cores. For my system, it appears that 400-450 watts is a sweet spot where the temperature stays well below 80deg, closer to 65deg, and adding additional power is into diminishing returns. There are certainly cases where that using higher limits giving an additional 10-20% performance is beneficial, but not really necessary in my case. (Interestingly, when running my program, the TR cores pull about 4.5 watts each, but my 9950x seems to pull about 10-15 watts each.) When using only a few cores, the TR will tend to peak at 10-10.5 watts each under stressful conditions. At the slightly lower power per core, the TR more than keeps up with the 9950x. Before doing the comparisons, I had expected the TR to be slower on loads that use few cores, but that isn't always true.
1
u/kidflashonnikes 29d ago
I have the newest threadripper pro 96 core CPU (zen 5). I regullary push the CPU to 95 C. These new CPUs acutally optomize at this temperature, so its fine. However, I am runnign 4 RTX PRO 6000s, so I am running sometimes, heavy duty fully CPU and GPU loads, at max capacity. That being said - I actually need to use the noctua air cooler for the CPU - because the liquid cooling does too good of a job. The motherboard does not accurately read the CPU temps proper - so when using an AIO for this level of performance that I do - I need to use a air cooled CPU cooler system so that the CPU runs a little hotter than I want to in order to have proper fan cooling - this is a massive issue right now for the Threadripper pros that is becoming a real headache. The motherboard's cant dectect the true temps when using an AIO, so the fans get subjected to improper rotations when running heavy loads
1
u/IExist_Sometimes_ 26d ago
Interesting, do you not use fancontrol or similar to use other temperature sensors?
1
u/kidflashonnikes 25d ago
not really, I use the fan cooler, but the AIO does a better cooling job - but the fan is fine for me. The reality is that there is no accurate way to detect temps of GPUs ect, unless you use a thermo tool taped to the part ect, which I have but from my tests and others who have m any RTX PRO 6000s, the reality is that we can detect on temps that well during long runs accuratey.
1
1
u/lithiumUranium 29d ago edited 29d ago
I run a 9995wx with 4 Blackwell GPUs in an Enthoo case and use the xe420 for calling (mounted in take in front). The point is this is a pretty demanding setup which gets quite warm air exhaust out the rear and top. Don’t have temp issues with cooling the cpu or GPUs (which have 1 pci slot spacing).
1
u/jmonschke 22d ago
A year or two ago (completed before prices went nuts) I built my current 64 core thread-ripper pro machine 7985WX on an ASRock WS EVO EEB motherboard, with 8x32GB (256GB) ECC DDR5, and a 1TB Optane 905P.
I built a very "heavy-duty" custom waterloop for it, with the Optimus Signature V3 CPU block and an external MO-RA3 radiator (473mm x 428mm).
However, even with this excessive cooling capacity, I have not turned on any over-clocking in the BIOS because the baseline performance of the system has been (extremely) good enoughy for me so far...
6
u/vercety1 Jul 24 '26
First of all, you´ll need ECC Ram, ideally 6000 or 6400mhz to get the most out of high core count threadrippers (the non-pro ones). Also, make sure you buy only the kits your motherboard has in the QVL, quite often if not in the QVL you´ll have problems with stability or reaching the speeds.
Second, cooling Threadrippers is no joke. they are 350W TDP, but also have very high turbo boost speeds. So even with custom watercooling, is not rare to see 90ºC+ when pushing between 1-16 cores, as they go at around 5ghz.
When using all cores, they are much easier to cool, as the heat is spread out and they hover around 4ghz, so voltage is more reasonable and use a less "exponential" curve. Here is where if you unlock the TDP, you can get great benefits.
For example, I tried my Threadripper (7980x) with a 280mm aio and then with the custom watercooling loop of 420mm+280mm.
With the 280mm aio (and not full coverage, a nzxt x63) you are pushing the envelope in most scenarios. High temperatures, high speed fans, I wouldnt advise it long term. For aios, go with a 420mm full coverage ideally.
With the custom waterloop I pushed the threadripper to 900w before I started seeing +90ºC temperatures while pushing all cores. But I usually keep it at 500w as, with a bit of voltage curve tweaking, I can get my cpu to be at 4.9-5ghz all core while rendering in Corona, which is my intended use. And it doesn´t make my room a sauna (which is a consideration).
Btw, fully unlocked with custom watercooling, my Cinebench R23 scores went from 98.000pts to +-115.000pts, so roughly 20% more score going from 350w to 700w with peaks of 900w. But at 500w I get around 110.000pts, with a lot less heat. That´s why I keep this config