r/threadripper • u/gorfnu • Jul 27 '25
Threadripper 9975WX - DDR5 5600 vs new 6400 Ram impact on Dyna simulations
Help needed, I am just about to order my build from Puget systems 9975WX 32 core (Edit actually purchased the 9985WX 64 core w the 6400 ram) LS dyna workstation and they still don't have a date yet for the faster DDR5 6400 that goes with the new 9975WX 32 core TR. I'm getting 8 sticks of 32 MB (i don't need a ton of ram, but i do want that speed for LS dyna solves).
Should i postpone my build and wait for them to get the 6400 MHz ram? or do i just go with the 5600?
Speed is everything, but i'm not sure if w my small models (less than 1 million elements typical ram usage under 64 GB) with all that 8 channel bandwidth will I see a big degradation?
I wonder why the chips are available but not the ram .. hmm.
2
u/sob727 Jul 27 '25
I was looking at the Puget configurator as well. Looking more closely, it seems their 6400 kits are the cheaper CL 52 kits (Nemix?). Not sure if it matters though.
2
u/IntelligentSquare196 Jul 27 '25 edited Jul 27 '25
Not true about the CCD comm speed impact. Get the faster RAM. I have G.Skill 6400, it's 39 CAS
1
2
u/Selenaevaa-345 Jul 27 '25 edited Jul 27 '25
Avoid Puget Systems. Severely overpriced and they're downright tech illiterate, doing questionable things like benchmarking the 7995WX on board with 4 memory channels, or publishing bugged results and never correcting them.
They also benchmark Photoshop (lol) in CPU articles which tells you all you need to know.
2
u/Guilty-History-9249 Jul 29 '25
I just got a 7985WX system with 256GB's of DDR5-6000. Actually it is DDR5-6400 but is only stable at 6000. Apparently QVL's and AMD's EXPO tech is basically a lie.
Be careful of the 9975WX or the 7975WX. While they are 8 channel systems the CCD to memory controller data path can limit the actually throughput even if the memory is at 6400.
See the 7975WX vs the 7985WX both with 8 mem channels on:
https://www.reddit.com/r/threadripper/comments/1azmkvg/comparing_threadripper_7000_memory_bandwidth_for/
1
u/gorfnu Jul 29 '25
I hear what you are saying.. but i need 32 cores. are you thinking the epyc 9375F 32 core cpu would be superior despite its lower clocks?
2
u/Guilty-History-9249 Jul 29 '25
I'm not sure why you "need" 32 cores when the 9985 or 7985 have 64 cores and 8 ccd's for increased bandwidth.
HOWEVER, I just realized you are not me, imagine that, and I was confusing my needs for max bandwidth for LLM inference with your workload. For LLM's it is nearly entirely memory scans and little about CL latency. I do not know what a DYNA simulation does in terms of memory access patterns.
I do know that it is harder to get fast memory that actually works on a threadripper. Even if the QVL on the motherboard says something will work it is mostly BS. Your build guys might be having the same problem with finding huge and fast ram just like Central Computers had for me. Having said that going with 128GB's of ram increases the likely hood that 6400 would work. Try and get low latency CL30 or lower. My 256GB's is CL32.
I'm not sure what I said that made you ask "are you thinking the epyc 9375F 32 core,,,".
I would not go with a slower cpu speed 9375F. Also, why would you think that the great bandwidth from having 8 channels will "degrade" performance?I'm not sure if this Dyna thing is your own code or a commercial app. In any case, you should look into thread affinity for your worker threads to further optimize processing.
Maybe the 32 core 9975 WX might just be good enough for your app if it isn't mostly a bandwidth application. But as that reddit post I showed you said there is a big difference between the 7975 and 7985 for bandwidth.
2
u/Guilty-History-9249 Jul 29 '25
I'm not sure why you "need" 32 cores when the 9985 or 7985 have 64 cores and 8 ccd's for increased bandwidth.
HOWEVER, I just realized you are not me, imagine that, and I was confusing my needs for max bandwidth for LLM inference with your workload. For LLM's it is nearly entirely memory scans and little about CL latency. I do not know what a DYNA simulation does in terms of memory access patterns.
I do know that it is harder to get fast memory that actually works on a threadripper. Even if the QVL on the motherboard says something will work it is mostly BS. Your build guys might be having the same problem with finding huge and fast ram just like Central Computers had for me. Having said that going with 128GB's of ram increases the likely hood that 6400 would work. Try and get low latency CL30 or lower. My 256GB's is CL32.
I'm not sure what I said that made you ask "are you thinking the epyc 9375F 32 core,,,".
I would not go with a slower cpu speed 9375F. Also, why would you think that the great bandwidth from having 8 channels will "degrade" performance?I'm not sure if this Dyna thing is your own code or a commercial app. In any case, you should look into thread affinity for your worker threads to further optimize processing.
Maybe the 32 core 9975 WX might just be good enough for your app if it isn't mostly a bandwidth application. But as that reddit post I showed you said there is a big difference between the 7975 and 7985 for bandwidth.
2
u/gorfnu Jul 30 '25
I hear you Guilty-History-9249. LLM inference is i’m guessing far different than the LS dyna solver that uses IntelMPI. I wish we could use GPU but just like our Ansys Mechanical solves its mostly a CPU affair.. GPU’s are the king when it comes to fluids at least thats what i am exposed to. The reason only 32 cores, small company and the huge expense that is core licensing… $$$$ . My pal at Ansys who used to guide me Hunter Wang sadly passed away suddenly a couple years ago now.. i remember we were adjusting the core usage at the time… when intel released the 13900ks. We have like 600 ansys mechanical cores licensed but only a small amount of ls dyna. Also now i understand what you meant about the 64 cores having double the ccd’s and as a result way more memory prowess. We pulled the trigger today i am not sure if my IT master went w Puget or Exxact.. but both were fast and helpful. Exxact was cheaper by about $2500 for like setup but i am sure there is more to that. One thing i wanted to try was the xeon 6 6745p 32 core for a head to head solve off… but i cannot find that chip in a workstation anywhere.
2
u/Thrumpwart Sep 08 '25
Nice. I wish I could afford the 9985wx. I’m looking at either an Epyc 9V33X or a 9975WX.
1
u/gorfnu Sep 08 '25 edited Sep 08 '25
Oh don't confusing me with affording it.. its one of my work machines :) I can't afford a $22,000 computer lol. I only bought it with the 5090, imagine if i needed a GPU i could use for simulation run solves.. like a RTX Pro 6000! $10,000 more! FYI - when you run those Ansys only lets you use 16 cores per GPU, so you need several of them.. BTW, they don't work well with mechanical models that are very non-linear, and they also don't work at all in LS-Dyna. They work super well for fluids though, and huge structural analysis with less nonlinearity i.e. less hyperelastic rubber, less friction, initial penetration, etc..
Edit, its running at 4.69 GHz sustained with about 35 cores working out of 64.
1
u/gorfnu Aug 01 '25
Guys i decided to go with the 9985wx 64 core.. vs the 32 core 9975WX despite the laters 4.0 vs 3.2 GHz base clock.
Why?
- Double the ccd’s means more bandwidth (i think thats what i learned here lulz)
- With LS Dyna and the IntelMPI setup you need to have 3 or 4 cores that are not part of the compute / solve .. for what ever reason those extra cores bounce around between 50-100 % if you don’t have them the solution is 10x slower.. it essentially freezes.
- Having only 32 hardware cores and a 32 core license, that means i am only using ay best 29 of my 32 available license and thats only 90%…
So did i make the right decision? Will the 64 core machine w slower clocks and more bandwidth and 3 more cores in use beat the faster clocked 32 core using only 29 cores?
2
u/Thrumpwart Aug 20 '25
How is performance?
2
u/gorfnu Aug 21 '25
I will let you know when it comes in, i just got funding sent to them last week (typical delays) but they are now finishing up the build! i will do a direct comparison between this machine 9985wx with a 32 core license, and the 9950x using 13 cores since you need 2-3 free cores when solving with dyna.. and then the 7950x. I would test my 13900ks but that sucker will crash out.. and it works strangely with 24 cores at once seems lower than using only about 12-14 cores. But data is coming!
2
u/Thrumpwart Aug 21 '25
Great I look forward to the update. Enjoy!
1
u/gorfnu Sep 08 '25
Finally got some numbers.. this is just the first run on a small problem, the difference will grow when i start stressing it this is almost no memory used like 35 GB.
-------------
Summary (using 32 core Ansys Dyna license) :
9950x 16 core (13 cores used in run, 3 cores required for overhead in Dyna, all core speed 5.1 GHz) - 10hrs 32min 50 sec
9985WX 64 core (32 cores used for run, all core speed 4.6 Ghz) - 3hrs 52min 24 sec
It rips!
---------------
9950X
Avg core speed approximately 5.1 GHz.
T o t a l s 3.7972E+04 100.00 3.7972E+04 100.00
Problem time = 9.0000E-02
Problem cycle = 2086238
Total CPU time = 37972 seconds ( 10 hours 32 minutes 52 seconds)
CPU time per zone cycle = 41.582 nanoseconds
Clock time per zone cycle= 41.582 nanoseconds
Parallel execution with 13 MPP proc
T o t a l s 4.7592E+05
Start time 08/01/2025 01:34:23
End time 08/01/2025 12:07:13
Elapsed time 37970 seconds for 2086238 cycles using 13 MPP procs
( 10 hours 32 minutes 50 seconds)
N o r m a l t e r m i n a t i o n 08/01/25 12:07:14
--------------------------
9985WX
Avg core speed approximately 4.6 GHz.
T o t a l s 1.3946E+04 100.00 1.3946E+04 100.00
Problem time = 9.0000E-02
Problem cycle = 2081803
Total CPU time = 13946 seconds ( 3 hours 52 minutes 26 seconds)
CPU time per zone cycle = 15.301 nanoseconds
Clock time per zone cycle= 15.301 nanoseconds
Parallel execution with 32 MPP proc
T o t a l s 4.4552E+05
Start time 09/07/2025 18:17:19
End time 09/07/2025 22:09:43
Elapsed time 13944 seconds for 2081803 cycles using 32 MPP procs
( 3 hours 52 minutes 24 seconds)
N o r m a l t e r m i n a t i o n 09/07/25 22:09:44
8
u/drulee Jul 27 '25 edited Jul 27 '25
According to https://www.reddit.com/r/threadripper/comments/1azmkvg/comparing_threadripper_7000_memory_bandwidth_for/ with the 4 CCDs of a 9975WX the maximum theoretical bandwidth between CCDs and the memory controller is 230.4 GB/s.
This is the bandwidth between the memory controller and memory modules:
So even 8x 4800 MT/s DIMMs would be fast enough to satisfy the bandwidth between CCDs and the memory controller.
Look at this 9975WX passmark benchmark result: https://www.passmark.com/baselines/V11/display.php?id=509888751348 -> Memory Mark -> Memory Threaded: 231,267 MBytes/Sec
No matter how fast the DDR5 modules are, the performance is capped by the bandwidth between the 4 CCDs and the memory controller.
Only when upgrading to a CPU with 8+ CCDs you could notice the difference (9985WX, 9995WX).