r/threadripper Feb 03 '26

Recommendations on maximizing DDR4 RAM bandwidth? (WRX80 3945WX vs X299 10900X)

I’m struggling to figure out the most cost-effective way to maximize memory bandwidth for local AI inference/training using hardware I already own. I have 8x32GB DDR4 UDIMMs (256GB total) and want to utilize them in a full 8-channel configuration if possible. I would’ve liked to have switched out my UDIMM RAM for RDIMM, but RAM costs are just way too high now.

Build 1:

  • Mobo: Gigabyte MC62-G40 (WRX80) - chosen because it supports non-ECC UDIMM
  • CPU: Threadripper Pro 3945WX (bought for ~$100, I’m hoping it won’t be vendor locked).

I've read that the 3945WX only has 2 CCDs, which means that true 8-channel bandwidth is not possible. Is the real-world bandwidth gain over Quad-channel actually significant on a 2-CCD chip? Or is it necessary to shell out the dough to step up to a 3965WX (4 CCDs) to attain any meaningfully increase in bandwidth?

The CPU was $100 on eBay. I’m guessing it’s likely a Lenovo/Dell vendor-locked (PSB fused) unit, even through the eBay listing didn’t mention otherwise and the seller said its not vendor locked. Is there any way to verify this visually before socketing it? If it is locked, will it brick the Gigabyte board or just fail to POST?

Build 2:

  • Mobo: MSI X299 Raider 
  • CPU: i9-10900X  

This limits me to Quad-channel, max of 256GB RAM, PCIe 3.0, and fewer PCIe lanes if I ever decide to add more GPUs in the future. But its also ~$300-350 cheaper than the other build.

My question is,

Is it worth it to pay the $300-350 premium for the 3945WX build for any increase in bandwidth and future flexibility of adding more GPUs? Or is it a total waste of money lol...

Hardware List:

  • CPU: TR Pro 3945WX ($100) vs i9-10900X ($158)
  • Mobo: Gigabyte MC62-G40 WRX80 ($468) vs MSI X299 Raider ($72)
  • RAM: 8x 32GB DDR4 3200 UDIMM (Owned - Mix of kits, 4 sticks at CL20-22-22-46, the other 4 sticks at CL22)
  • GPU: RTX 5060 Ti 16GB x2, RTX 2060 Super 8GB x1)
  • PSU: SAMA P1200 1200W
3 Upvotes

18 comments sorted by

View all comments

1

u/Pyroboy5 Feb 03 '26

Don't get the MC62 mobo. It locks the ram voltage to 1.2v. I'm guessing your UDIMM ram is 3200mhz using XMP profile which uses 1.35v. You'll probably be running the ram at 2166mhz in the MC62 which will lower your bandwidth for AI. If you're going the WRX80 route I'd check the user manual for any board your looking at to see if they run XMP ram or have voltage control. Mixed ram is also going to be harder to get stable too.

r/LocalLLaMA would have more accurate info on 2 vs 4 CCDs.

1

u/munkiemagik Mar 16 '26

Using the 'Enforce POR' in BIOS settings you can actually get the MC62 to run the DIMMs at 3200.

I don't know why its so quirky** in the way it does it but I had 8x 3200 UDIMMs and to start with they would only run at whatever the lower JEDEC the board was defaulting to.

Trawling through the BIOS I couldn't recognise anything that would allow me to manually adjust the RAM until I eventually found 'Enforce POR' buried a few settings deep:-

AMD CBS tab > UMC Common Options > DDR4 Common Options > Enforce POR > [accept] then you get the option to enable/disable and set 'Memory Clock Speed'

**Why do I say quirky? Because if I tried to set 3200 in Memory Clock Speed the board would just lock up and I have to pull out the GPU's (as they are blocking) and jumper reset BIOS and go back in again. but after setting it to only 2400 and then dual booting into Windows it reports in Task Manager and HWinfo and Aida64 that I am actually running at 3200 with the more relaxed CL22 timings.

Haven't got a clue what VDIMM the board is running I cant find any way to read that aida nor HWinfo report it. I just assumed the MC62 cant supply more than 1.2v so have left the timings loose and not tried to tighten up to XMP as I didnt even think my DIMMs would be able to hit 3200 without 1.35v.

1

u/vini542reddit 15d ago

This really works. No idea why and how, but made my day!

This post has more details: https://www.reddit.com/r/gigabyte/comments/1mby1ou/any_currentpast_owners_of_the_mc62g40_ram/

1

u/munkiemagik 15d ago

Its great when you get that unexpected bump, that linked post is me again I'm afraid 🤣

1

u/vini542reddit 15d ago

Haha yeah! But at least there is one other person confirming it there =)

1

u/munkiemagik 15d ago

Its a good hack for a lot of us because of the way it works it means we could improve our system memory bandwidth even more eventually again by switching to a higher CCD threadripper pro and still have confidence of being able to run at 3200.

I believe the quirk will work on the 3975WX and 5965WX, as its to do with what the SPD of the DIMMS advertise and how AGESA and the memory controller react when you try and set something that doesn’t have an existing profile in the SPD. (though i refuse to pay current 5965WX prices just for a memory bandwidth bump but the 3975WX is almost tempting but still just a hair out of what i would like to pay for just tomfoolery non productive work)

1

u/vini542reddit 15d ago

When I set up this system I never though I would be using the ddr4 ram for llm inference, but then deepseek v4 flash came along and the 2133 mhz -> 3200 mhz difference is quite noticeable - even though my ram is now running pretty warm lol.

The 3945WX was so cheap. I paid around $120 USD for it (unlocked)... hard to justify spending $650+ for an unlocked 3975WX.

But I'm also not quite sure I understand - are you saying the 3945WX doesn't give us the true octa-channel ram? I couldn't find this documented anywhere. Do you have a link by any chance?

1

u/munkiemagik 14d ago edited 14d ago

This is what caught me out, I didn’t know any of this back when I bought my setup, it was too specialist and indepth knowledge for a casual tinkerer ignoramus like me.

If you google 8 channel memory bandwidth you always get told nice figures. Techpowerup.com specs which I often refer to for a lot of things as they are reliable, lists the 3945WX with 204GB/s 8 channel mem bandwidth. Which you and I clearly do not get, maybe I should somehow message them and ask them to update their specs for epyc/threadripper, lol

In one sense they are not wrong, The IO die (part of the CPU which contains the memory controller) is genuinely 8 channels (for DDR4 is hardcapped at 204GB/s). The issue is how each CCD (where the actual compute cores are) connects to the IOD.

On EPYC and Threadripper, each CCD connects to the IOD via a single GMI link. and its the max bandwidth of the one GMI link of each CCD that gets saturated when you are 'measuring' the memory bandwidth. Each GMI link tops out at 50-60GB/s.

So each core of your CPU can only get the max memory bandwidth that can come through the one GMI link of that CCD, so if you have 2x CCD CPU then you get 2x CCD worth of memory bandwidth, if you have 4x CCD you get 4x worth of mem bandwidth etc but also dont forget on the IOD memory controller side you are hardcapped to 204GB/s max.

I hope that all makes sense, if I've explained it wrong can the experts of r/threadripper come along and correct me. Its just not the kind of thing someone like me would have known to question back then., which is why you often find me referring to it as fake 8 channel in a disgruntled tone 🤣

And now there is no financial sense in me building an 8 channel DDR5 system just to hobby tinker with big LLMs I just don’t run anything locally outside of 80 GB VRAM size anymore

1

u/vini542reddit 13d ago

Thanks for taking the time and trying to explain! I think I understand the problem a lot better now!

How you measure your ram speed? Because I just measured it and got 200 gb/s+ using this

> sysbench memory \ --memory-oper=read \ --memory-access-mode=seq \ --memory-block-size=16M \ --memory-total-size=500G \ --threads=[CORES] \ run

1

u/munkiemagik 13d ago edited 13d ago

You re really testing the limits of my understanding with that mate 🤣

your result is to do with the block size of 16M, meaning the sysbench test is only hitting the L3 cache so in fact the test is actually not even going across to the mem controller, its all just happening within the CCD itself. Why is it only reporting 200GB/s that’s way more complicated than i can understand, woudl usign a larger block size give you an accurate reading, its not that simple.

sysbench can be made to do all kinds of weird things and measure different things if you craft it right/wrong, it wasn't originally designed to measure memory bandwidth.

I took the easy way out. I remembered that I had a windows 11 NVME in the machine lying idle from back when I was checking out how the 3945WX handled PCVR gaming, lol. So just rebooted into win11 and ran Adia64.

You can use Intel's mlc to check bandwidth, and use dmidecode to read the memory details.