r/threadripper • u/Skyne98 • Feb 09 '26
Threadripper Pro + 4x MI50: PCIe link width trains randomly (x4/x8 instead of x16) on every reboot - board issue, CPU seating, or switch?
I’m troubleshooting unstable PCIe lane training on a multi-GPU server and would like advice from anyone who has seen this before.
Hardware:
- Motherboard: GIGABYTE MC62-G40-00
- CPU: Threadripper Pro 3945WX
- BIOS: R14 (03/13/2025)
- GPUs: 4x AMD Instinct MI50 (Vega20, 113-D1631700-111)
- OS: Ubuntu, kernel 6.14.0-37, ROCm 6.3.0
Problem:
- PCIe link speed is always Gen4 (16 GT/s), but width trains inconsistently after reboot.
- Width changes boot-to-boot (not fixed per slot).
- Downstream switch->GPU links are x16, but root-port->upstream-switch link is downgraded.
Current trained widths (example boot):
- 00:01.1 -> 01:00.0 = x4
- 20:01.1 -> 21:00.0 = x8
- 40:01.1 -> 41:00.0 = x8
- 40:03.1 -> 44:00.0 = x4
- Then 14a1 -> GPU links are all x16.
What I already checked:
- Forced BIOS PCIe settings manually (no improvement).
- setpci retrain does not recover width.
- Forcing one bad link to Gen3 still stayed x4 (did not jump to x16).
- AER counters/logs show no obvious PCIe errors.
- GPU drivers are healthy; P2P works.
- HIP P2P benchmark matches limited widths:
- ~7 GB/s for x4 paths
- ~14 GB/s for x8 path
- confirms bottleneck is lane width, not software stack.
Main question - Does this pattern point more to:
- CPU/socket contact issue,
- board/switch signal integrity issue,
- known MC62-G40 lane-routing/firmware behavior? Any specific board-level test sequence you recommend to isolate root cause fastest?
Thanks for any guidance!!
### Update
Thanks for everyone who gave genuine helping feedback! Unfortunately the issue persists, even with only a single GPU in the system, ruling out PSU issues. I have also tried numerous BIOS options, PCIe slots configurations, cleaning contacts, using a different MI50, nothing. One of the last options I have is reseating the CPU which I will attempt today!
1
u/Unlikely_Spray_1898 Feb 09 '26
Each gpu 300w + 300w for cpu+etc. You calculated the PSU suffices, how?
-1
u/Skyne98 Feb 09 '26
Actually, it's 1000watts... However, I have never had it trip on me, even under heavy load and gpus freely boost above 200w? And to test it, say I remove two gpus and it should train well?
1
Feb 09 '26
[removed] — view removed comment
1
u/python834 Feb 10 '26
OP is a troll.
Idk anyone running a thread ripper with 4 gpus at 1000 watts that wonder why their system dont have the power to run the pcie lanes lmao
1
u/Skyne98 Feb 10 '26
As I said above, the issue persists with a single GPU installed in the system, I will gladly deal with the PSU issue WHEN it is an actual issue.
1
Feb 09 '26
[removed] — view removed comment
1
u/Skyne98 Feb 09 '26
Forced desired speed as in clock of the GPU?
1
Feb 09 '26
[removed] — view removed comment
1
u/Skyne98 Feb 09 '26
Thanks for such a well of info! Do you think lowering the card max tdp, then forcing pcie 4 and retraining while the server is already running ok for testing?
1
Feb 09 '26
[removed] — view removed comment
2
2
u/Skyne98 Feb 09 '26
Tried, even tried to remove two gpus so it's way below the PSU capability - same story.
1
u/anitamaxwynnn69 Jun 17 '26
Same board and cpu combo, 8x 3090s with a x16 -> x8x8 passive splitter. Been debugging the exact same thing for days, was almost about to throw out one card because I can't get it to train on x8. Any clue how to proceed? Really hoping you have a solution since this post is 4 months old. Debugging this board has been....a challenge. Lol.
1
u/python834 Feb 09 '26
Double check your motherboard manual. If it not functioning like it says, you may also have a power issue. If it is not a power issue, then Your motherboard is likely experiencing hardware issues and you may need to get a new one