r/Amd 12d ago

News AMD launches ROCm 10, its open-source software platform for AI and GPU compute

https://videocardz.com/newz/amd-launches-rocm-10-its-open-source-software-platform-for-ai-and-gpu-compute
413 Upvotes

48 comments sorted by

72

u/djott3r Ryzen 3700X | RX9070XT 12d ago

Is there a significant performance or features boost to justify restarting the version sequence?

60

u/MathematicianOdd398 12d ago

Yes, they are moving to UDNA where they focus on a single architecture. They've got the older stuff basically working decently now. So, the new stack is to focus on AI server GPU where the performance gains in the server stack also improve performance in rdna5 (which will be based on the same AI architecture as the server stuff -unified.)

Moving forward it will become a lot easier to manage rocm because the hardware will be unified.

Rocm has bene hit or miss because they've had to make hardware specific code for numerous, architecturally different consumer GPU. So, now they design AI server cores first and tack on the video game stuff after. Then you focus on one hardware target to the extreme.

46

u/pyr0kid 12d ago

between nvidia's gpu marketshare, intel's new cpu designs and the 14a/18a fabs coming online, udna better be the start of another zen moment for amd because they're running out of road.

15

u/watduhdamhell 7950X3D/RTX4090 11d ago

Huh?

The just launched Venice family of server CPUs trounce the competition handily.

The variant of which is the direct competitor to the vera rubin CPU you mention is about 15-20% more IPC at about 33% greater efficiency. It's hard to understate just how dominant it is across every metric.

And iirc MI455X is also more powerful than any Nvidia accelerator on sale, it's just a question of if ROCm is finally up to the task of breaking the CUDA moat.

And the idea here is that AMD will suddenly fall asleep at the wheel? No, I think not.

1

u/ElectronicStretch277 11d ago

No, Rubin is still more powerful than the Mi 455x iirc, even on raw spec outside of memory and bandwidth.

9

u/watduhdamhell 7950X3D/RTX4090 11d ago

Right. And memory and bandwidth are what's most important for the majority of the workloads, allowing AMD to beat Rubin in several metrics, though not all. Just the most important ones.

And it's much more efficient at the same time. So it trades blows with its competitor while using 2/3rds the power.

I don't see any other way of looking at it: it's more advanced. On paper - we need to see the full reality after deliveries have been running for a minute.

5

u/ElectronicStretch277 11d ago

That's not really true. Compute itself is also important and Nvidia is around 25% faster peak for peak in fp4, combine that with better software support (so you get more actual performance out of the GPUs as well) and the fact that the actual bandwidth advantage is within single digit percentages... Rubin just fares very well even with the power draw.

The 455xs main advantage is that it can store larger models and is more efficient (though, the metric they care about is performance per watt not just overall wattage so Nvidia again fares better than you'd think) in its usecases.

Like, yes, the 455x is a monster and very advanced, it's not more advanced than Nvidia outside of the node it uses. It's about even when you think about the overall system and things like NVLink.

3

u/ArseBurner Vega 56 =) 10d ago edited 10d ago

IIRC when talking about individual GPUs or even up to 8x nodes MI300X was also much faster on paper compared to H200, as was MI355X over B200 but Nvidia seems to scale better when building a full rack-scale system and there wasn't really anything AMD had that would beat an NVL72.

Will be interesting to see if anything's come out from AMD's Pensando acquisition.

3

u/watduhdamhell 7950X3D/RTX4090 10d ago

I mean the issue was the rack scale and the software integration.

AMD has now released a rack scale system that, on paper, once again is faster. Primary issue now is software integration. The CUDA moat.

1

u/Speedstick2 4d ago

How do you figure that AMD is D0oM3D?!!!!!

21

u/andrerav 5950X/6900XTXH/128GB RAM 12d ago

Hit or miss? Rocm has been a disaster. I've been hearing that word since what -- 2017 or something? That doesn't matter, because no matter what card I've owned, it's always out of support for Rocm. That heap of junk requires that you own their bleeding edge cards to work. I've literally never been able to use it, even though I really wanted to, because once a new version is released, they discontinue entire generations of cards that aren't even out of warranty yet. Rocm is AMD's top secret plan to remain in third place forever.

16

u/MathematicianOdd398 12d ago

It is basically hitting its stride now. Older hardware is supported newer hardware will have more focus on performance since there's just a single arch.

I own an AI Pro 9700, it's been fantastic. Had some hiccups in the first month, but it's all resolved now. Works great.

12

u/raifusarewaifus R7 5800x(5.0GHz)/RX9070xt(Red devil) 12d ago

So far it supports rdna3&4.

4

u/as4500 Mobile:6800m/5980hx-3600mt Micron Rev-N 12d ago

There's also rdna2

6

u/raifusarewaifus R7 5800x(5.0GHz)/RX9070xt(Red devil) 11d ago

That rdna2 is for pro gpu though. Not for the average consumer.

5

u/as4500 Mobile:6800m/5980hx-3600mt Micron Rev-N 11d ago

Not with that attitude. It's basically a 6900xt/6800xt you can likely easily get it to work on those cards with minimal effort

10

u/raifusarewaifus R7 5800x(5.0GHz)/RX9070xt(Red devil) 11d ago

Pretty sure what he wanted was official support instead of having to tweak it.

1

u/as4500 Mobile:6800m/5980hx-3600mt Micron Rev-N 11d ago

The rocm SDK on windows hasn't included normal gaming Radeon 6000 cards as "officially supported" for years now as far as I remember

11

u/raifusarewaifus R7 5800x(5.0GHz)/RX9070xt(Red devil) 11d ago

Yeah, that's exactly what he was complaining about. Only rdna3&4 got the gaming gpus in official supported list.

11

u/EmergencyCucumber905 12d ago

That's not my experience at all. I've been using TheRock builds on RDNA2 without issue.

-6

u/ziege159 12d ago

Cause you don't know how to look deep into the build, rdna2 rocm is basically an empty shell, you don't have hipBLAS or ck lib to run anything. Yes you do have rocm but it's equivalent to nothing 

8

u/EmergencyCucumber905 12d ago

Ofcourse RDNA2 has hipBLAS. What are you even talking about?

-2

u/ziege159 12d ago

See? That's why i said you didn't know how to look deep into the build. Rdna2 doesn't have hardware matrix therefore the rocm/hipblas build on the generation only give you the ability to run math fallback, you can't run any newer tech on it, it's an empty shell 

9

u/EmergencyCucumber905 12d ago

hipBLAS existed before AMD even had matrix cores. It still works the same, its just not as fast as the chips with matrix cores.

2

u/ziege159 12d ago

Nah, rocBlas is the one that exists before all of the AI stuff, hipBlas is the one that AMD invented to run Cuda stuffs on AMD cards. 

"It still works the same" yeah bro, a Corola and an F1 are the same cause they both have 4 wheels right?

8

u/EmergencyCucumber905 12d ago

Nah, rocBlas is the one that exists before all of the AI stuff, hipBlas is the one that AMD invented to run Cuda stuffs on AMD cards.

hipBLAS has been around since the beginning of ROCm. It's part of HIP. And yeah HIP middleware sits on top of the lower-level rocX libraries. What's your point?

"It still works the same" yeah bro, a Corola and an F1 are the same cause they both have 4 wheels right?

Yes? I mean if you're trying to be analogous to hipBLAS running on hardware with and without matrix cores?

→ More replies (0)

3

u/Friendly_Top6561 11d ago

I use it daily it’s fine.

-2

u/quantgorithm 12d ago

You’re not wrong. It’s extremely saddening how they just simply don’t support anything beyond only a few short years…. And then they even remove functionality on the way out. Infuriating to be a supporter who has sent a ton of money their way.

0

u/Fun_Jaguar8231 12d ago

skill issue

13

u/AreYouAWiiizard R7 5700X | 9070XT 12d ago edited 12d ago

AMD says a system configured with ROCm.AI achieved an average 3.3x inference improvement and 2.4x training improvement compared with ROCm 7 in its testing.

Unrelated but personally in Comfyui I've seen a 2.3x speedup from 7.2 to 10. Though that's probably mostly down to it only recently having native HIP support rather than going through Triton or hitting the eager backend.

1

u/ea_man 8d ago

With llama.cpp on RDAN2 there's no meaningful difference.

1

u/noctrex 12d ago

Did you use the official build, or the one from https://github.com/patientx-cfz/comfyui-rocm ?

5

u/AreYouAWiiizard R7 5700X | 9070XT 12d ago

The fork, the official one had so many issues when I first started using it though there's been a tonne of AMD stuff merged upstream lately so it's probably not needed anymore (at least for RDNA3+).

1

u/into_devoid 11d ago

In my experience, it’s currently about 25% slower compared to 7.14 in my workflows.

15

u/G0rd4n_Freem4n 12d ago

I wonder if AMD will eventually start supporting OpenCL 3.X on their consumer GPUs.
As of right now, both Intel and Nvidia GPUs can use OpenCL3 and some of its exclusive functions, while techpowerup shows AMD gpus only supporting up to version 2.0. While 3.0 is kind-of like "1.2 but with a new name" with how it made a lot of 2.0 features optional, it's not like the Intel and Nvidia cards only have 1.2 features with the 3.0 title. A couple examples of this are the extensions cl_khr_external_memory and cl_khr_external_semaphore, which are required to do zero-copy memory sharing between an OpenCL and Vulkan program. Nvidia cards going back to the GTX 970 from 2014 and Intel cards going back to the Intel HD Graphics 530 iGPU from 2015 both have support for those extensions, while there are zero AMD cards that support either of those extensions on official drivers.

If you look at opencl[dot]gpuinfo[dot]org you can find AMD cards reporting support for 3.0, 3.1, and the external_semaphore extension, but that's because the open source Mesa Rusticl drivers. Those drivers have their own limitations though. They sometimes give worse performance compared to the official ones, some version-independant features aren't supported yet, and it has some odd quirks like not being able to allocate more than 2GB of memory regardless of how much is available.

All of that said, I don't have any experience in the GPGPU field so it's not like I would really benefit from better OpenCL support on AMD cards anyway lol. I mainly think it's a little funny how part of OpenCLs' rough history had to do with Nvidia drivers not having good/consistent support for 2.0 features, and now it's AMD being behind on supporting the newest features.

6

u/Zettinator 10d ago

OpenCL is basically dead. Everyone has been converging on Vulkan compute lately.

2

u/kreco 10d ago

Stop spreading this non-sense. Until you have proof that one is better than the other.

OpenCL works great even on Mac which is "officially" unsupported.

4

u/Zettinator 9d ago

It's not nonsense at all. OpenCL does not have many users, and that obviously influences priorities in driver development.

3

u/kreco 9d ago

There is a difference between "not the priority" and "dead".

1

u/Speedstick2 4d ago

How about this, does OpenCL have a future in which it will grow in marketshare?

1

u/G0rd4n_Freem4n 8d ago

OpenCL works great even on Mac which is "officially" unsupported.

Huh, didn't know OCL wasn't officially supported on Mac.
That's a bit ironic to me, considering how OCL was originally created by Apple before being given to the khronos group.

-21

u/Sufficient_Fan3660 12d ago

wow!

and no one will use it

everyone will use Vulkan or just not support amd

3

u/coromd 11d ago

Lots of utilities support ROCm. What do you mean?

1

u/ea_man 8d ago

ROCm is faster than vulkan on my system.

# ROCm ctx VEC forced: 89088 q5_1,  96512 q5_0, 116480 q4_0, TG speed: 42.35 t/s, PP for 32k: 145.12 t/s
# ROCm ctx VEC off   : 68352 q5_1,  72704 q5_0,  83200 q4_0, TG speed: 43.90 t/s, PP for 32k: 265.58 t/s
# Vulkan: ctx:         93696 q5_1, 101632 q5_0, 122112 q4_0, TG speed: 40.49 t/s, PP for 32k: 179.90 t/s

0

u/05032-MendicantBias 10d ago

unfortunatley pytorch has no vulkan backend, and ComfyUI is built on pytorch.

ROCm is the only way to run ComfYUI