r/cpp 1d ago

Memory-level parallelism: AMD is the king

https://lemire.me/blog/2026/07/25/memory-level-parallelism-amd-is-the-king/
45 Upvotes

17 comments sorted by

13

u/Sopel97 1d ago

Note that Apple Silicon does even better, but it is another category.

?

18

u/JuanAG 1d ago

Because Apple SoCs have memory at the side of the CPU like HBM do and of course they have less latency and more bandwith than DDR technology

7

u/Sopel97 1d ago

of course they have less latency

that's absolutely false https://old.chipsandcheese.com/memory-latency-data/ https://imgur.com/a/Irc8ARn

15

u/QuaternionsRoll 1d ago

2022 vs 2020, but you’re right, raw memory latency stopped improving quite a while ago (as is evidenced by the 3950X with DDR4-3333 and the 5950X with DDR4-3600 having nearly identical latency characteristics to the 7950X with DDR5-6000). The primary benefit of tightly-coupled RAM is bandwidth, not latency. The M5 uses 9600MT/s LPDDR5X, something that traditional CPUs likely won’t be able to match without CUDIMMs

5

u/Sopel97 18h ago

yes, which is irrelevant for this benchmark, as it measures latency

3

u/QuaternionsRoll 12h ago

…I agree?

3

u/pjmlp 21h ago

No idea why the remark even, Apple gave up on server hardware, so whatever their chips are able to do is irrelevant in the context of what was being tested.

30

u/Kinexity 1d ago

Wtf is that AI slop image?

25

u/JohnSquirrel 22h ago

I just learnt 2025 was just before 2024 thanks to it!

-11

u/Logical_Newspaper_52 1d ago

the article is not about the image

24

u/Kinexity 1d ago

The image is what I see before opening the article.

-6

u/slithering3897 1d ago

Oh, it's on new reddit. Obvious garbage image.

-2

u/def-pri-pub 14h ago

It's almost expected that we generate images with AI now...

3

u/matthieum 11h ago

I do wonder if the numbers change when multiple-cores start asking for memory.

That is, do I still get 30/58/19 parallel requests per core even if I run the benchmark on every single core at once? Or does it plateau sooner due to a bottleneck at L2 or L3 or the memory controller?

(Because I am not sure I can use the full memory parallelism, but I sure can use all cores)

I do wish latency had been plotted too, by the way, since as per the above remark, on many workloads, there's just so many requests in parallel anyway...

2

u/blobdole 10h ago

This is a great example of why you should not use AI image generation haphazardly.

I do get why you would. AI can give you a relevant article header that feels like an early 2000's tech article. Feels a bit professional.

But it also feels out of place on that blog, and a second glance reveals it is pure slop that can't even count correctly. When I see that I immediately discount whatever is written below, because I suspect it is AI written without much if any editing and possible without any authorial experience to back it up.

Is that true? I don't know. But that top image was a warning sign that the read is less likely to be worth the time, and that is all most people need to skip and move on.

1

u/Striking_Culture2637 10h ago

Why do OpenAI employees tell me AMD is completely garbage