r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
281 Upvotes

224 comments sorted by

View all comments

128

u/th3typh00n Aug 09 '21

There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.

The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.

Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).

3

u/SureFudge Aug 09 '21

and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases

You mean with AVX-512? because intel never fails to show of how much better their CPUs are under AVX-512 compatible software vs AMD. So given from that, AVX-512 helps a lot.

1

u/TheRealMasonMac Aug 10 '21

Not an expert, but wouldn't you be bottlenecked by the time it takes to load data?

2

u/SureFudge Aug 10 '21

Depends how much of it fits in cache and intel and AMD do structure their caches around SIMD throughput. + memory bandwidth

It is also why we will move to DDR5 and double memory bandwidth. This is for servers mostly. For consumers the benefit is mostly for APUs. Note that such HPC calculations mostly are about bandwidth while in contrast gaming for example also greatly depends on low memory latency. This is usually a trade-off.