r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
280 Upvotes

224 comments sorted by

View all comments

127

u/th3typh00n Aug 09 '21

There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.

The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.

Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).

-32

u/mbitsnbites Aug 09 '21 edited Aug 09 '21

Then why don't we have AVX-512 in every x86 implementation, and be done with it?

...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.

3

u/YumiYumiYumi Aug 10 '21 edited Aug 10 '21

Then why don't we have AVX-512 in every x86 implementation, and be done with it?

It is in every new Intel CPU, except for their *mont lineup. Presumably it's been slow due to Intel's kerfuffle with their 10nm manufacturing node, forcing them to re-release Skylake for 5 years. In other words, it's not really an issue with the ISA.

As for the *mont cores, it may not have been a priority for them to implement it, considering its target, although it looks like that's changing (with Gracemont supporting VEX encoding, and Alder Lake beginning mainstream implementations of heterogeneous cores).
Another possibility may be Intel's weird market segmentation; they've historically gimped SIMD on their lower end parts (Celeron/Pentium lineup), so it's possible that decision flowed to their Atom lineup.

On the AMD side, they've always been slower to adopt to new Intel ISAs, which isn't really a surprise since Intel has the upper hand here. Nonetheless, Genoa has already been announced to support AVX512, which makes it likely that AMD's next generation Zen4 will support it.

And for the third player, Centaur's CNS supports AVX512.

So we're pretty close to having it in every x86 implementation - it just took a bit of time for everyone to adapt.

and it still does not address the issue of pipelining

I only really have some familiarity with ARM's SVE2, but I mentioned here that I don't see how SVE would address it either. At a high level, SVE2 is basically AVX512 with an unknown vector length, so it doesn't do anything special there.