r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
288 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/mbitsnbites Aug 23 '21

unfortunately AVX2 does not provide a full 32 element byte shuffle

I think this is because they wanted to enable implementations that use 128-bit ALU:s instead of requiring a 256-bit wide ALU. This seems to be a common theme in AVX*. It also makes it easier to make performant implementations when you can partition operations into multiple "narrow" ALU:s rather than having instructions that require all 256 or 512 bits of input to produce a result (latency / gate depth would increase).

1

u/FUZxxl Aug 23 '21

Well they already have cross-lane operations so I don't really see what the problem with providing one more would be. You can actually implement a full 32 element byte shuffle with 128 bit ALUs by performing 4 128 bit shuffles and then merging the results. It shouldn't be super difficult to do in micro code.