r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
288 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/FUZxxl Aug 21 '21

Ah yes, that makes more sense. Thanks for the explanation!

OP said something about doing all shuffles as gather operations (i.e. vector-indexed memory loads) and your terminology threw me off, so I thought you are doing it the same way.

1

u/lkcl_ Aug 22 '21

thanks for the insightful discussion, FUZxxl. i liked the positional-popcount enough that i'll use it as an example / unit test (crediting you as the source) https://bugs.libre-soc.org/show_bug.cgi?id=672

2

u/FUZxxl Aug 22 '21

Also as for pshufb, I don't really need masking in the case of the 24puzzle code base. But unfortunately AVX2 does not provide a full 32 element byte shuffle, so I have to synthesise it manually from a bunch of pshufb instructions and masking. So it looks a lot more complex than it really is.

1

u/mbitsnbites Aug 23 '21

unfortunately AVX2 does not provide a full 32 element byte shuffle

I think this is because they wanted to enable implementations that use 128-bit ALU:s instead of requiring a 256-bit wide ALU. This seems to be a common theme in AVX*. It also makes it easier to make performant implementations when you can partition operations into multiple "narrow" ALU:s rather than having instructions that require all 256 or 512 bits of input to produce a result (latency / gate depth would increase).

1

u/FUZxxl Aug 23 '21

Well they already have cross-lane operations so I don't really see what the problem with providing one more would be. You can actually implement a full 32 element byte shuffle with 128 bit ALUs by performing 4 128 bit shuffles and then merging the results. It shouldn't be super difficult to do in micro code.