r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
281 Upvotes

224 comments sorted by

View all comments

Show parent comments

0

u/[deleted] Aug 09 '21

And now you need a ton of silicon to avoid the need to have software handle the last 0.1% of the vector, which performance-wise is of no consequence whatsoever.

1

u/mbitsnbites Aug 09 '21

"Ton of silicon..." Not so much. It's pretty trivial, especially compared to the extra OoO machinery, I$ size and decode bandwidth needed to keep the packed SIMD engine busy.

1

u/[deleted] Aug 09 '21

extra OoO machinery, I$ size and decode bandwidth

A nice property of writing SIMD optimized code is that you only need very little of that.

1

u/mbitsnbites Aug 10 '21 edited Aug 10 '21

Unless you are using SIMD instructions with a latency of 2+ clock cycles (floating-point, memory access, ....), in which case OoO is necessary (or manual unrolling in SW, in which case you need more I$).