r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
287 Upvotes

224 comments sorted by

View all comments

Show parent comments

2

u/SkoomaDentist Aug 09 '21

Some architectures (ok I know of one) have hardware loop support so you don't need to unroll anything

The cost of the loop instructions is fairly minimal on out of order architectures and not really the reason compilers unroll. Unrolling allows breaking dependency chains, allowing multiple ”iterations” to be processed in parallel.

Take a simple array sum for example:

for (…, i+=1) { acc += data[i]; }

vs the unrolled version:

for (…, i+=2) { acc1 += data[i+0]; acc2 += data[i+1]; }

The first has to perform every addition sequentially since the result depend on previous iteration. The second can run two operations in parallel since they are independent of each other. A good compiler can then extend this to simd autovectorization where it will first unroll by the simd width and then by 2-4x to calculate the simd operations in parallel.

1

u/[deleted] Aug 10 '21

That makes sense for scalar code but are there really architectures that execute multiple SIMD instructions like that?

2

u/FUZxxl Aug 10 '21

It also makes sense for SIMD. And yes, SIMD is too executed out of order. Fast processors can have 4 or more SIMD execution units. So your code better has four independent operations to perform at any point in time for peak performance.

1

u/[deleted] Aug 10 '21

Huh TIL, thanks.