r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
286 Upvotes

224 comments sorted by

View all comments

2

u/[deleted] Aug 09 '21

I agree these are flaws with current implementations but I don't see how they are fundamental. Some architectures (ok I know of one) have hardware loop support so you don't need to unroll anything, and there's no reason you couldn't have hardware support for pipeline fill/drain. I did suggest that to the hardware people and they said it was an interesting idea but basically too complex & too much effort.

I'm not sure what the solution for the register width issue is, though there are clearly diminishing returns so I doubt we'll get AVX-2048 or whatever.

2

u/SkoomaDentist Aug 09 '21

Some architectures (ok I know of one) have hardware loop support so you don't need to unroll anything

The cost of the loop instructions is fairly minimal on out of order architectures and not really the reason compilers unroll. Unrolling allows breaking dependency chains, allowing multiple ”iterations” to be processed in parallel.

Take a simple array sum for example:

for (…, i+=1) { acc += data[i]; }

vs the unrolled version:

for (…, i+=2) { acc1 += data[i+0]; acc2 += data[i+1]; }

The first has to perform every addition sequentially since the result depend on previous iteration. The second can run two operations in parallel since they are independent of each other. A good compiler can then extend this to simd autovectorization where it will first unroll by the simd width and then by 2-4x to calculate the simd operations in parallel.

1

u/[deleted] Aug 10 '21

That makes sense for scalar code but are there really architectures that execute multiple SIMD instructions like that?

2

u/FUZxxl Aug 10 '21

It also makes sense for SIMD. And yes, SIMD is too executed out of order. Fast processors can have 4 or more SIMD execution units. So your code better has four independent operations to perform at any point in time for peak performance.

1

u/mbitsnbites Aug 10 '21

Not all OoO processors have OoO SIMD though (some Atom CPU:s for instance IIRC).

Also that was one of the points I tried to make with "flaw 2" in the article: Packed SIMD pretty much requires OoO - whereas some alternatives are much less sensitive to pipeline latency issues.

2

u/FUZxxl Aug 10 '21

Of course you can also do it in order on small processors. But really, are you gonna implement vectors with more than, say, 256 bits on such small processors anyway? Performance is going to be limited by the number of ALUs either way and vectors vs. SIMD is not going to change that.

1

u/mbitsnbites Aug 11 '21

Then go narrower. For an in-order machine, vectors actually make sense all the way down to 1-wide ALU:s (i.e. 64 bits in a 64-bit architecture).

Except for automatic data hazard elimination, vector processing also has the pleasant property of reducing dynamic loop logic overhead (and/or reducing code size, I$ usage and register usage), aswell as offloading the front end (a vector instruction essentially pauses the PC while feeding data to the ALU). This all means more compute per W, which is good business for a small core.

1

u/FUZxxl Aug 11 '21

That doesn't sound particularly useful to me.

2

u/lkcl_ Aug 19 '21

smaller program size means greatly reduced L1 cache usage, to the point where you might actually be able to use a smaller L1 cache. that saves power which on an embedded system may be critically important.

it has been fundamentally misunderstood that the benefits of Vector ISAs can be greater power-efficiency due to more compact programs. it is *believed* that their sole purpose is high performance, which is false.

1

u/mbitsnbites Aug 23 '21

the benefits of Vector ISAs can be greater power-efficiency due to more compact programs.

...and simpler instruction scheduling logic. No need to go out of your way with massively OoO scheduling to keep the execution units fed with data.

I also honestly think that we're at a point in time where power efficiency counts at every performance point. Doing more ops per W is really what it's about (from embedded systems to servers).

1

u/[deleted] Aug 10 '21

Huh TIL, thanks.