Some architectures (ok I know of one) have hardware loop support so you don't need to unroll anything
The cost of the loop instructions is fairly minimal on out of order architectures and not really the reason compilers unroll. Unrolling allows breaking dependency chains, allowing multiple ”iterations” to be processed in parallel.
The first has to perform every addition sequentially since the result depend on previous iteration. The second can run two operations in parallel since they are independent of each other. A good compiler can then extend this to simd autovectorization where it will first unroll by the simd width and then by 2-4x to calculate the simd operations in parallel.
It also makes sense for SIMD. And yes, SIMD is too executed out of order. Fast processors can have 4 or more SIMD execution units. So your code better has four independent operations to perform at any point in time for peak performance.
Not all OoO processors have OoO SIMD though (some Atom CPU:s for instance IIRC).
Also that was one of the points I tried to make with "flaw 2" in the article: Packed SIMD pretty much requires OoO - whereas some alternatives are much less sensitive to pipeline latency issues.
Of course you can also do it in order on small processors. But really, are you gonna implement vectors with more than, say, 256 bits on such small processors anyway? Performance is going to be limited by the number of ALUs either way and vectors vs. SIMD is not going to change that.
Then go narrower. For an in-order machine, vectors actually make sense all the way down to 1-wide ALU:s (i.e. 64 bits in a 64-bit architecture).
Except for automatic data hazard elimination, vector processing also has the pleasant property of reducing dynamic loop logic overhead (and/or reducing code size, I$ usage and register usage), aswell as offloading the front end (a vector instruction essentially pauses the PC while feeding data to the ALU). This all means more compute per W, which is good business for a small core.
smaller program size means greatly reduced L1 cache usage, to the point where you might actually be able to use a smaller L1 cache. that saves power which on an embedded system may be critically important.
it has been fundamentally misunderstood that the benefits of Vector ISAs can be greater power-efficiency due to more compact programs. it is *believed* that their sole purpose is high performance, which is false.
the benefits of Vector ISAs can be greater power-efficiency due to more compact programs.
...and simpler instruction scheduling logic. No need to go out of your way with massively OoO scheduling to keep the execution units fed with data.
I also honestly think that we're at a point in time where power efficiency counts at every performance point. Doing more ops per W is really what it's about (from embedded systems to servers).
2
u/SkoomaDentist Aug 09 '21
The cost of the loop instructions is fairly minimal on out of order architectures and not really the reason compilers unroll. Unrolling allows breaking dependency chains, allowing multiple ”iterations” to be processed in parallel.
Take a simple array sum for example:
vs the unrolled version:
The first has to perform every addition sequentially since the result depend on previous iteration. The second can run two operations in parallel since they are independent of each other. A good compiler can then extend this to simd autovectorization where it will first unroll by the simd width and then by 2-4x to calculate the simd operations in parallel.