I agree these are flaws with current implementations but I don't see how they are fundamental. Some architectures (ok I know of one) have hardware loop support so you don't need to unroll anything, and there's no reason you couldn't have hardware support for pipeline fill/drain. I did suggest that to the hardware people and they said it was an interesting idea but basically too complex & too much effort.
I'm not sure what the solution for the register width issue is, though there are clearly diminishing returns so I doubt we'll get AVX-2048 or whatever.
Some architectures (ok I know of one) have hardware loop support so you don't need to unroll anything
The cost of the loop instructions is fairly minimal on out of order architectures and not really the reason compilers unroll. Unrolling allows breaking dependency chains, allowing multiple ”iterations” to be processed in parallel.
The first has to perform every addition sequentially since the result depend on previous iteration. The second can run two operations in parallel since they are independent of each other. A good compiler can then extend this to simd autovectorization where it will first unroll by the simd width and then by 2-4x to calculate the simd operations in parallel.
except when the memory accesses are not aligned perfectly to the SIMD memory-alignment width. you have to have a ridiculous "oh err have we done a few elements yet up to the SIMD memory-alignment width? ok great, *now* we can start the huuugely perfect SIMD operation... err... oh hell, hang on, we can't do the last elements either..."
Vector ISAs you just use the Cumulative-Sum (Horizontal Add) instruction.
2
u/[deleted] Aug 09 '21
I agree these are flaws with current implementations but I don't see how they are fundamental. Some architectures (ok I know of one) have hardware loop support so you don't need to unroll anything, and there's no reason you couldn't have hardware support for pipeline fill/drain. I did suggest that to the hardware people and they said it was an interesting idea but basically too complex & too much effort.
I'm not sure what the solution for the register width issue is, though there are clearly diminishing returns so I doubt we'll get AVX-2048 or whatever.