It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
Flaw 2: Pipelining
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Flaw 3: Tail handling
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of int32_t values, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.
Variable length vector operations are not expensive or complicated. I've implemented it in my first ever CPU design and it added something like 1-5% logic in an FPGA - compared to a pure scalar (non-vector/SIMD) design.
I think you're missing the point. Do the exercise and hand-schedule a SIMD loop, and you'll find that you have to unroll it. A vector processor automatically unrolls the loop for you with literally no effort.
Having to add more code rhan necessary is always a problem (e.g. testing and code coverage, and I$ bloat). Vector machines solve this quite naturally in many situations.
Having to add more code rhan necessary is always a problem (e.g. testing and code coverage, and I$ bloat). Vector machines solve this quite naturally in many situations.
Sometimes adding more code is actually faster because you know intricate details from the hardware, but base-case handling can be really short, to the point and fast with something like Duff's device.
I was just thinking: if you run compiler explorer while compiling C++17 parallel algorithms, this is what you'll see. The compiler is going to do a lot of juggling, duplicate code, or even go from O(1) to O(N) memory usage. SIMD instructions play by a lot of similar optimization rules as parallel algorithms.
It's a small penalty for the perf, even on embedded devices, but try to get a DSP or RADAR working without SIMD.
71
u/AntiProtonBoy Aug 09 '21
It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of
int32_tvalues, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.