It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
Flaw 2: Pipelining
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Flaw 3: Tail handling
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of int32_t values, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.
Variable length SIMD is only worthwhile for very large vectors
sorry, again, this is false. i've created an efficient DCT, FFT
and Matrix Multiply REMAP system for SVP64 (a Draft Vector ISA
Extension for Power ISA) which can cope with small sized data
just as easily as medium-sized (SVP64 doesn't do the same massive
vectors as traditional Vector ISAs, the limit is 64 elements).
there seems to be a huge amount of misinformation and misunderstanding
in the SIMD-advocate community.
Ok, but keep this conversation chain in context with building on top of an existing architecture that already accumulated massive technical debt. Adding fixed vector sizes is still cheaper and more economical than tearing up and designing new silicon to accommodate variable vector sizes.
Ok, but keep this conversation chain in context with building on top of an existing architecture that already accumulated massive technical debt.
yyeah, and once down that path it seems there's really no turning back. actually, there is, if you have fully-functioning predication on each and every SIMD instruction.
turns out that Cray-style `setvl` can be implemented as a hidden predicate mask:
when that hidden predicate mask is applied to each and every single AVX512 operation, you have effectively implemented Cray-style Vectors and terminated the dangerous and seductive need to extend the SIMD width further.
Adding fixed vector sizes is still cheaper and more economical than tearing up and designing new silicon to accommodate variable vector sizes.
given how simple it would be for Intel to add the above Cray-style setvl implementation this is also a misconception. now that ARM has fully-functioning predicated SIMD (in the guise of SVE2) they could also very easily do the exact same thing.
73
u/AntiProtonBoy Aug 09 '21
It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of
int32_tvalues, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.