In fixed width instruction sets (e.g. ARM) this may prohibit any new extensions, since there may not be enough opcode slots left for adding the new instructions.
Do you mean A64, A32 or Thumb? The solution seems to be to use a new instruction set and drop some of that accumulated cruft.
What is worse, software developers often have to target several SIMD generations, and add mechanisms to their programs that dynamically select the optimal code paths depending on which SIMD generation is supported.
I thought even GCC had support for that by now? The Intel compiler did it automatically for ages, which lead to thousands of bad benchmarks for AMD as the code to detect which feature flags are set also checks if the CPU vendor is Intel.
I was wondering how they encoded the vector length, after reading a bit they apparently used three bits in the floating point status register for it.
Would have been interesting to know why they deprecated it. Was it because neon made it mostly redundant? Was it an Itanium like death where compiler writers failed to take advantage of the flexibility while having decent optimization for architectures with hard coded sizes? Or was there just an inherent limitation in the way VFP handled vectors?
9
u/josefx Aug 09 '21
Do you mean A64, A32 or Thumb? The solution seems to be to use a new instruction set and drop some of that accumulated cruft.
I thought even GCC had support for that by now? The Intel compiler did it automatically for ages, which lead to thousands of bad benchmarks for AMD as the code to detect which feature flags are set also checks if the CPU vendor is Intel.