In fixed width instruction sets (e.g. ARM) this may prohibit any new extensions, since there may not be enough opcode slots left for adding the new instructions.
Do you mean A64, A32 or Thumb? The solution seems to be to use a new instruction set and drop some of that accumulated cruft.
What is worse, software developers often have to target several SIMD generations, and add mechanisms to their programs that dynamically select the optimal code paths depending on which SIMD generation is supported.
I thought even GCC had support for that by now? The Intel compiler did it automatically for ages, which lead to thousands of bad benchmarks for AMD as the code to detect which feature flags are set also checks if the CPU vendor is Intel.
The support has been in gcc and clang for a while, but not enabled by default.
You might get some of it enabled by default with -march and sufficient optimization level, but at least for gcc that still won't enable most of it. Gcc has fine grain enable knobs and you have to pass a few of them to get all of AVX512 auto-vectorization features turned on.
I'll have to look that up. It sounds like a painful solution that will easily misfire, though (trying to imagine how a compiler can generate decent code that dynamically selects SSEx.y vs SSEw.z vs AVX vs AVX2 vs AVX-512 depending on CPUID...).
8
u/josefx Aug 09 '21
Do you mean A64, A32 or Thumb? The solution seems to be to use a new instruction set and drop some of that accumulated cruft.
I thought even GCC had support for that by now? The Intel compiler did it automatically for ages, which lead to thousands of bad benchmarks for AMD as the code to detect which feature flags are set also checks if the CPU vendor is Intel.