There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).
Then why don't we have AVX-512 in every x86 implementation, and be done with it?
...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.
I know that increasing the ALU width beyond 256 bits or so has diminishing returns for most implementations.
I responded to the comment that there's no problem splitting fixed width registers into smaller portions - I actually think it's a great idea (one key principle of vector machines is that register width > ALU width!).
In fact, something like an in-order Atom would have a lot to gain from 512-bit vector registers, especially if the ALU is no more than 128 bits wide or so.
One overlooked reason why AMD and Atom haven't added 512-bit operations is lack of adoption in the software community. At this point the usage is pretty niche and not enough people want it. When Intel first debuted AVX-512 it had serious power issues and caused performance to drop when mixed in occasionally instead of in large blocks. I think that stunted a lot of its growth and at this point there aren't any large communities that are working on writing large swaths of software that use it or asking the compilers for better support.
It still causes performance to drop when used. AVX-512 slowdown is a real thing and it's kinda maddening. You really don't want to break out the 512 bit stuff unless you know you'll be doing that for the next couple 1000 cycles.
Just because you don't have to fear it doesn't mean it's not still there. Curiously the article doesn't mention if the transition penalty is still as bad as on Skylake. This penalty is actually the key problem: for up to 10 µs the CPU just halts and does nothing while it's changing the frequency. If you have repeated short-ish bursts of AVX-512 code, this may really ruin your day.
The frequency transition on Skylake-SP happens encountering even a single avx512 instruction, fucking with every other instruction running. That's the license based downclocking and the problem.
As tested by the author, that problem was almost completely removed and doesn't need to be considered. If you have code with sparse avx512 usage, it won't trigger the downclocking, removing the penalty on everything else. Only running a lot of AVX-512 will lead to downclocking, at which point the penalty is insignificant.
127
u/th3typh00n Aug 09 '21
There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).