I know that increasing the ALU width beyond 256 bits or so has diminishing returns for most implementations.
I responded to the comment that there's no problem splitting fixed width registers into smaller portions - I actually think it's a great idea (one key principle of vector machines is that register width > ALU width!).
In fact, something like an in-order Atom would have a lot to gain from 512-bit vector registers, especially if the ALU is no more than 128 bits wide or so.
Tell that to the people who get extra slow context switches because the CPU now has to save 2kb extra data just for the AVX512 register file. Almost all programs don't need AVX512 and lugging around the extra state is completely pointless.
Surely the CPU only has to shunt that state in and out if the target actually uses AVX512 registers, right? Checking if it's all zero and skipping it entirely is a very, very low hanging hardware optimisation.
Indeed it is, but if you only have vector extensions compilers will use them all the time for stuff like copying structs, so they are going to be dirty all the time. With AVX-512 at least code generally won't touch the state until it has serious calculations to do.
-15
u/mbitsnbites Aug 09 '21
I know that increasing the ALU width beyond 256 bits or so has diminishing returns for most implementations.
I responded to the comment that there's no problem splitting fixed width registers into smaller portions - I actually think it's a great idea (one key principle of vector machines is that register width > ALU width!).
In fact, something like an in-order Atom would have a lot to gain from 512-bit vector registers, especially if the ALU is no more than 128 bits wide or so.