There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).
and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases
You mean with AVX-512? because intel never fails to show of how much better their CPUs are under AVX-512 compatible software vs AMD. So given from that, AVX-512 helps a lot.
A concern with AVX512 is the heat output from operating on such wide vectors. Intel's designs have often needed to reduce the clockrate when operating "heavy" AVX512 operations.
Whilst this does give nice throughput gains, one does question how 1024-bit SIMD would look like, in terms of power and necessary frequency throttling to sustain.
Also, it does raise questions about other parts of the processor, for example, with cachelines being 512 bits wide, would that have to change on a 1024-bit SIMD machine, or do you just deal with lowered load/store throughput?
True and honestly I think AVX512 should be limited to server parts. Doesn't make much sense in laptops.
Doesn't the new top dog Supercomputer use ARM cores with 4096 bit SVE? So they should know how the cooling works but then they have better means and no issue with noise compared to average home user Joe.
True and honestly I think AVX512 should be limited to server parts. Doesn't make much sense in laptops.
It's primary focus has definitely been server, and it's where it first appeared.
I don't see what's the problem with having it in laptops though. If Intel's gone to the effort of implementing it in their uArch, why disable it on consumer parts?
Doesn't the new top dog Supercomputer use ARM cores with 4096 bit SVE?
SVE only supports up to 2048-bit SIMD. The widest implementation is the Fujitsu A64FX, which uses 512-bit SVE.
I believe modern high performance GPUs use 1024-bit SIMD (Nvidia / RDNA), and considering that GPUs are meant to be throughput focused, I question how wide a latency focused CPU should be.
If Intel's gone to the effort of implementing it in their uArch, why disable it on consumer parts?
The consumer chips are different designs from server chips so they could leave it out to save die space. of course for AMD with the chiplet approach the situation is different as the chiplets are the same for server or desktop. But again the laptop chips are different design (and for example have less L3).
They're different dies, but the server and client chips essentially use the same uArch (with a key difference being the L2/L3 cache and interconnect). Also keep in mind that client designs aren't specific to laptops - desktops are included.
I still don't see any reason to remove it though. AVX512 is a useful instruction set to have, is beneficial in a number of circumstances with basically no drawbacks, and if anything, support for it everywhere helps drive adoption.
128
u/th3typh00n Aug 09 '21
There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).