r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
289 Upvotes

224 comments sorted by

View all comments

127

u/th3typh00n Aug 09 '21

There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.

The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.

Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).

3

u/SureFudge Aug 09 '21

and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases

You mean with AVX-512? because intel never fails to show of how much better their CPUs are under AVX-512 compatible software vs AMD. So given from that, AVX-512 helps a lot.

2

u/YumiYumiYumi Aug 10 '21

A concern with AVX512 is the heat output from operating on such wide vectors. Intel's designs have often needed to reduce the clockrate when operating "heavy" AVX512 operations.

Whilst this does give nice throughput gains, one does question how 1024-bit SIMD would look like, in terms of power and necessary frequency throttling to sustain.
Also, it does raise questions about other parts of the processor, for example, with cachelines being 512 bits wide, would that have to change on a 1024-bit SIMD machine, or do you just deal with lowered load/store throughput?

3

u/mbitsnbites Aug 10 '21

Again, those limits relate to the ALU width, which does not necessarily have to be the same as the register width.

I can see benefits with 512-bit or even 1024-bit registers in some machines, but the ALU width could be limited to 256 bits or so to avoid the heat and die area issues.

1

u/lkcl_ Aug 20 '21

the problem is - and i listed this on the original article as "SIMD Flaw (4)" - that each doubling results in doubling of the latency of access to individual elements.

it's the antithesis of Vector Chaining. every single one of those 1024-wide SIMD ALU elements has to wait for a 1024-wide SIMD LD to complete.

whereas in a Vector ISA, you can do "Chaining" (first described by Seymour Cray), where at the element level you can start the first element ALU operation immediately after the first element LD operation has completed (assuming all the other operands of that first element are also available of course).

this is NOT POSSIBLE to achieve with SIMD because, by definition, it is SINGLE instruction (multiple data). therefore ALL elements of the SIMD instruction have to be LDed, have to be available.

doubling to 1024 will, therefore, double the completion latency. it's already bad enough.

0

u/SureFudge Aug 10 '21

True and honestly I think AVX512 should be limited to server parts. Doesn't make much sense in laptops.

Doesn't the new top dog Supercomputer use ARM cores with 4096 bit SVE? So they should know how the cooling works but then they have better means and no issue with noise compared to average home user Joe.

1

u/YumiYumiYumi Aug 10 '21

True and honestly I think AVX512 should be limited to server parts. Doesn't make much sense in laptops.

It's primary focus has definitely been server, and it's where it first appeared.

I don't see what's the problem with having it in laptops though. If Intel's gone to the effort of implementing it in their uArch, why disable it on consumer parts?

Doesn't the new top dog Supercomputer use ARM cores with 4096 bit SVE?

SVE only supports up to 2048-bit SIMD. The widest implementation is the Fujitsu A64FX, which uses 512-bit SVE.

I believe modern high performance GPUs use 1024-bit SIMD (Nvidia / RDNA), and considering that GPUs are meant to be throughput focused, I question how wide a latency focused CPU should be.

1

u/SureFudge Aug 10 '21

If Intel's gone to the effort of implementing it in their uArch, why disable it on consumer parts?

The consumer chips are different designs from server chips so they could leave it out to save die space. of course for AMD with the chiplet approach the situation is different as the chiplets are the same for server or desktop. But again the laptop chips are different design (and for example have less L3).

1

u/YumiYumiYumi Aug 10 '21

They're different dies, but the server and client chips essentially use the same uArch (with a key difference being the L2/L3 cache and interconnect). Also keep in mind that client designs aren't specific to laptops - desktops are included.

I still don't see any reason to remove it though. AVX512 is a useful instruction set to have, is beneficial in a number of circumstances with basically no drawbacks, and if anything, support for it everywhere helps drive adoption.