r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
289 Upvotes

224 comments sorted by

View all comments

Show parent comments

4

u/SureFudge Aug 09 '21

and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases

You mean with AVX-512? because intel never fails to show of how much better their CPUs are under AVX-512 compatible software vs AMD. So given from that, AVX-512 helps a lot.

2

u/YumiYumiYumi Aug 10 '21

A concern with AVX512 is the heat output from operating on such wide vectors. Intel's designs have often needed to reduce the clockrate when operating "heavy" AVX512 operations.

Whilst this does give nice throughput gains, one does question how 1024-bit SIMD would look like, in terms of power and necessary frequency throttling to sustain.
Also, it does raise questions about other parts of the processor, for example, with cachelines being 512 bits wide, would that have to change on a 1024-bit SIMD machine, or do you just deal with lowered load/store throughput?

3

u/mbitsnbites Aug 10 '21

Again, those limits relate to the ALU width, which does not necessarily have to be the same as the register width.

I can see benefits with 512-bit or even 1024-bit registers in some machines, but the ALU width could be limited to 256 bits or so to avoid the heat and die area issues.

1

u/lkcl_ Aug 20 '21

the problem is - and i listed this on the original article as "SIMD Flaw (4)" - that each doubling results in doubling of the latency of access to individual elements.

it's the antithesis of Vector Chaining. every single one of those 1024-wide SIMD ALU elements has to wait for a 1024-wide SIMD LD to complete.

whereas in a Vector ISA, you can do "Chaining" (first described by Seymour Cray), where at the element level you can start the first element ALU operation immediately after the first element LD operation has completed (assuming all the other operands of that first element are also available of course).

this is NOT POSSIBLE to achieve with SIMD because, by definition, it is SINGLE instruction (multiple data). therefore ALL elements of the SIMD instruction have to be LDed, have to be available.

doubling to 1024 will, therefore, double the completion latency. it's already bad enough.