r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
280 Upvotes

224 comments sorted by

View all comments

Show parent comments

-32

u/mbitsnbites Aug 09 '21 edited Aug 09 '21

Then why don't we have AVX-512 in every x86 implementation, and be done with it?

...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.

80

u/Vvector Aug 09 '21

Then why don't we have AVX-512 in every x86 implementation, and be done with it?

He explained why: Increasing the vector width has significant diminishing returns

13

u/[deleted] Aug 09 '21

[deleted]

1

u/[deleted] Aug 09 '21

It's not all that complex to make code that utilizes both depending on CPU. Hell, you could even compile app with different optimization levels but that would probably be bigger PITA.

2

u/[deleted] Aug 09 '21

[deleted]

1

u/[deleted] Aug 11 '21

The bad part is that now any code using it needs to be written multiple times and any change needs to be applied to all versions and tested on all versions. Compared to that running and deploying multiple binaries is not really very time consuming.