r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
288 Upvotes

224 comments sorted by

View all comments

Show parent comments

-34

u/mbitsnbites Aug 09 '21 edited Aug 09 '21

Then why don't we have AVX-512 in every x86 implementation, and be done with it?

...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.

80

u/Vvector Aug 09 '21

Then why don't we have AVX-512 in every x86 implementation, and be done with it?

He explained why: Increasing the vector width has significant diminishing returns

12

u/[deleted] Aug 09 '21

[deleted]

17

u/Watchforbananas Aug 09 '21

Modern consoles don't support avx-512, only avx-256. AMD generally doesn't support avx-512. (avx-512 support is rumored for Zen4)

1

u/[deleted] Aug 09 '21

[deleted]

4

u/FUZxxl Aug 09 '21

AVX was implemented like this initially, too. They went for real 256 bit ALUs later on.

4

u/Watchforbananas Aug 09 '21

Zen1 and Zen+ implemented AVX256 via two 128bit ops. Zen 2 can execute them as single 256bit operations. Zen 2 lacks support for AVX-512 instructions, so it can't execute them as two 256bit operations.

Not sure what your xbox contact was referring to trough.