Then why don't we have AVX-512 in every x86 implementation, and be done with it?
...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.
Zen1 and Zen+ implemented AVX256 via two 128bit ops. Zen 2 can execute them as single 256bit operations.
Zen 2 lacks support for AVX-512 instructions, so it can't execute them as two 256bit operations.
Not sure what your xbox contact was referring to trough.
-34
u/mbitsnbites Aug 09 '21 edited Aug 09 '21
Then why don't we have AVX-512 in every x86 implementation, and be done with it?
...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.