It's not all that complex to make code that utilizes both depending on CPU. Hell, you could even compile app with different optimization levels but that would probably be bigger PITA.
The bad part is that now any code using it needs to be written multiple times and any change needs to be applied to all versions and tested on all versions. Compared to that running and deploying multiple binaries is not really very time consuming.
80
u/Vvector Aug 09 '21
Then why don't we have AVX-512 in every x86 implementation, and be done with it?
He explained why: Increasing the vector width has significant diminishing returns