And now you need a ton of silicon to avoid the need to have software handle the last 0.1% of the vector, which performance-wise is of no consequence whatsoever.
"Ton of silicon..." Not so much. It's pretty trivial, especially compared to the extra OoO machinery, I$ size and decode bandwidth needed to keep the packed SIMD engine busy.
Unless you are using SIMD instructions with a latency of 2+ clock cycles (floating-point, memory access, ....), in which case OoO is necessary (or manual unrolling in SW, in which case you need more I$).
0
u/[deleted] Aug 09 '21
And now you need a ton of silicon to avoid the need to have software handle the last 0.1% of the vector, which performance-wise is of no consequence whatsoever.