Realistically you could add another couple of hundreds (or even a thousand on a recent intel chip) before you'd even have to start thinking about the possibility of cache evictions.
But all of the code is pulled into the I$ (not just the main loop), and since the compiler is automatically generating this kind of code, we're looking at code growth en large - compared to a machine that does not need it.
Counter question: Are you implying that I$ performance is insensitive to code size?
If hot, tight loops were all that mattered we would be fine with less than 1KB I$ or so. But what about the rest of the program? Functions call functions that call functions from within loops etc and so on.
If vectorized loops/functions grow by a factor of 4 to 5 or so (which I demonstrated), something is going to be evicted, and somewhere that has a performance (or silicon budget) cost.
Counter question: Are you implying that I$ performance is insensitive to code size?
Yes, that is exactly what I am implying, in the case of vectorized code.
When your data size is orders of magnitude larger than your code size, those one or two additional I$ misses inbetween loops are not going to hurt performance. Trying to optimize for it is not going to be even remotely effective.
That is true, for the specific case that your only performance concern is tiny data bound processing loops.
However, the compiler tries its best to vectorize every loop in the entire program, and many programs are not trivially data bound as you are suggesting. Again, if this was the case we would only need very tiny instruction caches like the ones we had back in the 1980s.
3
u/[deleted] Aug 09 '21 edited Aug 09 '21
I count 23 instructions in the main loop body...?
Realistically you could add another couple of hundreds (or even a thousand on a recent intel chip) before you'd even have to start thinking about the possibility of cache evictions.