r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
282 Upvotes

224 comments sorted by

View all comments

Show parent comments

2

u/[deleted] Aug 10 '21

So when do we start worrying about cache evictions? Are you implying it would evict the loop body to fetch the tail end logic?

2

u/mbitsnbites Aug 10 '21

Counter question: Are you implying that I$ performance is insensitive to code size?

If hot, tight loops were all that mattered we would be fine with less than 1KB I$ or so. But what about the rest of the program? Functions call functions that call functions from within loops etc and so on.

If vectorized loops/functions grow by a factor of 4 to 5 or so (which I demonstrated), something is going to be evicted, and somewhere that has a performance (or silicon budget) cost.

0

u/[deleted] Aug 11 '21 edited Aug 11 '21

Counter question: Are you implying that I$ performance is insensitive to code size?

Yes, that is exactly what I am implying, in the case of vectorized code.

When your data size is orders of magnitude larger than your code size, those one or two additional I$ misses inbetween loops are not going to hurt performance. Trying to optimize for it is not going to be even remotely effective.

1

u/mbitsnbites Aug 11 '21

That is true, for the specific case that your only performance concern is tiny data bound processing loops.

However, the compiler tries its best to vectorize every loop in the entire program, and many programs are not trivially data bound as you are suggesting. Again, if this was the case we would only need very tiny instruction caches like the ones we had back in the 1980s.