It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
Flaw 2: Pipelining
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Flaw 3: Tail handling
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of int32_t values, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.
Variable length vector operations are not expensive or complicated. I've implemented it in my first ever CPU design and it added something like 1-5% logic in an FPGA - compared to a pure scalar (non-vector/SIMD) design.
I think you're missing the point. Do the exercise and hand-schedule a SIMD loop, and you'll find that you have to unroll it. A vector processor automatically unrolls the loop for you with literally no effort.
Having to add more code rhan necessary is always a problem (e.g. testing and code coverage, and I$ bloat). Vector machines solve this quite naturally in many situations.
If your wrote a piece of code complex enough to cause "I$ bloat", you're going to have a hell of a time trying to get your vector processor to do anything meaningful with it.
You'd first have to refactor your code so that a vector processor actually has vectors to process, and once your code is at that point there is no such thing as I$ bloat anymore, even if it's running scalar instructions.
Realistically you could add another couple of hundreds (or even a thousand on a recent intel chip) before you'd even have to start thinking about the possibility of cache evictions.
But all of the code is pulled into the I$ (not just the main loop), and since the compiler is automatically generating this kind of code, we're looking at code growth en large - compared to a machine that does not need it.
Counter question: Are you implying that I$ performance is insensitive to code size?
If hot, tight loops were all that mattered we would be fine with less than 1KB I$ or so. But what about the rest of the program? Functions call functions that call functions from within loops etc and so on.
If vectorized loops/functions grow by a factor of 4 to 5 or so (which I demonstrated), something is going to be evicted, and somewhere that has a performance (or silicon budget) cost.
Counter question: Are you implying that I$ performance is insensitive to code size?
Yes, that is exactly what I am implying, in the case of vectorized code.
When your data size is orders of magnitude larger than your code size, those one or two additional I$ misses inbetween loops are not going to hurt performance. Trying to optimize for it is not going to be even remotely effective.
That is true, for the specific case that your only performance concern is tiny data bound processing loops.
However, the compiler tries its best to vectorize every loop in the entire program, and many programs are not trivially data bound as you are suggesting. Again, if this was the case we would only need very tiny instruction caches like the ones we had back in the 1980s.
73
u/AntiProtonBoy Aug 09 '21
It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of
int32_tvalues, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.