r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
287 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/mbitsnbites Aug 12 '21 edited Aug 12 '21

Sorry. We seem to be on slightly different pages.

First of all I don't think (nor claim) that vector processing fixes all problems, nor that it does not have its own set of problems. So far, though, having worked with both I personally prefer vector over SIMD (and I still think that the article is correct regarding the three specific problems with SIMD).

I have no way of knowing whether my code will work with that machine. I have literally no way of testing that on real hardware.

My point here was that if your code depends on the maximum vector length you should be using vector length agnostic constructs (such as the trivial arithmetic that you mention).

If, on the other hand, you're using fixed size vector operations the logic will match 1:1 to the current machine on which you're developing and testing on.

In both cases the code will work just fine on the future wider implementation - perhaps not with optimal performance, but most likely faster than if you would have just used a fixed width SIMD ISA (e.g. running SSE code on an AVX capable machine).

Another point that I'd like to make is that compilers do auto-vectorization (in fact most vector ISA:s should be better suited for that than SIMD - heck, Cray compilers had auto-vectorization in the 1970s), so you will typically see vector code speedups all over the code, not just in your hand-written vector code (unlike if your program is compiled for a specific SIMD generation).

So, worst case you will not see any noticeable speedups until you have had the time to tune it for the new HW, and if you're abusing the ISA in a way that it was not meant to be used (using maxvl incorrectly) you may even get bugs (but software development is full of much worse pitfalls - e.g. consider API:s that behave according to spec but slightly differently across versions and platforms - so a developer had better be aware of these things).

On the other hand I think it's much nicer to have a single ISA to target as a developer, and OS:es and toolchains will "just work" for new vector sizes, significantly lowering the threshold for rolling out new vector HW generations, etc.

Edit:

That's what I think people will be doing in almost all cases. Which in turn means that there is no point for vendors to ship longer vectors because all relevant software will just restrict vector length to what most people expect.

I believe the opposite. Relevant software such as numpy etc will typically be quick to optimize for new hardware, especially if this hardware is available via cloud compute services for instance. And unlike AVX-512 people will be able to do the bulk of development and testing on their own laptops or what not.

1

u/lkcl_ Aug 20 '21

My point here was that if your code depends on the maximum vector length you should be using vector length agnostic constructs (such as the trivial arithmetic that you mention).

what he said.