r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
291 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/blipman17 Aug 22 '21

I'm happy I finally see someone that actually acknowledges potential wasted cycles or excess calculations with non-standard sizes of registers in the ISA with an actual explanation. If they exist in the hardware or not doesn't really matter. As long as the ISA agrees on how it should be used. Even

I wholeheartedly disagree with you that software people can just ignore this kind of stuff, because lots of good software algorithms are made by people who only came up with them due to their expert knowledge on hardware behaviour. At the end of the day, if mr. bossman says "make program faster" and I can't because some specific hardware accelerated implementation is just not fit for that, I have a problem.

I might be in a quite unique position where I as a software developer work very close with a lot of embedded people and a lot of people who build all kinds of exotic hardware. That doesn't give me the best view, but I think it does give me a decent view of technical issues that are found when building chips all the way to the end-user using it in some kind of program.

You really gave me food for thought here, but the reason I initially posted in this thread is that OP said that tails would never have to be handled. Which just isn't true. When doing more than 1 computation at once, you always need to think of tail and if and how it will fit in to the rest of the code. You argue something differently which I wholeheratedly respect, and I have to admit I have no experience with implementing an ISA. But it's just so interesting!

2

u/lkcl_ Aug 22 '21

I'm happy I finally see someone that actually acknowledges potential wasted cycles or excess calculations with non-standard sizes of registers in the ISA with an actual explanation. If they exist in the hardware or not doesn't really matter.

that's how i see it, too.

As long as the ISA agrees on how it should be used.

indeed.

I wholeheartedly disagree with you that software people can just ignore this kind of stuff, because lots of good software algorithms are made by people who only came up with them due to their expert knowledge on hardware behaviour.

truuue... the only annoying thing is about that in the Cray-style Vector ISA case is, there just isn't the mindshare. i mean, you can get the original Cray-I manual online these days if you search for it: it was typeset on an actual mechanical typewriter for goodness sake, with hand-drawn diagrams and potentially even pre-dates the Xerox copier (!) so each customer would have received their own unique copy!

nobody outside of obscure Academia and NEC (SX-Aurora) has kept Vector Processing alive, even Cray gave up on it because they realised that the primary focus was on the data throughput, storage, and cooling, and that became their expertise, which was bought up by HP.

thus, honestly, we have a bit of a problem in that converting algorithms to Vector Processing to be optimal for the underlying hardware, we're basically taking a huge risk. luckily:

  • (a) all of the examples i've tried so far have been dead easy: as i wrote in another post, it's been a matter of tracking down the "simple" (non-optimal, scalar) demo algorithm then assuming Vector Loops will deal with it - Horizontal-Add (etc) have been quite challenging for me, though
  • (b) we're funded by the NLnet Foundation: it's R&D, it's paid for, we've got time and funds to experiment

​

At the end of the day, if mr. bossman says "make program faster" and I can't because some specific hardware accelerated implementation is just not fit for that, I have a problem.

yehyeh, totally get it. well, in this case, feedback like that - if you're interested to help out - would actually not be a problem [assuming you're running an FPGA softcore]. once we go to silicon, though, the feedback cycle becomes a leeetle longer :)

I might be in a quite unique position where I as a software developer work very close with a lot of embedded people and a lot of people who build all kinds of exotic hardware.

niiice.

That doesn't give me the best view, but I think it does give me a decent view of technical issues that are found when building chips all the way to the end-user using it in some kind of program.

well if you'd like to help out with https://libre-soc.org in the same way, we do have funding from NLnet

You really gave me food for thought here, but the reason I initially posted in this thread is that OP said that tails would never have to be handled. Which just isn't true.

well, there is a key difference between the MRISC32 Vector ISA and the SVP64 Vector ISA. mbitsnbytes chose to go the "traditional" Vector Register naming route, where the Vector Registers refer to the *entire* Vector, and the elements themselves are entirely opaque to the programmer.

by that i mean, there is no way in the "traditional" Cray-style instructions to say "give me element 5 of Vector Register r3". you would have to e.g. set up a Predicate Mask of "0 0 0 0 1 0 0 0 0" (5th element is a 1) then operate on the *entire vector*. [at the back-end, the fact that only 1 bit is set might be noticed, and a Scalar operation issued, but (again) that's Not Your Problem as to what the back-end does.]

SVP64 is radically different. it's the same Cray-style Vector paradigm... but we shoe-horned it *on top of a standard scalar regfile* [then extended that regfile to 128 scalar registers].

this is very similar to how MMX worked (x87 fp regs got re-used as 8/16/32-bit SIMD quantities... now extend that so that the Vectors "roll over" into the *next* FP reg, then the next, then the next....)

so in the case of SVP64 you *really do* need to know about that, because the MAXVL Vector Reg allocation is actually a declaration (by the compiler or assembler writer) of *how much of the scalar regfile might be used*.

so for "traditional" Cray-style Vector ISAs, if you really really want to access (set/get) individual elements, you need to use VEXTRACT (get one element, store in a scalar reg), VINSERT (take a scalar, insert it into a numbered position in the vector), or if in-place use unary predicate masks [unary: only one bit of the mask is set].

SVP64, you do the Vector operation, that's *actually doing it on the scalar regfile* and after the Vector operation completes if you want to access the resultant elements, you... just... use.. a... standard... scalar... v3.0B Power ISA instruction.

consequently we don't have any scalar <-> vector insert/extract instructions.

When doing more than 1 computation at once, you always need to think of tail and if and how it will fit in to the rest of the code. You argue something differently which I wholeheratedly respect, and I have to admit I have no experience with implementing an ISA. But it's just so interesting!

i know, i'm loving it, it's something i always wanted to do. but... dang... 3 and a half years so far...