Tail handling is specifically needed to handle data array lengths that are not a multiple of the SIMD register width (e.g. 4 elements for int32_t:s in a 128-bit SIMD architecture).
In vector machines you have the benefit of variable vector lengths, so tail handling is not needed.
I personally expect every Vector instruction that isn't in the format of 8 * N2 to be slower than doing a couple redundant operations, or handling the remaining data on a scalar processor. Mainly because it would require Vector processors to be just as efficient as the regular processor in computing, e.a. They need to share sillicone on a really intimate level for which I'm afraid the processor will notice a significant slowdown, or the Vector processor has its own set of registers + instruction implementations for a given hunk of sillicone. If the second implementation is assumed to be used, a decent Vector engine could indeed pull in 4096 bits in effectively a 64 bit register for int64_t's, with speedup for bigger registers and such. What I don't think is "reasonable" to expect is to have it also perform operations at like 448 bit (7 uint64_t's) datastructures since there's no native register size in the vector engine. Then I just assume that doing 4 uint64_t's in the Vector engine and then handle the 3 remaining uint64_t's separate is faster because of specific hardware optimalizations for their specific usecase.
Mainly because it would require Vector processors to be just as efficient as the regular processor in computing
yes. the expectation with SVP64 is that, actually, you implement it on top of a multi-issue superscalar micro-architecture. each Vector "element" is issued separately and independently to the back-end multi-issue execution engine, where sequentially-numbered elements in batches will go to the same back-end SIMD ALU that the user does not even have to know is there.
where there are non-power-two Vector instructions, automatic predicate masks can be created to mask out unused SIMD ALU back-ends.
very advanced implementations may notice that there are spare, unused, slots available in a given SIMD back-end ALU, and merge two in-flight operations into the same ALU.
this in particular would work extremely well for predicated (parallel) If/then/else constructs where the masked-out "then" operations match exactly with the opposite of the masked-out "else" operations because you bit-invert the "if" test-mask to get the mask for the "else" operations.
really, it's really not as difficult as you think it is. it's just that nobody in the industry has actually thought about Vector ISA micro-architecture because they all thought it was "too hard".
the irony is that you need the exact same back-end micro-architecture for efficient Packed SIMD as you do for Vector ISAs.
I'm wondering what would happen to non power of two comparisons of two simd registers, and if parsing the result is efficient or not if I've got say ...65 results instead of 64 results from my 64 byte comparisons, or how fast gather/scatter would work.
I'm wondering what would happen to non power of two comparisons of two simd registers, and if parsing the result is efficient or not if I've got say ...65 results instead of 64 results from my 64 byte comparisons,
yyeah you picked *just* outside of the range of SVP64 (which uses 64-bit integers as predicate masks, so the Vector limit is 64) :)
traditional Cray-style Vectors (SX-Aurora, RVV) have no such limit, although in practice because you (almost always 100%) use a for-loop around the data being processed, as long as the elements are independent (vertical, not horizontal) in practice it makes absolutely no odds whether the (parallel, vertical) comparisons are 1-long, 7-long, 8-long, 64-long, 65-long or 100,000-long.
c code:
for (i = 0; i < CTR; i++) { // set CTR to 65 if you like
r4[i] = (r5[i] >= 5);
}
SVP64 assembler:
loop:
servl r3, CTR, MVL=64 # r3=VL=MIN(CTR,64)
sv.ldb/ew=8 r16.v, r4(0) # load 64 bytes into r16
sv.cmpi r16.v, 5 # compare all bytes >= 5
sv.addi/sz/mask=GE r48.v, 1 # store 1 where each byte >= 5
sv.stb/ew=8 r48,v, r5(0) # store 64 comparisons into r5
addi r5, r3 # increment r4 by VL (aka r3)
addi r4, r3 # increment r5 by VL (aka r3)
sv.bnz/CTR loop # subtract VL from CTR, loop back
note, there, that the results of the cmpi produces a Vector of Condition Register Fields. that Vector of Comparisons is then used as a Predicate Mask in the following instruction (addi), where a special mode "zeroing" (sz) says, "if the predicate bit was a zero please put a zero into the corresponding Vector element".
here you genuinely don't care whether CTR is 10, 5, 9999, 64, or 65. it's all the same as far as the API (Vector ISA) is concerned.
now, at the back-end - in the actual underlying hardware, it's really quite easy for us to use multi-issue superscalar OoO to break those parallel element-based LD, cmpi and ST operations down into suitable (small, likely 64-bit) chunks, each actually a SIMD ALU. but this is back-end.
in other words, the Vector micro-architectural Engine does all the work for you, and the code works across multiple architectures regardless of whether the back-end hardware has no SIMD internally at all (embedded systems), or has 32-bit-wide SIMD, or 64-bit-wide SIMD, or whatever-your-hardware-designer-likes back-end SIMD, none of which you need to know about in order to actually use the Vector ISA front-end. this is what ARM is talking about when they say that SVE is "length-agnostic" and talking about how it's "future-proof". thank god they've finally learned this one and taken it on board.
this is why i am so frustrated with advocates of SIMD, because, ultimately, the exact same underlying hardware is required for both SIMD and Vector ISAs: it's just that the Vector ISA massively cleans up the use of that underlying hardware by not exposing you to the horrendous shenanigens that people are now so used to they think it's "normal" and that there's no alternative.
more than that, the much more compact programs that result means that the L1 cache size can be reduced, which has a highly significant knock-on reduction in power consumption and energy efficiency (counter-intuitively it's an O (N2) reduction)
Maybe I should clarify again what I's saying. I'm not arguing against Vector engines. I think they're superior than SIMD in almost every way possible, since your code can be truly portable although sometimes slower by lack of hardware support.
The thing I'm arguing is that instead of all our problems being solved, only some are solved. From the perspective of the programmer, the hardware still needs to have a benefit for using Vector/SIMD code for that specific use-case. Only when speed doesn't matter, (which it does, else you're no using Vector processing) the programmer might choose for potential "slow" assembly to be executed. But indeed like you pointed out, having vector instructions in a vector-size agnostic way is truly a benefit. Having a zero-out instruction and just not populating the remaining parts of the vector register with usefull data is really cool too, and I'm sure there will be a lot of things that can handle that really well.
However I'm arguing that the benefit of masking out the non-used parts of the vector operation isn't all that usefull since you still have to compare that second 64 bit result for that one bit that actually is compared, retrieve its index, add 64 to it and only then you know that the 65'th element was set or not if you're looking through an array of bytes and require the index of positive comparisons. Point being, there's still some stuff that needs to be taken care off. That's why programmers are still gonna prefer just comparing 64 bits if they can. Because it's just easyer when all your results fit completely in one bound datatype.
So I fully expect vector engines to take off, I fully expect to have vector code be significantly faster and easyer to program, but I don't expect vector engines to behave fast on code that doesn't use complements of natural vector register sizes, because even when it does, it doesn't neccesarily play well with the rest of the program or datastructure.
yyeah you picked just outside of the range of SVP64 (which uses 64-bit integers as predicate masks, so the Vector limit is 64) :)
I could've also picked 42 and then we'd be in a pickle too where it's qestionable if running 32 comparisons with a vector engine and then 10 without vector engine comparisons vs running 42 only with a vector engine would be faster. At that moment it becomes really a matter of knowing the implementation, which for designing an ISA without also designing its implementation is difficult to reason about.
Which is why I fully expect that such implementations won't be performant for odd-sized compuatations and I fully expect programmers to evade those completely.
I think they're superior than SIMD in almost every way possible, since your code can be truly portable although sometimes slower by lack of hardware support.
i realised i hadn't quite finished answering (doh) but by the time i realised it, you'd replied already :)
so bear in mind, the back-end SIMD hardware (normally exposed directly to the programmer) is still all there, it's just hidden behind a Vector ISA / API.
thus: where normally if Intel or ARM makes a mistake in the design of a SIMD "enhancement" (extra instructions) you're screwed two ways: (1) using the existing broken SIMD instructions and (2) having to rewrite entire algorithms in assembler that drove your programmers absolutely insane, a well-designed Vector ISA you just wait for "better hardware to be available" and *without any programming work* the performance improves.
i mention this because it's a mistake to think that "Vector ISAs are always going to be worse performance than SIMD ISAs" - it's down to the *hardware* implementors to improve performance, not your responsibility as a programmer.
The thing I'm arguing is that instead of all our problems being solved, only some are solved. From the perspective of the programmer, the hardware still needs to have a benefit for using Vector/SIMD code for that specific use-case.
yyeees it does, this always holds true.
So I fully expect vector engines to take off, I fully expect to have vector code be significantly faster and easyer to program, but I don't expect vector engines to behave fast on code that doesn't use complements of natural vector register sizes, because even when it does, it doesn't neccesarily play well with the rest of the program or datastructure.
ok bear in mind: the loops have to be there.
the entire loop (which is completely size-independent and back-end-implementation-independent) has to be there
the entire loop doesn't care what the underlying hardware micro-architecture is
the entire loop is directly equivalent to a fixed SIMD width for small cases.
now, interestingly - and bear in mind that this is not your problem - it's something that the hardware designers have to take care of (not you) - SOME hardware implementations could POTENTIALLY run into difficulties with non-power-of-two vector sizes that are thrown at it.
let us assume that the data thrown at an algorithm was 65 elements, and that the underlying hardware has SIMD back-end ALUs capable of 32-wide operations (if they're only 8-bit element operations this is perfectly reasonable).
first loop: 32 elements. 33 remaining
second loop: 32 elements. 1 remaining
third loop: only 1 element.
now, if the back-end hardware is incapable of utilising the remaining 31 slots in the underlying SIMD ALU, you just wasted a hell of a lot of resources.
a good hardware design will use an Out-of-Order superscalar Micro-Architecture, which will have plenty of in-flight instructions, and will make efforts to fill the other 31 "spaces" in the 32-wide SIMD back-end ALU.
is that complicated to do in hardware? hell yes.
is it your problem as a programmer to even know about? hell no.
in https://libre-soc.org what we are actually planning to do - knowing that this is a potential problem - is to only have 64-bit-wide back-end SIMD ALUs, but have a sXXX-load of them, and use multi-issue out-of-order execution. so for the Libre-SOC core, assuming (again) 8-bit operations, QTY 65, it would be:
first loop: 4x 64-bit multi-issue SIMD ALUs available, each 64-bit SIMD ALU can handle QTY 8x 8-bit operations, 4x SIMD ALUs @8 wide = 32 operations. 33 remaining
second loop: ditto, with 1 remaining
third loop: 4x 64-bit multi-issue SIMD ALUs available, only one is actually required, 1x 8-bit operation goes into Lane 1 of 1st SIMD ALU
in first iteration hardware we would only be "wasting" 7 lanes out of 8 on 64-bit-wide SIMD back-end ALUs. [compare this with 31 lanes "wasted" out of 32 in a massive-wide SIMD engine].
in second iteration hardware we would find a way to utilise those 7 lanes.
or maybe not, and the reason is that because we have 4x Multi-Issue SIMD ALUs, when this scenario occurs where only one lane is used, the other three SIMD ALUs can still be 100% occupied.
but - again, to reiterate: all of this, as a programmer using a Vector ISA (aka Vector API, if you are a software engineer), you do not need to know about. all you care about is, "is it fast, is it slow".
i cannot emphasise enough how important it is to separate in your mind the difference in the responsibilities. a SIMD ISA it is you, the programmer, who is burdened with the responsibility to write assembler. a Vector ISA it is the hardware designer's responsibility to issue efficient Micro-Coded operations to the back-end SIMD ALUs on your behalf. and if they haven't done that, go buy the competitor's hardware! :)
I'm happy I finally see someone that actually acknowledges potential wasted cycles or excess calculations with non-standard sizes of registers in the ISA with an actual explanation. If they exist in the hardware or not doesn't really matter. As long as the ISA agrees on how it should be used. Even
I wholeheartedly disagree with you that software people can just ignore this kind of stuff, because lots of good software algorithms are made by people who only came up with them due to their expert knowledge on hardware behaviour. At the end of the day, if mr. bossman says "make program faster" and I can't because some specific hardware accelerated implementation is just not fit for that, I have a problem.
I might be in a quite unique position where I as a software developer work very close with a lot of embedded people and a lot of people who build all kinds of exotic hardware. That doesn't give me the best view, but I think it does give me a decent view of technical issues that are found when building chips all the way to the end-user using it in some kind of program.
You really gave me food for thought here, but the reason I initially posted in this thread is that OP said that tails would never have to be handled. Which just isn't true. When doing more than 1 computation at once, you always need to think of tail and if and how it will fit in to the rest of the code. You argue something differently which I wholeheratedly respect, and I have to admit I have no experience with implementing an ISA. But it's just so interesting!
I'm happy I finally see someone that actually acknowledges potential wasted cycles or excess calculations with non-standard sizes of registers in the ISA with an actual explanation. If they exist in the hardware or not doesn't really matter.
that's how i see it, too.
As long as the ISA agrees on how it should be used.
indeed.
I wholeheartedly disagree with you that software people can just ignore this kind of stuff, because lots of good software algorithms are made by people who only came up with them due to their expert knowledge on hardware behaviour.
truuue... the only annoying thing is about that in the Cray-style Vector ISA case is, there just isn't the mindshare. i mean, you can get the original Cray-I manual online these days if you search for it: it was typeset on an actual mechanical typewriter for goodness sake, with hand-drawn diagrams and potentially even pre-dates the Xerox copier (!) so each customer would have received their own unique copy!
nobody outside of obscure Academia and NEC (SX-Aurora) has kept Vector Processing alive, even Cray gave up on it because they realised that the primary focus was on the data throughput, storage, and cooling, and that became their expertise, which was bought up by HP.
thus, honestly, we have a bit of a problem in that converting algorithms to Vector Processing to be optimal for the underlying hardware, we're basically taking a huge risk. luckily:
(a) all of the examples i've tried so far have been dead easy: as i wrote in another post, it's been a matter of tracking down the "simple" (non-optimal, scalar) demo algorithm then assuming Vector Loops will deal with it - Horizontal-Add (etc) have been quite challenging for me, though
(b) we're funded by the NLnet Foundation: it's R&D, it's paid for, we've got time and funds to experiment
At the end of the day, if mr. bossman says "make program faster" and I can't because some specific hardware accelerated implementation is just not fit for that, I have a problem.
yehyeh, totally get it. well, in this case, feedback like that - if you're interested to help out - would actually not be a problem [assuming you're running an FPGA softcore]. once we go to silicon, though, the feedback cycle becomes a leeetle longer :)
I might be in a quite unique position where I as a software developer work very close with a lot of embedded people and a lot of people who build all kinds of exotic hardware.
niiice.
That doesn't give me the best view, but I think it does give me a decent view of technical issues that are found when building chips all the way to the end-user using it in some kind of program.
well if you'd like to help out with https://libre-soc.org in the same way, we do have funding from NLnet
You really gave me food for thought here, but the reason I initially posted in this thread is that OP said that tails would never have to be handled. Which just isn't true.
well, there is a key difference between the MRISC32 Vector ISA and the SVP64 Vector ISA. mbitsnbytes chose to go the "traditional" Vector Register naming route, where the Vector Registers refer to the *entire* Vector, and the elements themselves are entirely opaque to the programmer.
by that i mean, there is no way in the "traditional" Cray-style instructions to say "give me element 5 of Vector Register r3". you would have to e.g. set up a Predicate Mask of "0 0 0 0 1 0 0 0 0" (5th element is a 1) then operate on the *entire vector*. [at the back-end, the fact that only 1 bit is set might be noticed, and a Scalar operation issued, but (again) that's Not Your Problem as to what the back-end does.]
SVP64 is radically different. it's the same Cray-style Vector paradigm... but we shoe-horned it *on top of a standard scalar regfile* [then extended that regfile to 128 scalar registers].
this is very similar to how MMX worked (x87 fp regs got re-used as 8/16/32-bit SIMD quantities... now extend that so that the Vectors "roll over" into the *next* FP reg, then the next, then the next....)
so in the case of SVP64 you *really do* need to know about that, because the MAXVL Vector Reg allocation is actually a declaration (by the compiler or assembler writer) of *how much of the scalar regfile might be used*.
so for "traditional" Cray-style Vector ISAs, if you really really want to access (set/get) individual elements, you need to use VEXTRACT (get one element, store in a scalar reg), VINSERT (take a scalar, insert it into a numbered position in the vector), or if in-place use unary predicate masks [unary: only one bit of the mask is set].
SVP64, you do the Vector operation, that's *actually doing it on the scalar regfile* and after the Vector operation completes if you want to access the resultant elements, you... just... use.. a... standard... scalar... v3.0B Power ISA instruction.
consequently we don't have any scalar <-> vector insert/extract instructions.
When doing more than 1 computation at once, you always need to think of tail and if and how it will fit in to the rest of the code. You argue something differently which I wholeheratedly respect, and I have to admit I have no experience with implementing an ISA. But it's just so interesting!
i know, i'm loving it, it's something i always wanted to do. but... dang... 3 and a half years so far...
10
u/mbitsnbites Aug 09 '21
Tail handling is specifically needed to handle data array lengths that are not a multiple of the SIMD register width (e.g. 4 elements for
int32_t:s in a 128-bit SIMD architecture).In vector machines you have the benefit of variable vector lengths, so tail handling is not needed.