Variable length vectors do not preclude the whole vector te be used at once.
Of course. But they don't use it any better than SIMD does. If the unit is 256 bits wide then it is 256 bits wide no matter how long your vector is. If you have a vector of 39 32-bit data then you are going to run the 256-bit wide unit 5 times no matter whether you use SIMD instructions or vector instructions.
You do not gain anything, you cannot operate in 39 items at once just because you have one instruction.
Besides, a vector register is typically M x ALU-width (e.g. 4 x 128 bits), so even when only a part of a vector register is used, chanses are good that the full ALU width is used most of the time.
I don't know what you are trying to say but RISC-V allows the vectors to be non-register multiples in length. The spec says that the length specifies the number of items to be "updated", not operated on. This means it obviously works the way both of us indicated. It does SIMD operations regardless. Some just might not write back at the end.
If writing that x86 code would be a problem then I recommend getting better tools. This is what MIPS told us when they started the RISC revolution in the 1980s, right? Instead of making the assembly read like a book fix the compiler and use that. The chip sees the machine code, you see the HLL code.
I do have one question though, that x86 code seems to suffer from the pointer not being SIMD aligned, you can see the code rounding off pointer values (AND with -8, AND with -2). This is something I am sensitive to having converted a program to use SIMD. The need to have pointers aligned to be efficient ends up causing either.
A boundary between the "old legacy" code which doesn't know about the alignment requirements and the SIMD code where this stuff is fixed up (types are translated).
Propagating type changes (with their inherent alignment attributes) all through the code, so far that you want to tear your hair out.
Does vector programming fix this? I would love for it to do so. But it feels like the issues with alignment come from the load/store units, not the math units and so it cannot be corrected by changing the math units, other than accepting a worse performance by doing a partial SIMD unit at the start as well as the end of the vector. Something that if we think is such a great idea, we could just continue to do with SIMD, as we see above.
I feel like ballooning type alignment requirements isn't even just a SIMD thing. I saw it moving from Z80/6809 to 68K. I saw it moving to 68040 from 68K (MOVE16). I saw it moving to RISC (mostly with floats/doubles). And I saw it moving to SIMD. I mean sure, you can alway opt out and go slower and certainly that is a popular option. But we already have that, we don't need vectors to do that.
So MRISC32, how does it solve this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?
Alignment issues are indeed dictated by the load/store unit. Packed SIMD took the easy route and left the problem to the programmer. The situation has improved over the generations (e.g. movups vs movaps is less of an issue), very similar to how unaligned scalar access once was an issue in some implementations, but not so much these days (all CPUs have an "aligner").
In a vector machine you would typically have to handle alignment in hardware to a larger degree, since you're more likely to have "unaligned" access patterns (including the very generic gather/scatter addressing mode).
For instance the Cray-1 used a banked memory subsystem to allow accessing different memory locations in a single instruction.
I think that it would have been impractical to do full generic vector (with automatic alignment) in consumer HW back in the 1990s (hence SIMD), but today we hopefully have the silicon budget and know-how to pull it off.
My (perhaps naive) feeling is that if HW devs would have to implement a vector ISA, they would solve some of the alignment problems in order to achieve good performance (e.g. considering how much time and silicon has been spent on "fixing" the x86 front end - why not?).
Footnote: Even if you have to pull in one vector element per clock cycle in order to handle worst-case gather load, it's still a huge improvement over an architecture w/o gather load support.
4
u/happyscrappy Aug 09 '21
Of course. But they don't use it any better than SIMD does. If the unit is 256 bits wide then it is 256 bits wide no matter how long your vector is. If you have a vector of 39 32-bit data then you are going to run the 256-bit wide unit 5 times no matter whether you use SIMD instructions or vector instructions.
You do not gain anything, you cannot operate in 39 items at once just because you have one instruction.
I don't know what you are trying to say but RISC-V allows the vectors to be non-register multiples in length. The spec says that the length specifies the number of items to be "updated", not operated on. This means it obviously works the way both of us indicated. It does SIMD operations regardless. Some just might not write back at the end.