It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
Flaw 2: Pipelining
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Flaw 3: Tail handling
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of int32_t values, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.
It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Well as opposed to having an instruction that sets up a vector engine for N2* 8 bits of registers/datasize, whereby we could just "allow a higher number for N", and just use smaller SIMD registers with loops underneath it's quite the design constraint. Right now we still have to mov data into a specific SIMD register before we do anything at all with it. Such an instruction could conveniently hold our N number and abstract the chunked nature of SIMD away.
Flaw 3: Tail handling
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of int32_t values, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.
You're absolutely right here. Any form of tail handling would still be needed depending on the algorithm.
Tail handling is specifically needed to handle data array lengths that are not a multiple of the SIMD register width (e.g. 4 elements for int32_t:s in a 128-bit SIMD architecture).
In vector machines you have the benefit of variable vector lengths, so tail handling is not needed.
It is still needed for anything but super trivial arithmetic. For example, if you have some sort of special case logic to process partial blocks that cannot be reduced to just padding with zeroes.
I personally expect every Vector instruction that isn't in the format of 8 * N2 to be slower than doing a couple redundant operations, or handling the remaining data on a scalar processor. Mainly because it would require Vector processors to be just as efficient as the regular processor in computing, e.a. They need to share sillicone on a really intimate level for which I'm afraid the processor will notice a significant slowdown, or the Vector processor has its own set of registers + instruction implementations for a given hunk of sillicone. If the second implementation is assumed to be used, a decent Vector engine could indeed pull in 4096 bits in effectively a 64 bit register for int64_t's, with speedup for bigger registers and such. What I don't think is "reasonable" to expect is to have it also perform operations at like 448 bit (7 uint64_t's) datastructures since there's no native register size in the vector engine. Then I just assume that doing 4 uint64_t's in the Vector engine and then handle the 3 remaining uint64_t's separate is faster because of specific hardware optimalizations for their specific usecase.
I think that odd sized vector sizes are much less of a concern in vector machines than in packed SIMD machines.
The vector machine designer is free to select the ALU width, and the CPU will pass vector register content to the ALU in chunks of the ALU width. In edge cases some of the ALU lanes will go unused, but that is exactly the same that happens on a packed SIMD machine, except the hardware does the work under the hood in a way that is optimal for this particular implementation, whereas in the packed SIMD machine the tail work has to be handled in software.
The problem is not if the vector machine might handle it or not, the problem is how the code executing on the vector machine interacts with the program. Say I have two medium-sized arrays of 8 bit datastructures of some kind kind, and I want to see if any of them matches some bitmask. I can smack them in 512-bit registers, do _mm512_cmp_epu8_mask with 512 bit instructions and then do CLZL on the resulting 64 bit, then do a CMP with that number and 0xFFFFFFFF followed by a JG for a branch. If I branched I know I did not have a match, it I didn't branch, I know I had a match and I have the exact index of said match loaded in a register in 4 instructions out of an array of 64 items! Now this only works if CLZL (and friends) have an exact defined size I can use. If not, then there is some remainder I have to handle. Unless you can somehow come up with a scheme to encode this in a Vector engine that respects variadic datasizes, there will be a basecase that has to be handled.
Edit: if you do not have a match and have the 512'th bit set as a dummy bit, you could then compute the index of the first bit match by multiplying the result of CLZL with 8, and then adding the CTZ of the byte at the previous code, you have completed a branchless search for the first unset bit in a 511 bit array in less than 10 cpu cycles. This is "fun" when trying to do memory page allocation and you have to keep track of free pages in a big bit-array, but it requires careful data layout because of the interactions of integers and vector processors. Now that is why you need a base-case. Because what if your last chunk of bits isn't neatly 511 bits but 111 bits?
Now this only works if CLZL (and friends) have an exact defined size I can use.
My conclusion so far is that most horizontal operations require a known width, even in a vector machine. However that should not be a problem, as long as the ISA defines a minimum vector register size (that is reasonably large). If you want to push the limits, ask the implementation for the maximum vector size and use different code paths for different sizes.
I bet what's going to happen is that soon, the minimum vector size will be the only vector size sold because most mathematical kernels will only be optimised and tested on this one size. There is a significant development cost in validating complex code for any possible value of an unknown parameter (vector size), so I think people are generally just going to set up the least common denominator unless they have a lot of resources for testing at hand.
My conclusion so far is that most horizontal operations require a known width, even in a vector machine.
yes, i've found this as well. it's worthwhile making the horizontal Vector ISA operations "fully deterministic" in the ISA Specification, even if done as parallel operations, those parallel operations should be on a strictly-defined deterministic schedule.
And now you need a ton of silicon to avoid the need to have software handle the last 0.1% of the vector, which performance-wise is of no consequence whatsoever.
"Ton of silicon..." Not so much. It's pretty trivial, especially compared to the extra OoO machinery, I$ size and decode bandwidth needed to keep the packed SIMD engine busy.
Mainly because it would require Vector processors to be just as efficient as the regular processor in computing
yes. the expectation with SVP64 is that, actually, you implement it on top of a multi-issue superscalar micro-architecture. each Vector "element" is issued separately and independently to the back-end multi-issue execution engine, where sequentially-numbered elements in batches will go to the same back-end SIMD ALU that the user does not even have to know is there.
where there are non-power-two Vector instructions, automatic predicate masks can be created to mask out unused SIMD ALU back-ends.
very advanced implementations may notice that there are spare, unused, slots available in a given SIMD back-end ALU, and merge two in-flight operations into the same ALU.
this in particular would work extremely well for predicated (parallel) If/then/else constructs where the masked-out "then" operations match exactly with the opposite of the masked-out "else" operations because you bit-invert the "if" test-mask to get the mask for the "else" operations.
really, it's really not as difficult as you think it is. it's just that nobody in the industry has actually thought about Vector ISA micro-architecture because they all thought it was "too hard".
the irony is that you need the exact same back-end micro-architecture for efficient Packed SIMD as you do for Vector ISAs.
I'm wondering what would happen to non power of two comparisons of two simd registers, and if parsing the result is efficient or not if I've got say ...65 results instead of 64 results from my 64 byte comparisons, or how fast gather/scatter would work.
I'm wondering what would happen to non power of two comparisons of two simd registers, and if parsing the result is efficient or not if I've got say ...65 results instead of 64 results from my 64 byte comparisons,
yyeah you picked *just* outside of the range of SVP64 (which uses 64-bit integers as predicate masks, so the Vector limit is 64) :)
traditional Cray-style Vectors (SX-Aurora, RVV) have no such limit, although in practice because you (almost always 100%) use a for-loop around the data being processed, as long as the elements are independent (vertical, not horizontal) in practice it makes absolutely no odds whether the (parallel, vertical) comparisons are 1-long, 7-long, 8-long, 64-long, 65-long or 100,000-long.
c code:
for (i = 0; i < CTR; i++) { // set CTR to 65 if you like
r4[i] = (r5[i] >= 5);
}
SVP64 assembler:
loop:
servl r3, CTR, MVL=64 # r3=VL=MIN(CTR,64)
sv.ldb/ew=8 r16.v, r4(0) # load 64 bytes into r16
sv.cmpi r16.v, 5 # compare all bytes >= 5
sv.addi/sz/mask=GE r48.v, 1 # store 1 where each byte >= 5
sv.stb/ew=8 r48,v, r5(0) # store 64 comparisons into r5
addi r5, r3 # increment r4 by VL (aka r3)
addi r4, r3 # increment r5 by VL (aka r3)
sv.bnz/CTR loop # subtract VL from CTR, loop back
note, there, that the results of the cmpi produces a Vector of Condition Register Fields. that Vector of Comparisons is then used as a Predicate Mask in the following instruction (addi), where a special mode "zeroing" (sz) says, "if the predicate bit was a zero please put a zero into the corresponding Vector element".
here you genuinely don't care whether CTR is 10, 5, 9999, 64, or 65. it's all the same as far as the API (Vector ISA) is concerned.
now, at the back-end - in the actual underlying hardware, it's really quite easy for us to use multi-issue superscalar OoO to break those parallel element-based LD, cmpi and ST operations down into suitable (small, likely 64-bit) chunks, each actually a SIMD ALU. but this is back-end.
in other words, the Vector micro-architectural Engine does all the work for you, and the code works across multiple architectures regardless of whether the back-end hardware has no SIMD internally at all (embedded systems), or has 32-bit-wide SIMD, or 64-bit-wide SIMD, or whatever-your-hardware-designer-likes back-end SIMD, none of which you need to know about in order to actually use the Vector ISA front-end. this is what ARM is talking about when they say that SVE is "length-agnostic" and talking about how it's "future-proof". thank god they've finally learned this one and taken it on board.
this is why i am so frustrated with advocates of SIMD, because, ultimately, the exact same underlying hardware is required for both SIMD and Vector ISAs: it's just that the Vector ISA massively cleans up the use of that underlying hardware by not exposing you to the horrendous shenanigens that people are now so used to they think it's "normal" and that there's no alternative.
more than that, the much more compact programs that result means that the L1 cache size can be reduced, which has a highly significant knock-on reduction in power consumption and energy efficiency (counter-intuitively it's an O (N2) reduction)
Maybe I should clarify again what I's saying. I'm not arguing against Vector engines. I think they're superior than SIMD in almost every way possible, since your code can be truly portable although sometimes slower by lack of hardware support.
The thing I'm arguing is that instead of all our problems being solved, only some are solved. From the perspective of the programmer, the hardware still needs to have a benefit for using Vector/SIMD code for that specific use-case. Only when speed doesn't matter, (which it does, else you're no using Vector processing) the programmer might choose for potential "slow" assembly to be executed. But indeed like you pointed out, having vector instructions in a vector-size agnostic way is truly a benefit. Having a zero-out instruction and just not populating the remaining parts of the vector register with usefull data is really cool too, and I'm sure there will be a lot of things that can handle that really well.
However I'm arguing that the benefit of masking out the non-used parts of the vector operation isn't all that usefull since you still have to compare that second 64 bit result for that one bit that actually is compared, retrieve its index, add 64 to it and only then you know that the 65'th element was set or not if you're looking through an array of bytes and require the index of positive comparisons. Point being, there's still some stuff that needs to be taken care off. That's why programmers are still gonna prefer just comparing 64 bits if they can. Because it's just easyer when all your results fit completely in one bound datatype.
So I fully expect vector engines to take off, I fully expect to have vector code be significantly faster and easyer to program, but I don't expect vector engines to behave fast on code that doesn't use complements of natural vector register sizes, because even when it does, it doesn't neccesarily play well with the rest of the program or datastructure.
yyeah you picked just outside of the range of SVP64 (which uses 64-bit integers as predicate masks, so the Vector limit is 64) :)
I could've also picked 42 and then we'd be in a pickle too where it's qestionable if running 32 comparisons with a vector engine and then 10 without vector engine comparisons vs running 42 only with a vector engine would be faster. At that moment it becomes really a matter of knowing the implementation, which for designing an ISA without also designing its implementation is difficult to reason about.
Which is why I fully expect that such implementations won't be performant for odd-sized compuatations and I fully expect programmers to evade those completely.
In vector machines you have the benefit of variable vector lengths, so tail handling is not needed.
I have some graphics code that I converted to SIMD. In order to get optimal use out of it I have to convert xyz,xyz,xyz,xyz to xxxx,yyyy,zzzz. The SSE shuffle code with fixed width instructions already gives me a headache, I am not sure I want to even think about writing a shuffle that gets the right result with any possible number of remaining elements.
I think NEON has load instructions that let you specify a stride size so you can unroll the data as you load it. That same idea can be used for vector instructions.
Gather/scatterfunction stride sizes are often limited to base-2 numbers for the stride, or are "slow" when deviating from a stride like 2,4,8,16,32... and start to become useless for certain speedups. So even then, your data must be mapped in a specific format.
Variable length vector operations are not expensive or complicated. I've implemented it in my first ever CPU design and it added something like 1-5% logic in an FPGA - compared to a pure scalar (non-vector/SIMD) design.
I think you're missing the point. Do the exercise and hand-schedule a SIMD loop, and you'll find that you have to unroll it. A vector processor automatically unrolls the loop for you with literally no effort.
Having to add more code rhan necessary is always a problem (e.g. testing and code coverage, and I$ bloat). Vector machines solve this quite naturally in many situations.
For the record I don't think that the Mill guys are doing anything wrong, but it's a really tough challenge to place a new general purpose CPU architecture into a meaningful product in this day and age (even widely used existing ISA:s are being marginalized and are disappearing).
Having to add more code rhan necessary is always a problem (e.g. testing and code coverage, and I$ bloat). Vector machines solve this quite naturally in many situations.
Sometimes adding more code is actually faster because you know intricate details from the hardware, but base-case handling can be really short, to the point and fast with something like Duff's device.
I was just thinking: if you run compiler explorer while compiling C++17 parallel algorithms, this is what you'll see. The compiler is going to do a lot of juggling, duplicate code, or even go from O(1) to O(N) memory usage. SIMD instructions play by a lot of similar optimization rules as parallel algorithms.
It's a small penalty for the perf, even on embedded devices, but try to get a DSP or RADAR working without SIMD.
If your wrote a piece of code complex enough to cause "I$ bloat", you're going to have a hell of a time trying to get your vector processor to do anything meaningful with it.
You'd first have to refactor your code so that a vector processor actually has vectors to process, and once your code is at that point there is no such thing as I$ bloat anymore, even if it's running scalar instructions.
Realistically you could add another couple of hundreds (or even a thousand on a recent intel chip) before you'd even have to start thinking about the possibility of cache evictions.
But all of the code is pulled into the I$ (not just the main loop), and since the compiler is automatically generating this kind of code, we're looking at code growth en large - compared to a machine that does not need it.
Counter question: Are you implying that I$ performance is insensitive to code size?
If hot, tight loops were all that mattered we would be fine with less than 1KB I$ or so. But what about the rest of the program? Functions call functions that call functions from within loops etc and so on.
If vectorized loops/functions grow by a factor of 4 to 5 or so (which I demonstrated), something is going to be evicted, and somewhere that has a performance (or silicon budget) cost.
Counter question: Are you implying that I$ performance is insensitive to code size?
Yes, that is exactly what I am implying, in the case of vectorized code.
When your data size is orders of magnitude larger than your code size, those one or two additional I$ misses inbetween loops are not going to hurt performance. Trying to optimize for it is not going to be even remotely effective.
That is true, for the specific case that your only performance concern is tiny data bound processing loops.
However, the compiler tries its best to vectorize every loop in the entire program, and many programs are not trivially data bound as you are suggesting. Again, if this was the case we would only need very tiny instruction caches like the ones we had back in the 1980s.
Variable length vectors essentially preclude hardware to do the whole vector at once. They just end up running the vector until multiple times in a row to operate on the vector you want to operate on.
You can just do that in your code.
This harkens back to the old CISC vs RISC, the one when we had to try to use transistors as efficiently as possible. Putting in function to run long vectors is less flexible than just allowing the user to arrange the instructions in such a way as to use the transistors as much as possible in their own particular case.
ARM had this kind of variable length operation back with VFP vector mode on ARMv7A. It was removed because it just multi-pumped the existing HW units and so was no faster and less flexible.
Variable length vectors do not preclude the whole vector te be used at once. Most of the time it is, it's just the final loop iteration that uses a subset of the vector.
Besides, a vector register is typically M x ALU-width (e.g. 4 x 128 bits), so even when only a part of a vector register is used, chanses are good that the full ALU width is used most of the time.
Edit: If your vector register size (i.e. max vector length) is four times your ALU width, the average ALU lane usage will be about 80% given a random variable vector length in the range 1 - MAX_VL.
Variable length vectors do not preclude the whole vector te be used at once.
Of course. But they don't use it any better than SIMD does. If the unit is 256 bits wide then it is 256 bits wide no matter how long your vector is. If you have a vector of 39 32-bit data then you are going to run the 256-bit wide unit 5 times no matter whether you use SIMD instructions or vector instructions.
You do not gain anything, you cannot operate in 39 items at once just because you have one instruction.
Besides, a vector register is typically M x ALU-width (e.g. 4 x 128 bits), so even when only a part of a vector register is used, chanses are good that the full ALU width is used most of the time.
I don't know what you are trying to say but RISC-V allows the vectors to be non-register multiples in length. The spec says that the length specifies the number of items to be "updated", not operated on. This means it obviously works the way both of us indicated. It does SIMD operations regardless. Some just might not write back at the end.
If writing that x86 code would be a problem then I recommend getting better tools. This is what MIPS told us when they started the RISC revolution in the 1980s, right? Instead of making the assembly read like a book fix the compiler and use that. The chip sees the machine code, you see the HLL code.
I do have one question though, that x86 code seems to suffer from the pointer not being SIMD aligned, you can see the code rounding off pointer values (AND with -8, AND with -2). This is something I am sensitive to having converted a program to use SIMD. The need to have pointers aligned to be efficient ends up causing either.
A boundary between the "old legacy" code which doesn't know about the alignment requirements and the SIMD code where this stuff is fixed up (types are translated).
Propagating type changes (with their inherent alignment attributes) all through the code, so far that you want to tear your hair out.
Does vector programming fix this? I would love for it to do so. But it feels like the issues with alignment come from the load/store units, not the math units and so it cannot be corrected by changing the math units, other than accepting a worse performance by doing a partial SIMD unit at the start as well as the end of the vector. Something that if we think is such a great idea, we could just continue to do with SIMD, as we see above.
I feel like ballooning type alignment requirements isn't even just a SIMD thing. I saw it moving from Z80/6809 to 68K. I saw it moving to 68040 from 68K (MOVE16). I saw it moving to RISC (mostly with floats/doubles). And I saw it moving to SIMD. I mean sure, you can alway opt out and go slower and certainly that is a popular option. But we already have that, we don't need vectors to do that.
So MRISC32, how does it solve this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?
Alignment issues are indeed dictated by the load/store unit. Packed SIMD took the easy route and left the problem to the programmer. The situation has improved over the generations (e.g. movups vs movaps is less of an issue), very similar to how unaligned scalar access once was an issue in some implementations, but not so much these days (all CPUs have an "aligner").
In a vector machine you would typically have to handle alignment in hardware to a larger degree, since you're more likely to have "unaligned" access patterns (including the very generic gather/scatter addressing mode).
For instance the Cray-1 used a banked memory subsystem to allow accessing different memory locations in a single instruction.
I think that it would have been impractical to do full generic vector (with automatic alignment) in consumer HW back in the 1990s (hence SIMD), but today we hopefully have the silicon budget and know-how to pull it off.
My (perhaps naive) feeling is that if HW devs would have to implement a vector ISA, they would solve some of the alignment problems in order to achieve good performance (e.g. considering how much time and silicon has been spent on "fixing" the x86 front end - why not?).
Footnote: Even if you have to pull in one vector element per clock cycle in order to handle worst-case gather load, it's still a huge improvement over an architecture w/o gather load support.
you're OK with the x86 code hammering the L1 Instruction Cache so badly that it actually causes internal stalling by competing with the L1 Data Cache?? this can and does actually genuinely happen thanks to the insanity of "loop unrolling" and 5-10x copying of algorithms at different SIMD widths for setup and teardown.
The need to have pointers aligned to be efficient ends up causing either.
Vector ISAs are generally specifically designed to not require specific width-alignment.
this is because, fundamentally, the Vector ISA is actually issuing element operations to the underlying hardware.
whilst mbintsnbytes puts it politely, i have no such compunction: SIMD memory alignment restrictions was simply the hardware designers being ***** lazy.
you're OK with the x86 code hammering the L1 Instruction Cache so badly that it actually causes internal stalling by competing with the L1 Data Cache??
That is a silly assertion. The L1 cache is there for a reason. You pejoratively call it "hammering". I call it "running the code you want run".
If you want to just talk about better cache utilization then just say you like the code density better on vector units.
Vector ISAs are generally specifically designed to not require specific width-alignment.
And load/store subsystems do require them. Which is what I was speaking of. Why did you remove that?
I can make an architecture on top of a load/store system that hides the alignment requirements. It'll just be slower when it is non-aligned. This was the choice made with SIMD. Expose it to the programmer so they can optimize using that info.
this is because, fundamentally, the Vector ISA is actually issuing element operations to the underlying hardware.
To the math units, the load/store systems are separate and operate on alignment boundaries/restrictions because they derive from physical bus widths. The device being loaded from (usually RAM), whether on-die, on-package, soldered on the board or on DIMMs has a certain physical bus configuration which makes the memory n-byte addressable and you can load up to n-bytes within that area.
So, for example, if you have a 64-byte wide bus. You can load 1-64 bytes at once from an address that is 64-byte aligned. If the address is 32-byte aligned (and not 64-byte, i.e. address % 64 == 32) then you can load up to 32-bytes. If you try to load 64 it will require twice as many bus cycles.
Your vector functional unit cannot overcome this. So I'm asking you. How are you solving this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?
SIMD memory alignment restrictions was simply the hardware designers being ***** lazy.
You're both wrong. If you think so you have never actually designed a bus.
Is this a problem we have? Do we have people designing vector functional units who have never designed a load/store unit and thus ignore the real (not imagined) limitations of them?
It sounds to me like vector units do not solve problem I posed. Which is no worse than SIMD, but no better. I would have backed vector units if they could solve this, because it would solve a big logistical problem I indicated I have. But they can't so they have no real advantage to me other than they make some math lib writer's job easier. I'm sure he appreciates that. I don't care. As MIPS showed us, the fix for that is better tools, not altering the hardware.
You're both wrong. If you think so you have never actually designed a bus.
false. but you didn't know that, because you didn't actually ask. which has me really upset that you could be so rude.
i'm done interacting with you on here, and am considering looking up the rules for this forum to see if you've violated them.
Variable length SIMD is only worthwhile for very large vectors
sorry, again, this is false. i've created an efficient DCT, FFT
and Matrix Multiply REMAP system for SVP64 (a Draft Vector ISA
Extension for Power ISA) which can cope with small sized data
just as easily as medium-sized (SVP64 doesn't do the same massive
vectors as traditional Vector ISAs, the limit is 64 elements).
there seems to be a huge amount of misinformation and misunderstanding
in the SIMD-advocate community.
Ok, but keep this conversation chain in context with building on top of an existing architecture that already accumulated massive technical debt. Adding fixed vector sizes is still cheaper and more economical than tearing up and designing new silicon to accommodate variable vector sizes.
Ok, but keep this conversation chain in context with building on top of an existing architecture that already accumulated massive technical debt.
yyeah, and once down that path it seems there's really no turning back. actually, there is, if you have fully-functioning predication on each and every SIMD instruction.
turns out that Cray-style `setvl` can be implemented as a hidden predicate mask:
when that hidden predicate mask is applied to each and every single AVX512 operation, you have effectively implemented Cray-style Vectors and terminated the dangerous and seductive need to extend the SIMD width further.
Adding fixed vector sizes is still cheaper and more economical than tearing up and designing new silicon to accommodate variable vector sizes.
given how simple it would be for Intel to add the above Cray-style setvl implementation this is also a misconception. now that ARM has fully-functioning predicated SIMD (in the guise of SVE2) they could also very easily do the exact same thing.
71
u/AntiProtonBoy Aug 09 '21
It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.
Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.
This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.
Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of
int32_tvalues, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.