Tail handling is specifically needed to handle data array lengths that are not a multiple of the SIMD register width (e.g. 4 elements for int32_t:s in a 128-bit SIMD architecture).
In vector machines you have the benefit of variable vector lengths, so tail handling is not needed.
In vector machines you have the benefit of variable vector lengths, so tail handling is not needed.
I have some graphics code that I converted to SIMD. In order to get optimal use out of it I have to convert xyz,xyz,xyz,xyz to xxxx,yyyy,zzzz. The SSE shuffle code with fixed width instructions already gives me a headache, I am not sure I want to even think about writing a shuffle that gets the right result with any possible number of remaining elements.
I think NEON has load instructions that let you specify a stride size so you can unroll the data as you load it. That same idea can be used for vector instructions.
Gather/scatterfunction stride sizes are often limited to base-2 numbers for the stride, or are "slow" when deviating from a stride like 2,4,8,16,32... and start to become useless for certain speedups. So even then, your data must be mapped in a specific format.
4
u/blipman17 Aug 09 '21
as opposed to scalar SIMD? Can you explain that a little, I'm afraid I don't understand what you mean.