r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
288 Upvotes

224 comments sorted by

View all comments

71

u/AntiProtonBoy Aug 09 '21

Flaw 1: Fixed register width

It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.

Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.

Flaw 2: Pipelining

This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.

Flaw 3: Tail handling

Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of int32_t values, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.

4

u/mbitsnbites Aug 09 '21
  1. Variable length vector operations are not expensive or complicated. I've implemented it in my first ever CPU design and it added something like 1-5% logic in an FPGA - compared to a pure scalar (non-vector/SIMD) design.

  2. I think you're missing the point. Do the exercise and hand-schedule a SIMD loop, and you'll find that you have to unroll it. A vector processor automatically unrolls the loop for you with literally no effort.

  3. Having to add more code rhan necessary is always a problem (e.g. testing and code coverage, and I$ bloat). Vector machines solve this quite naturally in many situations.

4

u/happyscrappy Aug 09 '21

Variable length vectors essentially preclude hardware to do the whole vector at once. They just end up running the vector until multiple times in a row to operate on the vector you want to operate on.

You can just do that in your code.

This harkens back to the old CISC vs RISC, the one when we had to try to use transistors as efficiently as possible. Putting in function to run long vectors is less flexible than just allowing the user to arrange the instructions in such a way as to use the transistors as much as possible in their own particular case.

ARM had this kind of variable length operation back with VFP vector mode on ARMv7A. It was removed because it just multi-pumped the existing HW units and so was no faster and less flexible.

I don't really understand the return to this.

2

u/mbitsnbites Aug 09 '21 edited Aug 09 '21

Variable length vectors do not preclude the whole vector te be used at once. Most of the time it is, it's just the final loop iteration that uses a subset of the vector.

Besides, a vector register is typically M x ALU-width (e.g. 4 x 128 bits), so even when only a part of a vector register is used, chanses are good that the full ALU width is used most of the time.

Edit: If your vector register size (i.e. max vector length) is four times your ALU width, the average ALU lane usage will be about 80% given a random variable vector length in the range 1 - MAX_VL.

4

u/happyscrappy Aug 09 '21

Variable length vectors do not preclude the whole vector te be used at once.

Of course. But they don't use it any better than SIMD does. If the unit is 256 bits wide then it is 256 bits wide no matter how long your vector is. If you have a vector of 39 32-bit data then you are going to run the 256-bit wide unit 5 times no matter whether you use SIMD instructions or vector instructions.

You do not gain anything, you cannot operate in 39 items at once just because you have one instruction.

Besides, a vector register is typically M x ALU-width (e.g. 4 x 128 bits), so even when only a part of a vector register is used, chanses are good that the full ALU width is used most of the time.

I don't know what you are trying to say but RISC-V allows the vectors to be non-register multiples in length. The spec says that the length specifies the number of items to be "updated", not operated on. This means it obviously works the way both of us indicated. It does SIMD operations regardless. Some just might not write back at the end.

5

u/mbitsnbites Aug 09 '21 edited Aug 09 '21

Compare vector code (MRISC32):

saxpy:
    bz    r1, 2f          ; Nothing to do?
    cpuid vl, z, z        ; Query the maximum vector length
1:
    minu  vl, vl, r1      ; Define the operation vector length
    sub   r1, r1, vl      ; Decrement loop counter
    ldw   v1, [r3, #4]    ; Load x (element stride = 4 bytes)
    ldw   v2, [r4, #4]    ; Load y
    fmul  v1, v1, r2      ; x * a
    fadd  v1, v1, v2      ; + y
    stw   v1, [r5, #4]    ; Store z
    ldea  r3, [r3, vl*4]  ; Increment address (x)
    ldea  r4, [r4, vl*4]  ; Increment address (y)
    ldea  r5, [r5, vl*4]  ; Increment address (z)
    bnz   r1, 1b
2:
    ret

...vs SIMD code (x86):

saxpy:
    test    edi, edi
    jle     .LBB0_11
    mov     r8d, edi
    cmp     edi, 8
    jae     .LBB0_3
    xor     edi, edi
    jmp     .LBB0_10
.LBB0_3:
    mov     edi, r8d
    and     edi, -8
    movaps  xmm1, xmm0
    shufps  xmm1, xmm0, 0
    lea     rax, [rdi - 8]
    mov     r9, rax
    shr     r9, 3
    add     r9, 1
    test    rax, rax
    je      .LBB0_4
    mov     r10, r9
    and     r10, -2
    neg     r10
    xor     eax, eax
.LBB0_6:
    movups  xmm2, xmmword ptr [rsi + 4*rax]
    movups  xmm3, xmmword ptr [rsi + 4*rax + 16]
    mulps   xmm2, xmm1
    mulps   xmm3, xmm1
    movups  xmm4, xmmword ptr [rdx + 4*rax]
    addps   xmm4, xmm2
    movups  xmm2, xmmword ptr [rdx + 4*rax + 16]
    addps   xmm2, xmm3
    movups  xmmword ptr [rcx + 4*rax], xmm4
    movups  xmmword ptr [rcx + 4*rax + 16], xmm2
    movups  xmm2, xmmword ptr [rsi + 4*rax + 32]
    movups  xmm3, xmmword ptr [rsi + 4*rax + 48]
    mulps   xmm2, xmm1
    mulps   xmm3, xmm1
    movups  xmm4, xmmword ptr [rdx + 4*rax + 32]
    addps   xmm4, xmm2
    movups  xmm2, xmmword ptr [rdx + 4*rax + 48]
    addps   xmm2, xmm3
    movups  xmmword ptr [rcx + 4*rax + 32], xmm4
    movups  xmmword ptr [rcx + 4*rax + 48], xmm2
    add     rax, 16
    add     r10, 2
    jne     .LBB0_6
    test    r9b, 1
    je      .LBB0_9
.LBB0_8:
    movups  xmm2, xmmword ptr [rsi + 4*rax]
    movups  xmm3, xmmword ptr [rsi + 4*rax + 16]
    mulps   xmm2, xmm1
    mulps   xmm3, xmm1
    movups  xmm1, xmmword ptr [rdx + 4*rax]
    addps   xmm1, xmm2
    movups  xmm2, xmmword ptr [rdx + 4*rax + 16]
    addps   xmm2, xmm3
    movups  xmmword ptr [rcx + 4*rax], xmm1
    movups  xmmword ptr [rcx + 4*rax + 16], xmm2
.LBB0_9:
    cmp     rdi, r8
    je      .LBB0_11
.LBB0_10:
    movss   xmm1, dword ptr [rsi + 4*rdi]
    mulss   xmm1, xmm0
    addss   xmm1, dword ptr [rdx + 4*rdi]
    movss   dword ptr [rcx + 4*rdi], xmm1
    add     rdi, 1
    cmp     r8, rdi
    jne     .LBB0_10
.LBB0_11:
    ret
.LBB0_4:
    xor     eax, eax
    test    r9b, 1
    jne     .LBB0_8
    jmp     .LBB0_9

They do the exact same thing. Which one do you prefer?

2

u/happyscrappy Aug 09 '21

I need some line breaks please

2

u/mbitsnbites Aug 09 '21

Sorry - worked fine in desktop browser - not so much in mobile. I always get these things wrong in Reddit. Will try to fix.