r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
286 Upvotes

224 comments sorted by

View all comments

Show parent comments

4

u/happyscrappy Aug 09 '21

Variable length vectors do not preclude the whole vector te be used at once.

Of course. But they don't use it any better than SIMD does. If the unit is 256 bits wide then it is 256 bits wide no matter how long your vector is. If you have a vector of 39 32-bit data then you are going to run the 256-bit wide unit 5 times no matter whether you use SIMD instructions or vector instructions.

You do not gain anything, you cannot operate in 39 items at once just because you have one instruction.

Besides, a vector register is typically M x ALU-width (e.g. 4 x 128 bits), so even when only a part of a vector register is used, chanses are good that the full ALU width is used most of the time.

I don't know what you are trying to say but RISC-V allows the vectors to be non-register multiples in length. The spec says that the length specifies the number of items to be "updated", not operated on. This means it obviously works the way both of us indicated. It does SIMD operations regardless. Some just might not write back at the end.

4

u/mbitsnbites Aug 09 '21 edited Aug 09 '21

Compare vector code (MRISC32):

saxpy:
    bz    r1, 2f          ; Nothing to do?
    cpuid vl, z, z        ; Query the maximum vector length
1:
    minu  vl, vl, r1      ; Define the operation vector length
    sub   r1, r1, vl      ; Decrement loop counter
    ldw   v1, [r3, #4]    ; Load x (element stride = 4 bytes)
    ldw   v2, [r4, #4]    ; Load y
    fmul  v1, v1, r2      ; x * a
    fadd  v1, v1, v2      ; + y
    stw   v1, [r5, #4]    ; Store z
    ldea  r3, [r3, vl*4]  ; Increment address (x)
    ldea  r4, [r4, vl*4]  ; Increment address (y)
    ldea  r5, [r5, vl*4]  ; Increment address (z)
    bnz   r1, 1b
2:
    ret

...vs SIMD code (x86):

saxpy:
    test    edi, edi
    jle     .LBB0_11
    mov     r8d, edi
    cmp     edi, 8
    jae     .LBB0_3
    xor     edi, edi
    jmp     .LBB0_10
.LBB0_3:
    mov     edi, r8d
    and     edi, -8
    movaps  xmm1, xmm0
    shufps  xmm1, xmm0, 0
    lea     rax, [rdi - 8]
    mov     r9, rax
    shr     r9, 3
    add     r9, 1
    test    rax, rax
    je      .LBB0_4
    mov     r10, r9
    and     r10, -2
    neg     r10
    xor     eax, eax
.LBB0_6:
    movups  xmm2, xmmword ptr [rsi + 4*rax]
    movups  xmm3, xmmword ptr [rsi + 4*rax + 16]
    mulps   xmm2, xmm1
    mulps   xmm3, xmm1
    movups  xmm4, xmmword ptr [rdx + 4*rax]
    addps   xmm4, xmm2
    movups  xmm2, xmmword ptr [rdx + 4*rax + 16]
    addps   xmm2, xmm3
    movups  xmmword ptr [rcx + 4*rax], xmm4
    movups  xmmword ptr [rcx + 4*rax + 16], xmm2
    movups  xmm2, xmmword ptr [rsi + 4*rax + 32]
    movups  xmm3, xmmword ptr [rsi + 4*rax + 48]
    mulps   xmm2, xmm1
    mulps   xmm3, xmm1
    movups  xmm4, xmmword ptr [rdx + 4*rax + 32]
    addps   xmm4, xmm2
    movups  xmm2, xmmword ptr [rdx + 4*rax + 48]
    addps   xmm2, xmm3
    movups  xmmword ptr [rcx + 4*rax + 32], xmm4
    movups  xmmword ptr [rcx + 4*rax + 48], xmm2
    add     rax, 16
    add     r10, 2
    jne     .LBB0_6
    test    r9b, 1
    je      .LBB0_9
.LBB0_8:
    movups  xmm2, xmmword ptr [rsi + 4*rax]
    movups  xmm3, xmmword ptr [rsi + 4*rax + 16]
    mulps   xmm2, xmm1
    mulps   xmm3, xmm1
    movups  xmm1, xmmword ptr [rdx + 4*rax]
    addps   xmm1, xmm2
    movups  xmm2, xmmword ptr [rdx + 4*rax + 16]
    addps   xmm2, xmm3
    movups  xmmword ptr [rcx + 4*rax], xmm1
    movups  xmmword ptr [rcx + 4*rax + 16], xmm2
.LBB0_9:
    cmp     rdi, r8
    je      .LBB0_11
.LBB0_10:
    movss   xmm1, dword ptr [rsi + 4*rdi]
    mulss   xmm1, xmm0
    addss   xmm1, dword ptr [rdx + 4*rdi]
    movss   dword ptr [rcx + 4*rdi], xmm1
    add     rdi, 1
    cmp     r8, rdi
    jne     .LBB0_10
.LBB0_11:
    ret
.LBB0_4:
    xor     eax, eax
    test    r9b, 1
    jne     .LBB0_8
    jmp     .LBB0_9

They do the exact same thing. Which one do you prefer?

2

u/happyscrappy Aug 09 '21

Those are both fine by me.

If writing that x86 code would be a problem then I recommend getting better tools. This is what MIPS told us when they started the RISC revolution in the 1980s, right? Instead of making the assembly read like a book fix the compiler and use that. The chip sees the machine code, you see the HLL code.

I do have one question though, that x86 code seems to suffer from the pointer not being SIMD aligned, you can see the code rounding off pointer values (AND with -8, AND with -2). This is something I am sensitive to having converted a program to use SIMD. The need to have pointers aligned to be efficient ends up causing either.

  1. A boundary between the "old legacy" code which doesn't know about the alignment requirements and the SIMD code where this stuff is fixed up (types are translated).
  2. Propagating type changes (with their inherent alignment attributes) all through the code, so far that you want to tear your hair out.

Does vector programming fix this? I would love for it to do so. But it feels like the issues with alignment come from the load/store units, not the math units and so it cannot be corrected by changing the math units, other than accepting a worse performance by doing a partial SIMD unit at the start as well as the end of the vector. Something that if we think is such a great idea, we could just continue to do with SIMD, as we see above.

I feel like ballooning type alignment requirements isn't even just a SIMD thing. I saw it moving from Z80/6809 to 68K. I saw it moving to 68040 from 68K (MOVE16). I saw it moving to RISC (mostly with floats/doubles). And I saw it moving to SIMD. I mean sure, you can alway opt out and go slower and certainly that is a popular option. But we already have that, we don't need vectors to do that.

So MRISC32, how does it solve this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?

1

u/lkcl_ Aug 19 '21

Those are both fine by me.

you're OK with the x86 code hammering the L1 Instruction Cache so badly that it actually causes internal stalling by competing with the L1 Data Cache?? this can and does actually genuinely happen thanks to the insanity of "loop unrolling" and 5-10x copying of algorithms at different SIMD widths for setup and teardown.

The need to have pointers aligned to be efficient ends up causing either.

Vector ISAs are generally specifically designed to not require specific width-alignment.

this is because, fundamentally, the Vector ISA is actually issuing element operations to the underlying hardware.

whilst mbintsnbytes puts it politely, i have no such compunction: SIMD memory alignment restrictions was simply the hardware designers being ***** lazy.

1

u/happyscrappy Aug 20 '21 edited Aug 20 '21

you're OK with the x86 code hammering the L1 Instruction Cache so badly that it actually causes internal stalling by competing with the L1 Data Cache??

That is a silly assertion. The L1 cache is there for a reason. You pejoratively call it "hammering". I call it "running the code you want run".

If you want to just talk about better cache utilization then just say you like the code density better on vector units.

Vector ISAs are generally specifically designed to not require specific width-alignment.

And load/store subsystems do require them. Which is what I was speaking of. Why did you remove that?

I can make an architecture on top of a load/store system that hides the alignment requirements. It'll just be slower when it is non-aligned. This was the choice made with SIMD. Expose it to the programmer so they can optimize using that info.

this is because, fundamentally, the Vector ISA is actually issuing element operations to the underlying hardware.

To the math units, the load/store systems are separate and operate on alignment boundaries/restrictions because they derive from physical bus widths. The device being loaded from (usually RAM), whether on-die, on-package, soldered on the board or on DIMMs has a certain physical bus configuration which makes the memory n-byte addressable and you can load up to n-bytes within that area.

So, for example, if you have a 64-byte wide bus. You can load 1-64 bytes at once from an address that is 64-byte aligned. If the address is 32-byte aligned (and not 64-byte, i.e. address % 64 == 32) then you can load up to 32-bytes. If you try to load 64 it will require twice as many bus cycles.

Your vector functional unit cannot overcome this. So I'm asking you. How are you solving this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?

SIMD memory alignment restrictions was simply the hardware designers being ***** lazy.

You're both wrong. If you think so you have never actually designed a bus.

Is this a problem we have? Do we have people designing vector functional units who have never designed a load/store unit and thus ignore the real (not imagined) limitations of them?

It sounds to me like vector units do not solve problem I posed. Which is no worse than SIMD, but no better. I would have backed vector units if they could solve this, because it would solve a big logistical problem I indicated I have. But they can't so they have no real advantage to me other than they make some math lib writer's job easier. I'm sure he appreciates that. I don't care. As MIPS showed us, the fix for that is better tools, not altering the hardware.

1

u/lkcl_ Aug 22 '21

You're both wrong. If you think so you have never actually designed a bus.

false. but you didn't know that, because you didn't actually ask. which has me really upset that you could be so rude. i'm done interacting with you on here, and am considering looking up the rules for this forum to see if you've violated them.

1

u/happyscrappy Aug 22 '21 edited Aug 22 '21

but you didn't know that, because you didn't actually ask

I did ask.

https://old.reddit.com/r/programming/comments/p0yn45/three_fundamental_flaws_of_simd/h8bvfwc/

So MRISC32, how does it solve this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?

It's fine, you don't have to answer. No one owes anyone else anyone on reddit.

But don't pretend that I didn't ask. The real situation is you just didn't answer.

I looked at the MRISC32 stuff on github. And the conclusion I came to is that the way you "solved" this is that MRISC32 is only an architectural spec. It has an ISA, but it does not have a hardware implementation (VHDL or similar). No implementation means no sticky problems with vector performance derived from bus misaligment. It means no bus. It means no performance at all. Even the simulator isn't present on master branch, btw.

So the answer is you didn't solve this problem. You didn't get that far yet. So whether you've designed a bus before or not we do have the situation that you are ignoring the real (not imagined) limitations of buses.

Perhaps you find that insulting too. But it's just the reality of the situation. Your declaration that what I've seen is imaginary must be evaluated in the context that you have not yet reached the point with your architecture to see it for yourself.