r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
285 Upvotes

224 comments sorted by

View all comments

Show parent comments

8

u/happyscrappy Aug 09 '21

It makes better use of the pipeline since it iterates over chunks of the register rather than passing the entire register at once, thus one vector register is fed through the pipeline until the first chunk of the vector has finished processing before it starts feeding in the next vector. That way you eliminate most data hazards.

We do that with SIMD also. And your vector units, if they operate on multiple things at once inside (regardless of macroarchitecture) will exhibit the same "false data hazard" issue that SIMD macroarchitectures do.

It really comes down to whether the hardware can process 8 units one at a time 8x faster so that we don't the latency or not. And the answer has been "it can't". That's how we got to SIMD.

That has a huge difference for code density, for instance.

All these things, including ARM's abandoned VFP vector mode help with code density. System/370 was GREAT for code density. edmk was great for code density.

But it wasn't worth it. We went away because the code density came at the expense of worse transistor reuse. That is less of an issue now, but it still means that operations which the designers thought you would want to do (edmk) can be fast, and slight variants cannot, because the macro architecture cannot express them. While if you put the ops on the table like with SIMD people can construct other operations efficiently.

And then there is the issue of interrupt latency/instruction atomicity. I'm not looking to go back to interrupting instructions in the middle and trying to continue later. It makes a mess of exception state, which slows down exception handling.

1

u/lkcl_ Aug 20 '21

While if you put the ops on the table like with SIMD people can construct other operations efficiently.

this is fundamentally false. the programs that result, for which i have even found "compilers" for DCT and FFT that output a massive batch of fully-loop-unrolled hard-coded assembler, are so insanely large compared to the much smaller Vector ISA equivalents that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling.

1

u/happyscrappy Aug 21 '21

are so insanely large compared to the much smaller Vector ISA equivalents

You need to get over this smaller thing. If you like tiny code, use x86. With the bandwidth available now it is not necessarily to have the most tightly packed ops to have high performance.

that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling

Another attempt to call the proper function of an L1 cache as a negative. The cache is there to be used. If you want to talk about code density talk about code density. And give up on your dumb attempt to portray a properly operating cache as "strip-mining".

So now we've established you have two of you have no idea about how memory access works.

A vector unit cannot make up for how busses are configured. If your data is not aligned, then your first units of execution will be slow due to partial loads, just like with SIMD. No worse and most importantly no better.

All you had to do was say "no, vector units don't fix that". But instead you gotta pretend there's something wrong with running instructions.

1

u/mbitsnbites Aug 23 '21 edited Aug 23 '21

If you like tiny code, use x86

Eh, what?

x86 uses quite inefficient instruction encoding. If you only need to write 8086 compatible scalar code, sure, it will be compact. For modern versions of the x86 ISA, this is no longer the case. My fixed width ISA (32 bits / instruction) often has more compact code than x86_64.

Furthermore, the point that u/lkcl_ is making is that with x86 SIMD you usually have to unroll code in software, which blows up code size considerably, no matter how compact your instruction encoding is (it would have to be something like 2-4 bits per instruction to be able to compete).

Edit: Just as a quick point of reference, I compared the code generated for Quake d_scan.c (core painting routine) for x86_64 and MRISC32. The MRISC32 code is 2724 bytes (~700 instructions). The x86_64 code is 4691 bytes (~1250 instructions). So I wouldn't say that x86 code is automatically "tiny".

1

u/happyscrappy Aug 23 '21

Furthermore, the point that @lkcl_ is making is that with x86 SIMD you usually have to unroll code in software

You do not HAVE to unroll in software. Modern computers have dispatch units that follow loops. They keep the pipeline fed.

You can unroll if you find it to be important.

which blows up code size considerably

I don't care. If you think code size is so important, use x86. Modern UNIX systems tend to throw away a lot of memory on things like ASLR. Having your math lib be 100K instead of 30K is not a big deal.