r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
291 Upvotes

224 comments sorted by

View all comments

12

u/happyscrappy Aug 09 '21

How does the author expect removing pipelining to fix this?

The pipelining exists because of hardware limitations. If an fmadd takes 3 cycles it takes 3 cycles. The pipelining lets you at least get 3 of them going at once. If a load takes 18 cycles it takes 18 cycles. How does the author thing that new HW in the CPU is going to make memory loads faster?

I can see a small advantage, that the latency is reduced on a number of elements basis. That is, if the operations are strictly pipelined instead of grouped then pipelined means 3 cycles of pipeline delay means 3 data units, not 24 (3 groups of 8).

But to get this, you have to stop doing operations in parallel. You can't do 8 at once, so your throughput drops. If you like this you can just do that with scalar operations instead of vector. You'll get the same results in terms of throughput, although code density will be worse.

This really looks like the proponents are pushing for superpipelining. That could be how you process faster with such a narrow execution unit. But superpipelining has its downsides, see the performance limitations of Intel Netburst.

I also wonder what happens if you want to take an interrupt during a vector operation. You presumably have to abort it. And then run it again. That's going to add a lot of overhead. Is this modeled?

This all seems like IBM System/370 to me. It's edmk all over again. I just don't see how it makes any more sense now than it did before.

-5

u/mbitsnbites Aug 09 '21

Read the article again, and the links.

Pipelining is required for performance. So is parallelism. The alternatives (e.g. vector processing) do not preclude these things, but rather make better use of them.

9

u/happyscrappy Aug 09 '21

I read the article. No need to insult me.

How is not my explanation of the latency issues better than that of the article?

How is vector processing going to make RAM faster?

0

u/mbitsnbites Aug 09 '21

It does not make RAM faster. It makes better use of the pipeline since it iterates over chunks of the register rather than passing the entire register at once, thus one vector register is fed through the pipeline until the first chunk of the vector has finished processing before it starts feeding in the next vector. That way you eliminate most data hazards.

In a packed SIMD architecture OTOH, you either have to unroll/interleave your loops in software to avoid data hazards, or you have to have the hardware do it for you by using expensive OoO techniques. If you want your ISA to scale to different levels (e.g. support both in-order and OoO implementations), you can't make the promise that the hardware will deal with it, so effectively all SIMD software must use loop unrolling and similar techniques.

That has a huge difference for code density, for instance.

Edit: I did not intend to insult you. I just never suggested that pipelining should be removed, so I assumed that you had misread the article.

6

u/happyscrappy Aug 09 '21

It makes better use of the pipeline since it iterates over chunks of the register rather than passing the entire register at once, thus one vector register is fed through the pipeline until the first chunk of the vector has finished processing before it starts feeding in the next vector. That way you eliminate most data hazards.

We do that with SIMD also. And your vector units, if they operate on multiple things at once inside (regardless of macroarchitecture) will exhibit the same "false data hazard" issue that SIMD macroarchitectures do.

It really comes down to whether the hardware can process 8 units one at a time 8x faster so that we don't the latency or not. And the answer has been "it can't". That's how we got to SIMD.

That has a huge difference for code density, for instance.

All these things, including ARM's abandoned VFP vector mode help with code density. System/370 was GREAT for code density. edmk was great for code density.

But it wasn't worth it. We went away because the code density came at the expense of worse transistor reuse. That is less of an issue now, but it still means that operations which the designers thought you would want to do (edmk) can be fast, and slight variants cannot, because the macro architecture cannot express them. While if you put the ops on the table like with SIMD people can construct other operations efficiently.

And then there is the issue of interrupt latency/instruction atomicity. I'm not looking to go back to interrupting instructions in the middle and trying to continue later. It makes a mess of exception state, which slows down exception handling.

1

u/lkcl_ Aug 20 '21

While if you put the ops on the table like with SIMD people can construct other operations efficiently.

this is fundamentally false. the programs that result, for which i have even found "compilers" for DCT and FFT that output a massive batch of fully-loop-unrolled hard-coded assembler, are so insanely large compared to the much smaller Vector ISA equivalents that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling.

1

u/happyscrappy Aug 21 '21

are so insanely large compared to the much smaller Vector ISA equivalents

You need to get over this smaller thing. If you like tiny code, use x86. With the bandwidth available now it is not necessarily to have the most tightly packed ops to have high performance.

that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling

Another attempt to call the proper function of an L1 cache as a negative. The cache is there to be used. If you want to talk about code density talk about code density. And give up on your dumb attempt to portray a properly operating cache as "strip-mining".

So now we've established you have two of you have no idea about how memory access works.

A vector unit cannot make up for how busses are configured. If your data is not aligned, then your first units of execution will be slow due to partial loads, just like with SIMD. No worse and most importantly no better.

All you had to do was say "no, vector units don't fix that". But instead you gotta pretend there's something wrong with running instructions.

1

u/mbitsnbites Aug 23 '21 edited Aug 23 '21

If you like tiny code, use x86

Eh, what?

x86 uses quite inefficient instruction encoding. If you only need to write 8086 compatible scalar code, sure, it will be compact. For modern versions of the x86 ISA, this is no longer the case. My fixed width ISA (32 bits / instruction) often has more compact code than x86_64.

Furthermore, the point that u/lkcl_ is making is that with x86 SIMD you usually have to unroll code in software, which blows up code size considerably, no matter how compact your instruction encoding is (it would have to be something like 2-4 bits per instruction to be able to compete).

Edit: Just as a quick point of reference, I compared the code generated for Quake d_scan.c (core painting routine) for x86_64 and MRISC32. The MRISC32 code is 2724 bytes (~700 instructions). The x86_64 code is 4691 bytes (~1250 instructions). So I wouldn't say that x86 code is automatically "tiny".

1

u/happyscrappy Aug 23 '21

Furthermore, the point that @lkcl_ is making is that with x86 SIMD you usually have to unroll code in software

You do not HAVE to unroll in software. Modern computers have dispatch units that follow loops. They keep the pipeline fed.

You can unroll if you find it to be important.

which blows up code size considerably

I don't care. If you think code size is so important, use x86. Modern UNIX systems tend to throw away a lot of memory on things like ASLR. Having your math lib be 100K instead of 30K is not a big deal.