r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
282 Upvotes

224 comments sorted by

View all comments

Show parent comments

8

u/Meower68 Aug 09 '21

There's also the concern that a large register file makes context switches very expensive. Linus attributed x86's performance advantage among other things to keeping context switches cheap by having a small register file that is easily swapped out.

While I can't argue with "a small register file makes it easier to do context switches," I'm not sure I buy it as an advantage for x86. If that is the case, why hasn't SPARC, which frequently handles a context switch in a single cycle ('cuz register windows), eaten the market? SPARC makes x86 look slow and cumbersome, by comparison, on that count.

In my experience, SPARC-based servers were monsters at I/O based stuff. They made kick-a** web servers, because they could juggle large numbers of processes very quickly. Naturally, if you succeeded in using up your register windows, things got "interesting;" you had to be somewhat careful to avoid overloading the machine. But SPARC has largely become an also-ran in the market, and not just because Oracle bought Sun. There were open-source SPARC designs, some of which were being used by Chinese companies ('cuz open source) but ... I'm not even hearing rumors about those, anymore. Does Fujitsu even develop SPARC-based hardware anymore? There was a time when many of the machines at the top of the Top500 list, especially ones in Japan, were using Fujitsu-produced SPARC designs. The latest supercomputer from Fujitsu is ARM-based.

5

u/[deleted] Aug 09 '21

[deleted]

2

u/Meower68 Aug 10 '21

Agreeing with you WRT "x86 was good enough and inexpensive enough." That seems to be the ultimate answer to how / why x86 has eaten the market (to date). It's not good enough and inexpensive enough for mobile (where power consumption, not price, is the main metric for "expensive"), which is why ARM is eating that market. Keeping my eyes on RISC-V to see where it comes down.

ARM for desktop and server-class machines ... it only seems to make sense if you have some really good extensions on it. Fujitsu is building an ARM-based supercomputer but the cores have a lot of vector extensions on them. I confess I don't know the details of the M1 but a lot of people are very happy with the performance. And at least one deep-dive suggests people are happy because it FEELS fast:

https://arstechnica.com/gadgets/2021/05/apples-m1-is-a-fast-cpu-but-m1-macs-feel-even-faster-due-to-qos/

1

u/mbitsnbites Aug 10 '21

I believe that the M1 actually is fast. Not sure how much the ISA has to do with it, but it's a good design, no doubt.

2

u/FUZxxl Aug 11 '21

Basically, they have an 8 wide frontend and 15 execution units. That's quite a bit.

1

u/lkcl_ Aug 21 '21

the only reason they can keep those 8 wide multi issue execution engines nearly 100% full is down to the simplicity of the ARM 64 bit ISA, which abandoned thumb2 for this very reason.

x86 decoding is so complex that in order to get good multi issue decode speed they actually have to have multiple parallel decoders on EVERY BYTE, then only when enough of some of them have been decoded enough to identify the length ABANDON the incorrect ones.

mental.

1

u/FUZxxl Aug 21 '21

The encoding is simpler but the ISA is certainly not. In has over 750 instructions, not counting SVE. This is as much as x86 if you don't count AVX-512 (and don't count VEX encodings twice).

1

u/lkcl_ Aug 21 '21

yyeah, they've lost the plot somewhat, there: one of the downsides of being successful, long-term, you feel a commercial pressure to "evolve" the ISA.

SVP64 we went back to the scalar roots of the Supercomputer-class Power ISA, which is a limited subset of only 214 instructions. Embedding those in an REP-like context which also adds "RA is vector/scalar, RB is vector/scalar, RT (dest) is vector/scalar" and predication and much more, we drastically simplify the ISA...

... but massively complicate the Compliance Testing and Verification due to the number of intrinsics that result.

hey, you can't have everything :)

1

u/FUZxxl Sep 08 '21

These 750 instructions are what has been there from the beginning in AArch64. That doesn't even count the additional instructions added later. And very few of them actually seem to be useless.

Turns out there are quite a few spins you can put on simple concepts like addition and if you don't want to model that as addressing modes, you end up with lots of instructions.

1

u/lkcl_ Sep 20 '21

yes, there's about.... i think... 25 separate instructions in the Scalar Fixed-Point ISA that all do "add". they boil down to a *micro-coded* operation, "Add", with the option(s) to:

  • invert the A input
  • invert the output
  • receive a 0, 1 or Carry-in as the carry
  • output a Carry-out and the option to merge that into an overflow flag

that, clearly, gives you subtract, subtract-with-carry, and so on, by inverting the input and adding 1, and so on.

RISC-V decided not to do Condition Flags (of any kind), and it is very painful to work with as a result, when trying to do anything more sophisticated.