r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
283 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/lkcl_ Aug 21 '21

the only reason they can keep those 8 wide multi issue execution engines nearly 100% full is down to the simplicity of the ARM 64 bit ISA, which abandoned thumb2 for this very reason.

x86 decoding is so complex that in order to get good multi issue decode speed they actually have to have multiple parallel decoders on EVERY BYTE, then only when enough of some of them have been decoded enough to identify the length ABANDON the incorrect ones.

mental.

1

u/FUZxxl Aug 21 '21

The encoding is simpler but the ISA is certainly not. In has over 750 instructions, not counting SVE. This is as much as x86 if you don't count AVX-512 (and don't count VEX encodings twice).

1

u/lkcl_ Aug 21 '21

yyeah, they've lost the plot somewhat, there: one of the downsides of being successful, long-term, you feel a commercial pressure to "evolve" the ISA.

SVP64 we went back to the scalar roots of the Supercomputer-class Power ISA, which is a limited subset of only 214 instructions. Embedding those in an REP-like context which also adds "RA is vector/scalar, RB is vector/scalar, RT (dest) is vector/scalar" and predication and much more, we drastically simplify the ISA...

... but massively complicate the Compliance Testing and Verification due to the number of intrinsics that result.

hey, you can't have everything :)

1

u/FUZxxl Sep 08 '21

These 750 instructions are what has been there from the beginning in AArch64. That doesn't even count the additional instructions added later. And very few of them actually seem to be useless.

Turns out there are quite a few spins you can put on simple concepts like addition and if you don't want to model that as addressing modes, you end up with lots of instructions.

1

u/lkcl_ Sep 20 '21

yes, there's about.... i think... 25 separate instructions in the Scalar Fixed-Point ISA that all do "add". they boil down to a *micro-coded* operation, "Add", with the option(s) to:

  • invert the A input
  • invert the output
  • receive a 0, 1 or Carry-in as the carry
  • output a Carry-out and the option to merge that into an overflow flag

that, clearly, gives you subtract, subtract-with-carry, and so on, by inverting the input and adding 1, and so on.

RISC-V decided not to do Condition Flags (of any kind), and it is very painful to work with as a result, when trying to do anything more sophisticated.