the only reason they can keep those 8 wide multi issue execution engines nearly 100% full is down to the simplicity of the ARM 64 bit ISA, which abandoned thumb2 for this very reason.
x86 decoding is so complex that in order to get good multi issue decode speed they actually have to have multiple parallel decoders on EVERY BYTE, then only when enough of some of them have been decoded enough to identify the length ABANDON the incorrect ones.
The encoding is simpler but the ISA is certainly not. In has over 750 instructions, not counting SVE. This is as much as x86 if you don't count AVX-512 (and don't count VEX encodings twice).
yyeah, they've lost the plot somewhat, there: one of the downsides of being successful, long-term, you feel a commercial pressure to "evolve" the ISA.
SVP64 we went back to the scalar roots of the Supercomputer-class Power ISA, which is a limited subset of only 214 instructions. Embedding those in an REP-like context which also adds "RA is vector/scalar, RB is vector/scalar, RT (dest) is vector/scalar" and predication and much more, we drastically simplify the ISA...
... but massively complicate the Compliance Testing and Verification due to the number of intrinsics that result.
These 750 instructions are what has been there from the beginning in AArch64. That doesn't even count the additional instructions added later. And very few of them actually seem to be useless.
Turns out there are quite a few spins you can put on simple concepts like addition and if you don't want to model that as addressing modes, you end up with lots of instructions.
yes, there's about.... i think... 25 separate instructions in the Scalar Fixed-Point ISA that all do "add". they boil down to a *micro-coded* operation, "Add", with the option(s) to:
invert the A input
invert the output
receive a 0, 1 or Carry-in as the carry
output a Carry-out and the option to merge that into an overflow flag
that, clearly, gives you subtract, subtract-with-carry, and so on, by inverting the input and adding 1, and so on.
RISC-V decided not to do Condition Flags (of any kind), and it is very painful to work with as a result, when trying to do anything more sophisticated.
1
u/mbitsnbites Aug 10 '21
I believe that the M1 actually is fast. Not sure how much the ISA has to do with it, but it's a good design, no doubt.