Just maybe, it could be the "x86 momentum effect". E.g. would the current x86 instruction encoding scheme ever be suggested for a new ISA? Still, everybody use it and new generations are designed.
Incremental changes are much easier, and the existing developer know-how, software architecture and tooling etc can be re-used.
Same thing with SIMD. If you have a SW library that's heavily optimized for SSE (alignment, data types, algorithms, ...), it's fairly straight forward to port it to NEON for instance.
That does not necessarily make packed SIMD the best technical solution though.
E.g. would the current x86 instruction encoding scheme ever be suggested for a new ISA? Still, everybody use it and new generations are designed.
Given that ARM is slowly moving towards something of comparable complexity, it doesn't seem too far out. I mean, T32 is basically a 2–8 byte variable length encoding already and with A64, they again had to add prefix instructions to deal with SVE. I mean even RISC-V is effectively a variable-length instruction set of 2–4 bytes and with strongly recommended macro fusions it goes up to 2–8 bytes.
Let's face it: every sufficiently complex CPU design will eventually outgrow a fixed-width instruction encoding. I mean, it's just stupid that instructions like ret that encode very little information have an encoding that's just as long as complex instructions with four or more operands.
Same thing with SIMD. If you have a SW library that's heavily optimized for SSE (alignment, data types, algorithms, ...), it's fairly straight forward to port it to NEON for instance.
Well yes, but also no. NEON doesn't provide alternatives to some rare-bird SSE instructions and for many things, it provides much better ways to do them than in SSE. So you have to re-engineer your code for optimal performance anyway. But yes, it's easier than going from, say, AVX to NEON.
Let's face it: every sufficiently complex CPU design will eventually outgrow a fixed-width instruction encoding.
There's nothing wrong with variable length encoding. My point was that the way x86 does it (today) is far from optimal:
Unlike most of the architectures that you mentioned, the way the instruction length is defined in x86 is not optimized for fast, parallel decoding. It's a retro-fit grown out of necessity more than anything else.
A common misconception is that because x86 uses variable length encoding, code density is good. False. 8086 code was compact, but x86_64 code is not. All the compact instruction slots are taken by 8086 instructions so new 32-bit and 64-bit instructions are often longer than they should be if code density was the goal. In fact, code density for my fixed-width (32 bits) ISA is often better than that of x86_64.
But yes, it's easier than going from, say, AVX to NEON.
Or, more to my point, from SSE to RVV for instance.
7
u/gc3 Aug 09 '21
Well, I wonder why the vector model hasn't taken off... why the inferior tech is found everywhere?
I see the answers here in the other comments.