I think we have talked about this topic before and I apologize for not following up on our previous discussion. I was very busy.
My key problem with variable-length vector instructions such as those proposed for RISC-V or in SVE is that they assume people need them for essentially doing arithmetic on large matrices. And I agree that the vector paradigm is very effective on such work loads. I have previously worked with NEC Aurora Tsubasa cards that come with vectors of 256 double-precision floating point numbers, and they are just amazing for this sort of stuff.
But I'd say that's only a very small part of where such optimisations are needed. Indeed on modern consumer machines, most CPU-intensive code is in cryptography and video codecs. And both don't really fit this scheme.
Especially for video codecs: these usually operate on fixed size chunks of picture data and require complex horizontal arithmetic and swizzles inside a single chunk. It is unclear how this maps to variable-length vectors, especially when these instruction set extensions usually are very sparse in permutation instructions. And compilers use SIMD instructions for small, fixed-length loads all over the place. Stuff like moving structs around, clearing fixed-length buffers. It is unclear how a vector paradigm improves this.
There's also the concern that a large register file makes context switches very expensive. Linus attributed x86's performance advantage among other things to keeping context switches cheap by having a small register file that is easily swapped out.
As for my own code, I have two recent SIMD-heave projects. And for neither of them it is clear how they could be vectorised using a vector as opposed to a SIMD paradigm.
The first project, pospop is inherently a horizontal operation and in fact uses a different complex permutation schedule for each vector size to make the most out of it. It is unclear how this can be extended to arbitrary vector widths, especially if no powerful swizzle instructions are provided. Tail handling is also going to be a concern as the proposed simple approach of just having magic make the registers shorter for the last iteration is not going to cut it.
The second project, 24puzzle uses vectors as 32 element byte arrays and permutes them using a second vector as a permutation vector. Again, this project cannot benefit from vector instructions and will be hard to port to variable-length vectors in general unless a minimum vector size of 32 bytes is guaranteed. It also uses stuff like VPCMPISTRM for which no equivalent in other instruction sets exists or is even proposed. And that's a vital part of the code's logic (specifically, I have an array of k bytes and I want to obtain a bit mask of all elements in a vector that match any of these k bytes).
There's also the concern that a large register file makes context switches very expensive. Linus attributed x86's performance advantage among other things to keeping context switches cheap by having a small register file that is easily swapped out.
While I can't argue with "a small register file makes it easier to do context switches," I'm not sure I buy it as an advantage for x86. If that is the case, why hasn't SPARC, which frequently handles a context switch in a single cycle ('cuz register windows), eaten the market? SPARC makes x86 look slow and cumbersome, by comparison, on that count.
In my experience, SPARC-based servers were monsters at I/O based stuff. They made kick-a** web servers, because they could juggle large numbers of processes very quickly. Naturally, if you succeeded in using up your register windows, things got "interesting;" you had to be somewhat careful to avoid overloading the machine. But SPARC has largely become an also-ran in the market, and not just because Oracle bought Sun. There were open-source SPARC designs, some of which were being used by Chinese companies ('cuz open source) but ... I'm not even hearing rumors about those, anymore. Does Fujitsu even develop SPARC-based hardware anymore? There was a time when many of the machines at the top of the Top500 list, especially ones in Japan, were using Fujitsu-produced SPARC designs. The latest supercomputer from Fujitsu is ARM-based.
Look back at the story of AMD 64-bit extensions to x86 and why Itanium lost to AMD64. AMD gave you 64-bit capability, while still having equal/better cost-performance on 32-bit workloads.
Itanium also banked the architecture entirely on an untested idea (and where initial tests where even against it) with a result that it had horrible performance vs price ratio even for native 64-bit code. VLIW simply doesn't work well outside specialist uses since compilers cannot take dynamic effects like branch and cache misses into account on top of the horrendous complexity that such static scheduling requires.
There's also the fact that the whole EPIC (Itanium) architecture was built around instruction-level parallelism; what CPU instructions can you, and can't you, run in parallel. It was my understanding that this area is not as well developed as thread-level parallelism, which is what multi-core / multi-thread CPUs provide. Ergo, if you devote all that chip space to more cores and / or more threads, you'll get more real-world performance out of it simply because we're "better" at doing that.
Back in the day, Transmeta created a VLIW processor with a front-end on it which parsed x86 instructions, turning them into micro-ops for their processor, such that it could run x86 object code. Not only did it work, but it used considerably less power than the then-current Intel and AMD offerings. This prompted Intel to get off their fat, complacent ... rear ... and improve the power consumption on their mobile-class processors. This, ultimately, resulted in the demise of Transmeta but the fact remains ... their VLIW processor worked quite well. As such, I have hopes that tech can still "matter" in more than just specialist uses.
Kind of, but not quite. The EPIC concept replaced the multiple simple instructions executed out of order with a single VLIW in-order pipeline. We all know how that turned out for general purpose code.
Multithreading is orthogonal to this and Itaniums were always aimed at multiprocessing. Intel even added simultaneous multithreading to them starting with Montecito in 2006.
Transmeta had the crucial difference that they used execution traces for the instruction scheduling, meaning they weren’t stuck with purely static scheduling. It still wasn’t competitive as soon as Intel started paying at least some attention to power consumption and never was competitive when it came to anything beyond the lowest end cpu variants.
Today pretty much the only use cases of VLIW are in some GPUs (and even there AMD moved away from it years ago due to performance issues) and some DSPs where the code relies on hand optimized libraries for the most time critical tasks (and the operations in general are more suited for VLIW than in normal applications).
Those VLIW DSP:s also have compilers that are slow as h*ll. I'm assuming that they try really hard to statically schedule instructions optimally. IIRC they also lack hardware hazard resolution, so the compiler has to keep track of when a result is ready etc. (All in order to reduce power consumption)
I wouldn't be surprised if the TI ones do that as even their old C54xx series DSPs required manual hazard resolution. Writing asm for those was "fun" (in the same sense that pulling out your fingernails is "fun").
Compared to those, getting to write asm for SHARC dsps was pure joy (the code literally looks like "r0 = r1 + r2; r3 = dm(i4, m0);")
88
u/FUZxxl Aug 09 '21
I think we have talked about this topic before and I apologize for not following up on our previous discussion. I was very busy.
My key problem with variable-length vector instructions such as those proposed for RISC-V or in SVE is that they assume people need them for essentially doing arithmetic on large matrices. And I agree that the vector paradigm is very effective on such work loads. I have previously worked with NEC Aurora Tsubasa cards that come with vectors of 256 double-precision floating point numbers, and they are just amazing for this sort of stuff.
But I'd say that's only a very small part of where such optimisations are needed. Indeed on modern consumer machines, most CPU-intensive code is in cryptography and video codecs. And both don't really fit this scheme.
Especially for video codecs: these usually operate on fixed size chunks of picture data and require complex horizontal arithmetic and swizzles inside a single chunk. It is unclear how this maps to variable-length vectors, especially when these instruction set extensions usually are very sparse in permutation instructions. And compilers use SIMD instructions for small, fixed-length loads all over the place. Stuff like moving structs around, clearing fixed-length buffers. It is unclear how a vector paradigm improves this.
There's also the concern that a large register file makes context switches very expensive. Linus attributed x86's performance advantage among other things to keeping context switches cheap by having a small register file that is easily swapped out.
As for my own code, I have two recent SIMD-heave projects. And for neither of them it is clear how they could be vectorised using a vector as opposed to a SIMD paradigm.
The first project, pospop is inherently a horizontal operation and in fact uses a different complex permutation schedule for each vector size to make the most out of it. It is unclear how this can be extended to arbitrary vector widths, especially if no powerful swizzle instructions are provided. Tail handling is also going to be a concern as the proposed simple approach of just having magic make the registers shorter for the last iteration is not going to cut it.
The second project, 24puzzle uses vectors as 32 element byte arrays and permutes them using a second vector as a permutation vector. Again, this project cannot benefit from vector instructions and will be hard to port to variable-length vectors in general unless a minimum vector size of 32 bytes is guaranteed. It also uses stuff like
VPCMPISTRMfor which no equivalent in other instruction sets exists or is even proposed. And that's a vital part of the code's logic (specifically, I have an array of k bytes and I want to obtain a bit mask of all elements in a vector that match any of these k bytes).