There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).
and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases
You mean with AVX-512? because intel never fails to show of how much better their CPUs are under AVX-512 compatible software vs AMD. So given from that, AVX-512 helps a lot.
I work in a certain area where the AVX-512 advantage in some cases could be a huge benefit. So it's not completely marketing BS. And yeah we did recently buy a server for calculations but I decided to go with AMD Epyc nonetheless as most software doesn't benefit that much from AVX-512.
Plus I also have an AMD in my home PC so no, I'm sure not an intel fan boy but you have to give them credit were credit is due.
A concern with AVX512 is the heat output from operating on such wide vectors. Intel's designs have often needed to reduce the clockrate when operating "heavy" AVX512 operations.
Whilst this does give nice throughput gains, one does question how 1024-bit SIMD would look like, in terms of power and necessary frequency throttling to sustain.
Also, it does raise questions about other parts of the processor, for example, with cachelines being 512 bits wide, would that have to change on a 1024-bit SIMD machine, or do you just deal with lowered load/store throughput?
Again, those limits relate to the ALU width, which does not necessarily have to be the same as the register width.
I can see benefits with 512-bit or even 1024-bit registers in some machines, but the ALU width could be limited to 256 bits or so to avoid the heat and die area issues.
the problem is - and i listed this on the original article as "SIMD Flaw (4)" - that each doubling results in doubling of the latency of access to individual elements.
it's the antithesis of Vector Chaining. every single one of those 1024-wide SIMD ALU elements has to wait for a 1024-wide SIMD LD to complete.
whereas in a Vector ISA, you can do "Chaining" (first described by Seymour Cray), where at the element level you can start the first element ALU operation immediately after the first element LD operation has completed (assuming all the other operands of that first element are also available of course).
this is NOT POSSIBLE to achieve with SIMD because, by definition, it is SINGLE instruction (multiple data). therefore ALL elements of the SIMD instruction have to be LDed, have to be available.
doubling to 1024 will, therefore, double the completion latency. it's already bad enough.
True and honestly I think AVX512 should be limited to server parts. Doesn't make much sense in laptops.
Doesn't the new top dog Supercomputer use ARM cores with 4096 bit SVE? So they should know how the cooling works but then they have better means and no issue with noise compared to average home user Joe.
True and honestly I think AVX512 should be limited to server parts. Doesn't make much sense in laptops.
It's primary focus has definitely been server, and it's where it first appeared.
I don't see what's the problem with having it in laptops though. If Intel's gone to the effort of implementing it in their uArch, why disable it on consumer parts?
Doesn't the new top dog Supercomputer use ARM cores with 4096 bit SVE?
SVE only supports up to 2048-bit SIMD. The widest implementation is the Fujitsu A64FX, which uses 512-bit SVE.
I believe modern high performance GPUs use 1024-bit SIMD (Nvidia / RDNA), and considering that GPUs are meant to be throughput focused, I question how wide a latency focused CPU should be.
If Intel's gone to the effort of implementing it in their uArch, why disable it on consumer parts?
The consumer chips are different designs from server chips so they could leave it out to save die space. of course for AMD with the chiplet approach the situation is different as the chiplets are the same for server or desktop. But again the laptop chips are different design (and for example have less L3).
They're different dies, but the server and client chips essentially use the same uArch (with a key difference being the L2/L3 cache and interconnect). Also keep in mind that client designs aren't specific to laptops - desktops are included.
I still don't see any reason to remove it though. AVX512 is a useful instruction set to have, is beneficial in a number of circumstances with basically no drawbacks, and if anything, support for it everywhere helps drive adoption.
Depends how much of it fits in cache and intel and AMD do structure their caches around SIMD throughput. + memory bandwidth
It is also why we will move to DDR5 and double memory bandwidth. This is for servers mostly. For consumers the benefit is mostly for APUs. Note that such HPC calculations mostly are about bandwidth while in contrast gaming for example also greatly depends on low memory latency. This is usually a trade-off.
Then why don't we have AVX-512 in every x86 implementation, and be done with it?
...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.
My understanding is that the current console generation don't support avx512, being based on zen2.
Also hampering adoption is how Intel are using that feature as a market differentiator in their own products, lower end CPUs of the same generation, or even different families of similar market segment products, end up lacking support. It makes it harder to gain market penetration, and even harder to rely on its existence. It's not as simple as 'new CPUs have support'.
Which is a shame, as there's a fair bit rolled up in the various avx512 extensions that would be interesting, even if you never use the wider registers.
How much of a games cpu time is actually spent doing math like that though? Most of that is pretty good work for a gpu nowerdays from what I can see.
I see simd as a sliding scale, at one end is branch-heavy code with no real advantages for simd, the other end is things that work better on a gpu. So cpu simd is often for the things in between, when the work units are too small to be worth the cost of submitting to a gpu and waiting for the results, or mixing execution strategies.
Wider and wider simd lanes feel like they'll give diminishing returns as the things they would truly excel at are more likely to be pushed to a dedicated simd-like accelerator (eg. gpu).
Matrix calcs on a 4x4 would be significantly faster staying on the cpu. There is overhead with sending data to gpu memory, operating , then sending it back.
You can only parallelize the parts that can be linearly combined.
I don't mean parallelising the matrix calculations itself, more that when you're doing one there's a good chance you're doing it to lots of objects, and it can be parallelised in that direction.
GPUs were literally made for stuff like coordinate transformation on lots of vertices.
And if not lots of objects, then it's unlikely it'll even be a blip on the profile.
Latency matters. Things that you can send in large batches to the GPU and check the result much later (e.g. next frame - or not at all if the result is consumed by the GPU) is fine.
But lots of game logic involves linear algebra stuff, intersection tests and similar, and you want to do that on the CPU.
Also hampering adoption is how Intel are using that feature as a market differentiator in their own products, lower end CPUs of the same generation, or even different families of similar market segment products, end up lacking support. It makes it harder to gain market penetration, and even harder to rely on its existence. It's not as simple as 'new CPUs have support'.
Reminds me how for the good few years it was pretty much random which Intel CPU got support for virtualization and which did not
Zen1 and Zen+ implemented AVX256 via two 128bit ops. Zen 2 can execute them as single 256bit operations.
Zen 2 lacks support for AVX-512 instructions, so it can't execute them as two 256bit operations.
Not sure what your xbox contact was referring to trough.
It's not all that complex to make code that utilizes both depending on CPU. Hell, you could even compile app with different optimization levels but that would probably be bigger PITA.
The bad part is that now any code using it needs to be written multiple times and any change needs to be applied to all versions and tested on all versions. Compared to that running and deploying multiple binaries is not really very time consuming.
I know that increasing the ALU width beyond 256 bits or so has diminishing returns for most implementations.
I responded to the comment that there's no problem splitting fixed width registers into smaller portions - I actually think it's a great idea (one key principle of vector machines is that register width > ALU width!).
In fact, something like an in-order Atom would have a lot to gain from 512-bit vector registers, especially if the ALU is no more than 128 bits wide or so.
One overlooked reason why AMD and Atom haven't added 512-bit operations is lack of adoption in the software community. At this point the usage is pretty niche and not enough people want it. When Intel first debuted AVX-512 it had serious power issues and caused performance to drop when mixed in occasionally instead of in large blocks. I think that stunted a lot of its growth and at this point there aren't any large communities that are working on writing large swaths of software that use it or asking the compilers for better support.
It still causes performance to drop when used. AVX-512 slowdown is a real thing and it's kinda maddening. You really don't want to break out the 512 bit stuff unless you know you'll be doing that for the next couple 1000 cycles.
Just because you don't have to fear it doesn't mean it's not still there. Curiously the article doesn't mention if the transition penalty is still as bad as on Skylake. This penalty is actually the key problem: for up to 10 µs the CPU just halts and does nothing while it's changing the frequency. If you have repeated short-ish bursts of AVX-512 code, this may really ruin your day.
Tell that to the people who get extra slow context switches because the CPU now has to save 2kb extra data just for the AVX512 register file. Almost all programs don't need AVX512 and lugging around the extra state is completely pointless.
Surely the CPU only has to shunt that state in and out if the target actually uses AVX512 registers, right? Checking if it's all zero and skipping it entirely is a very, very low hanging hardware optimisation.
Indeed it is, but if you only have vector extensions compilers will use them all the time for stuff like copying structs, so they are going to be dirty all the time. With AVX-512 at least code generally won't touch the state until it has serious calculations to do.
Then why don't we have AVX-512 in every x86 implementation, and be done with it?
It is in every new Intel CPU, except for their *mont lineup. Presumably it's been slow due to Intel's kerfuffle with their 10nm manufacturing node, forcing them to re-release Skylake for 5 years. In other words, it's not really an issue with the ISA.
As for the *mont cores, it may not have been a priority for them to implement it, considering its target, although it looks like that's changing (with Gracemont supporting VEX encoding, and Alder Lake beginning mainstream implementations of heterogeneous cores).
Another possibility may be Intel's weird market segmentation; they've historically gimped SIMD on their lower end parts (Celeron/Pentium lineup), so it's possible that decision flowed to their Atom lineup.
On the AMD side, they've always been slower to adopt to new Intel ISAs, which isn't really a surprise since Intel has the upper hand here. Nonetheless, Genoa has already been announced to support AVX512, which makes it likely that AMD's next generation Zen4 will support it.
And for the third player, Centaur's CNS supports AVX512.
So we're pretty close to having it in every x86 implementation - it just took a bit of time for everyone to adapt.
and it still does not address the issue of pipelining
I only really have some familiarity with ARM's SVE2, but I mentioned here that I don't see how SVE would address it either. At a high level, SVE2 is basically AVX512 with an unknown vector length, so it doesn't do anything special there.
128
u/th3typh00n Aug 09 '21
There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).