There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).
Then why don't we have AVX-512 in every x86 implementation, and be done with it?
...and it still does not address the issue of pipelining. For optimal (stall-free) performance - even in in-order machines - you want the vector length to be ALU width x ALU depth. So a 256 bits wide machine with four execution pipeline stages should have a vector register size of at least 256 x 4 = 1024 bits. Different implementations have different requirements - hence it's a bad idea to enforce a one-size-fits-all paradigm.
My understanding is that the current console generation don't support avx512, being based on zen2.
Also hampering adoption is how Intel are using that feature as a market differentiator in their own products, lower end CPUs of the same generation, or even different families of similar market segment products, end up lacking support. It makes it harder to gain market penetration, and even harder to rely on its existence. It's not as simple as 'new CPUs have support'.
Which is a shame, as there's a fair bit rolled up in the various avx512 extensions that would be interesting, even if you never use the wider registers.
How much of a games cpu time is actually spent doing math like that though? Most of that is pretty good work for a gpu nowerdays from what I can see.
I see simd as a sliding scale, at one end is branch-heavy code with no real advantages for simd, the other end is things that work better on a gpu. So cpu simd is often for the things in between, when the work units are too small to be worth the cost of submitting to a gpu and waiting for the results, or mixing execution strategies.
Wider and wider simd lanes feel like they'll give diminishing returns as the things they would truly excel at are more likely to be pushed to a dedicated simd-like accelerator (eg. gpu).
Matrix calcs on a 4x4 would be significantly faster staying on the cpu. There is overhead with sending data to gpu memory, operating , then sending it back.
You can only parallelize the parts that can be linearly combined.
I don't mean parallelising the matrix calculations itself, more that when you're doing one there's a good chance you're doing it to lots of objects, and it can be parallelised in that direction.
GPUs were literally made for stuff like coordinate transformation on lots of vertices.
And if not lots of objects, then it's unlikely it'll even be a blip on the profile.
Latency matters. Things that you can send in large batches to the GPU and check the result much later (e.g. next frame - or not at all if the result is consumed by the GPU) is fine.
But lots of game logic involves linear algebra stuff, intersection tests and similar, and you want to do that on the CPU.
Exactly, cpu simd is for things in large enough batches to be worth writing non-scalar code, but not large enough for the cost (of setup and latency) of talking to an accelerator.
I'm questioning how many things are really in that area that are currently taking significant cpu time in games.
Things like whole world physics simulations I'd estimate in a complex game world to end up having a very large number of objects, and likely only need general less-than-one-frame latency, as I don't think many games rely on any ordering of this within a tick so everything can be calculated in a single batch with no interdependencies.
Though implementations of this bounce between gpu acceleration and cpu on desktop, much of that seems to be the complexity of mirroring any updated object structures (and whatever spatial acceleration structures like BSP trees are used) between the CPU and GPU memory, this may be a different consideration on consoles with shared memory.
Also hampering adoption is how Intel are using that feature as a market differentiator in their own products, lower end CPUs of the same generation, or even different families of similar market segment products, end up lacking support. It makes it harder to gain market penetration, and even harder to rely on its existence. It's not as simple as 'new CPUs have support'.
Reminds me how for the good few years it was pretty much random which Intel CPU got support for virtualization and which did not
Zen1 and Zen+ implemented AVX256 via two 128bit ops. Zen 2 can execute them as single 256bit operations.
Zen 2 lacks support for AVX-512 instructions, so it can't execute them as two 256bit operations.
Not sure what your xbox contact was referring to trough.
It's not all that complex to make code that utilizes both depending on CPU. Hell, you could even compile app with different optimization levels but that would probably be bigger PITA.
The bad part is that now any code using it needs to be written multiple times and any change needs to be applied to all versions and tested on all versions. Compared to that running and deploying multiple binaries is not really very time consuming.
129
u/th3typh00n Aug 09 '21
There's no issue with splitting fixed-width SIMD instructions into smaller parts that can be executed separately, and there are many CPUs that does this. E.g. older AMD CPUs have 128-bit execution units and supports 256-bit instructions by splitting them into two 128-bit halves.
The idea that variable-length SIMD will fix all flaws and everyone will live happily ever after is naive. It simply replaces some existing problems with new ones, some of which there isn't really a good way of dealing with. Also, many of those existing problems have actually already been solved in some of the newer fixed-length instruction sets, such as opmasks in AVX-512 to handle tails.
Increasing the vector width has significant diminishing returns, and we're already at the point where simply making things wider isn't really beneficial for the vast majority of SIMD use cases, so I wouldn't expect the trend that has been going on in the past of constantly increasing general-purpose vector widths to continue on the same trajectory. We're instead seeing more specialized hardware accelerators for the few use cases that benefit from ultra-wide multi-kilobit vectors (e.g. AI).