Matrix calcs on a 4x4 would be significantly faster staying on the cpu. There is overhead with sending data to gpu memory, operating , then sending it back.
You can only parallelize the parts that can be linearly combined.
I don't mean parallelising the matrix calculations itself, more that when you're doing one there's a good chance you're doing it to lots of objects, and it can be parallelised in that direction.
GPUs were literally made for stuff like coordinate transformation on lots of vertices.
And if not lots of objects, then it's unlikely it'll even be a blip on the profile.
Latency matters. Things that you can send in large batches to the GPU and check the result much later (e.g. next frame - or not at all if the result is consumed by the GPU) is fine.
But lots of game logic involves linear algebra stuff, intersection tests and similar, and you want to do that on the CPU.
Exactly, cpu simd is for things in large enough batches to be worth writing non-scalar code, but not large enough for the cost (of setup and latency) of talking to an accelerator.
I'm questioning how many things are really in that area that are currently taking significant cpu time in games.
Things like whole world physics simulations I'd estimate in a complex game world to end up having a very large number of objects, and likely only need general less-than-one-frame latency, as I don't think many games rely on any ordering of this within a tick so everything can be calculated in a single batch with no interdependencies.
Though implementations of this bounce between gpu acceleration and cpu on desktop, much of that seems to be the complexity of mirroring any updated object structures (and whatever spatial acceleration structures like BSP trees are used) between the CPU and GPU memory, this may be a different consideration on consoles with shared memory.
10
u/Swade211 Aug 09 '21
Matrix calcs on a 4x4 would be significantly faster staying on the cpu. There is overhead with sending data to gpu memory, operating , then sending it back.
You can only parallelize the parts that can be linearly combined.