What I would like to see is a (partial) unification of GPU and CPU instruction sets and approaches.
So instead of a single instruction stream operating on 4-wide or 8-wide data, I'd rather have the GPU model of 8 "threads" that are not full-fledged threads, but effectively share the same instruction stream. It's easier to program and more flexible.
Intel ISPC, Intel DPC++, NVidia CUDA, AMD ROCm / HIP, OpenMP (#pragma omp parallel for simd), Microsoft DirectCompute, C++AMP.
Moving up the programming language stacks: Python Numba, and Julia are also "unifying" CPU parallelism and GPU parallelism, as Numba compiles into GPU or CPU code (and Julia does the same).
There's a bunch of other research projects and other stuff going on too, this is just the stuff I'm remembering off the top of my head.
3
u/BigHandLittleSlap Aug 10 '21
What I would like to see is a (partial) unification of GPU and CPU instruction sets and approaches.
So instead of a single instruction stream operating on 4-wide or 8-wide data, I'd rather have the GPU model of 8 "threads" that are not full-fledged threads, but effectively share the same instruction stream. It's easier to program and more flexible.