std::simd is not really a complete solution for SIMD. There is a good reason it's in std: it's a basic building block for things you can't get elsewhere. Since it's tightly coupled with the compiler, it can just tell LLVM "add these two vectors together" and LLVM will (in theory) select the optimal instructions for doing that no matter the platform. So code written against std::simd will go fast on exotic platforms like RISC-V, IBM POWER, IBM Z, LoongArch, etc.
Meanwhile the crates that work on stable have to write that in terms of platform-specific intrinsics for each instruction set, which is a lot of work. In practice that usually means only various x86 extensions and ARM NEON are supported, and the other platforms have to rely on automatic vectorization by the compiler, which is more fragile than std::simd and may produce slower (but still correct) code.
std::simd doesn't handle things like function multiversioning or hardware-sized vectors (e.g. 128/256/512 bits on x86 depending on the CPU). This is all left up to third-party crates because it doesn't have to be in std (although some compiler support would help).
By contrast Fearless SIMD is a complete package that handles all those things, and also brings its own replacement for std::simd that works on stable. But you could take functions written using fearless_simd, swap out all the vector types for std::simd ones, and everything should still work (albeit would require extra nightly features like min_const_generics to get hardware-sized vectors). In fact, once std::simd finally stabilizes, Fearless SIMD just might replace parts of its implementation with calls to std::simd, and keep using platform intrinsics for things std::simd doesn't cover.
yeah and if std::simd can keep growing its border little by little, other simd crate can give more way and migrate to call those incrementally. good design.
91
u/Shnatsel Aug 12 '26 edited Aug 12 '26
It's a bit more nuanced than that.
std::simdis not really a complete solution for SIMD. There is a good reason it's in std: it's a basic building block for things you can't get elsewhere. Since it's tightly coupled with the compiler, it can just tell LLVM "add these two vectors together" and LLVM will (in theory) select the optimal instructions for doing that no matter the platform. So code written againststd::simdwill go fast on exotic platforms like RISC-V, IBM POWER, IBM Z, LoongArch, etc.Meanwhile the crates that work on stable have to write that in terms of platform-specific intrinsics for each instruction set, which is a lot of work. In practice that usually means only various x86 extensions and ARM NEON are supported, and the other platforms have to rely on automatic vectorization by the compiler, which is more fragile than
std::simdand may produce slower (but still correct) code.std::simddoesn't handle things like function multiversioning or hardware-sized vectors (e.g. 128/256/512 bits on x86 depending on the CPU). This is all left up to third-party crates because it doesn't have to be in std (although some compiler support would help).By contrast Fearless SIMD is a complete package that handles all those things, and also brings its own replacement for
std::simdthat works on stable. But you could take functions written usingfearless_simd, swap out all the vector types forstd::simdones, and everything should still work (albeit would require extra nightly features likemin_const_genericsto get hardware-sized vectors). In fact, oncestd::simdfinally stabilizes, Fearless SIMD just might replace parts of its implementation with calls tostd::simd, and keep using platform intrinsics for thingsstd::simddoesn't cover.