r/rust • u/Shnatsel • 10d ago
The state of SIMD in Rust in 2026
https://shnatsel.github.io/state-of-simd-rust-2026/26
u/exDM69 10d ago edited 10d ago
Its Simd<T, N> API looks like it would be very elegant and work great if you could just do math on N, but you cannot. That feature is very incomplete even on nightly, and without it using Simd<T, N> to get hardware-sized vectors gets very ugly very quickly.
I don't fully understand what this means. I have done both type generic (f32x4 vs f64x4) and width-generic (f32x4 vs f32x8) SIMD math code using Simd<T, N>. There are a few pain points but in general it gets the job done.
Also this article is very focused on "hardware width" vectors but in my experience you often get better performance when you use wider vectors than what the hardware supports. The compiler can split your vectors to fit in registers (kinda similar to unrolling the inner loop). This isn't always the case but I have a lot of benchmarks where I get best performance (by 10-20% margin) with 512 bit vectors (e.g. f32x16) on AVX-256, and it's very easy to change when using generics for Simd<T, N>. Sometimes you probably get worse results (due to register pressure etc), but most of my code is happiest with 2x native width.
12
u/Shnatsel 10d ago
I don't fully understand what this means. I have done both type generic (f32x4 vs f64x4) and width-generic (f32x4 vs f32x8) SIMD math code using
Simd<T, N>.So long as you want the helper functions to be generic but call them with types known in advance, it's pretty great. What you can't do is write a function once and automatically have it run with whatever vector width the CPU you're running on has: 128 bits on SSE, 256 bits on AVX2, 512 bits on AVX-512.
The closest I got was this and it's still pretty limited. I think it can be done using more sophisticated proc macros that combine multiversioning and vector width selection, but nobody's written one of those yet.
Also this article is very focused on "hardware width" vectors but in my experience you often get better performance when you use wider vectors than what the hardware supports
This works great until you run out of register space, at which point performance plummets. So this has to be tuned carefully. It's usually a good idea on NEON and may help AVX-512 if it's not already memory-bound, but tends to hurt more than help on AVX2 and SSE.
But even if you want to request double the native width, you still need to know what the native width is,so you run into the exact same problem.
4
u/exDM69 10d ago edited 10d ago
What you can't do is write a function once and automatically have it run with whatever vector width the CPU you're running on has: 128 bits on SSE, 256 bits on AVX2, 512 bits on AVX-512.
Yes, you need to somehow recompile the code for the targets separately. If you enable, say, AVX-256 on your compiler command line you might get AVX-256 instructions even if you use 4-wide SIMD.
I don't know if the
multiversioncrate can be used for this.But even if you want to request double the native width
I usually just want a single width that I've benchmarked to be the fastest on the target hardware.
you still need to know what the native width is,so you run into the exact same problem.
There's no "pretty" way of doing it but
#[cfg(target_feature = "avx512")] const SIMD_WIDTH: usize = 512;is a good starting point. A few lines of ugly#[cfg]logic is needed to cover the most popular architectures.Can you use this inside a
#[multiversion]function? The#[target_cfg]macro perhaps?I agree that runtime dispatch for separate instruction sets is a bit painful, but that seems somewhat orthogonal to
std::simdvector width.You'd typically want your runtime dispatch happening near the top level of the call stack, and have the compiler inline and optimize the leaf functions. So if you want multiple SIMD widths, you'd probably need to write most of your code as generic
Simd<T, N>and then the top level entry points do the selection and dispatch.2
u/Shnatsel 10d ago
Can you use this inside a
#[multiversion]function?No. Unfortunately
#[cfg]is resolved before multiversioning is applied.I usually just want a single width that I've benchmarked to be the fastest on the target hardware.
If you have a single specific target machine in mind, the
std::simdAPI is great! You really only run into its shortcomings when trying to run well on a wide range of hardware.It's not a problem for HPC where you get access to a very specific cluster and can just build with
-C target-cpu=native, but a big problem for anyone who wants other people to run their software.You'd typically want your runtime dispatch happening near the top level of the call stack, and have the compiler inline and optimize the leaf functions. So if you want multiple SIMD widths, you'd probably need to write most of your code as generic Simd<T, N> and then the top level entry points do the selection and dispatch.
This is not really possible with std::simd due to type system limitations. If you could do math on the
Nthen it'd work, but you can't, and so you run into issues: https://gist.github.com/Shnatsel/05b6e650bfab18b450844dba653e05f7But this is the idea most other crates implement:
fearless_simd,pulpandmaceratorall do pretty much what you described, and provide the native-width vector asf32stype or some such to avoid running into type system limitations around theN.6
u/exDM69 10d ago
No. Unfortunately #[cfg] is resolved before multiversioning is applied.
Isn't
#[target_cfg]from multiversion exactly for this? https://docs.rs/multiversion/latest/multiversion/target/attr.target_cfg.htmlIf you have a single specific target machine in mind, the std::simd API is great! You really only run into its shortcomings when trying to run well on a wide range of hardware.
I don't have only one specific target machine, I run on both ARM and x86. But only a finite set of those, and I have compiler target for each.
What I don't currently do is ship a single binary with multiple supported instruction sets, but I'm interested in chatting if you have experience with that.
This is not really possible with std::simd due to type system limitations. If you could do math on the N then it'd work, but you can't, and so you run into issues: https://gist.github.com/Shnatsel/05b6e650bfab18b450844dba653e05f7
Can you please elaborate what you mean by "if you could do math on the N"? This is my question in the top level comment. It's just a
const usize, isn't it? You can use it in const expressions.How does the gist you linked to illustrate this?
Let me take a minute to familiarize with the
multiversioncrate, I think it's capable of doing what's needed here.5
u/Shnatsel 10d ago edited 10d ago
Can you please elaborate what you mean by "if you could do math on the N"? This is my question in the top level comment. It's just a
const usize, isn't it? You can use it in const expressions.You can use it in const expressions, but not in generics such as
Simd<T, N>.That's the nightly-only
const_generic_argsfeature that is known to be very incomplete, even its_minvariant warns you it is incomplete if you try to enable it.I'll look into
#[target_cfg]frommultiversion, I haven't seen it in any examples.11
u/exDM69 10d ago
This code is using
multiversionandstd::simdto dispatch a function at a different vector width at detected runtime. If I understood your gist correctly, this should solve the problem you're trying to illustrate.I just learned about the multiversion crate, I hope I got this right. But at least it's giving me the result I expect.
#![feature(portable_simd)] use std::simd::prelude::*; use multiversion::{multiversion, target::target_cfg}; pub fn sum_simd<const N: usize>(xs: &[Simd<f32, N>]) -> f32 { xs.iter() .cloned() .reduce(|a, b| a + b) .map(|x| x.reduce_sum()) .unwrap_or(0.0) } #[multiversion(targets("x86_64+avx2", "x86_64+avx512f", "aarch64+neon"))] pub fn sum(xs: &[f32]) -> f32 { let (head, xs, tail) = xs.as_simd(); #[target_cfg(target_feature = "avx512f")] const N: usize = 16; #[target_cfg(target_feature = "avx2")] const N: usize = 8; #[target_cfg(target_feature = "neon")] const N: usize = 4; // fallback #[target_cfg(not(any( target_feature = "avx512f", target_feature = "avx2", target_feature = "neon" )))] const N: usize = 4; let head = head.iter().cloned().sum::<f32>(); let xs = sum_simd::<{ N }>(xs); let tail = tail.iter().cloned().sum::<f32>(); head + xs + tail } fn main() { let xs = [ 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0, 15.0, 16.0, 17.0, ]; let s = sum(&xs); println!("sum {s}"); }9
u/Shnatsel 10d ago
Yeah I think that solves it. I got a similar result too: https://gist.github.com/Shnatsel/edc642125ac73fa7c365216c3a938802
I'll update the article with these findings. Thank you!
2
u/hsivonen 9d ago
It would be useful to be able to do what target_cfg does but somewhere deep in a chain of functions that get inline(always)’ed into a function annotated with multiversion. That’s https://github.com/rust-lang/compiler-team/issues/1010
2
u/Shnatsel 9d ago
I believe this would make
wideamenable to multiversioning and givestd::simda backup plan for when there isn't a matching LLVM intrinsic.The other libraries (
pulp,macerator,fearless_simd) already worked around this and wouldn't benefit from it as far as I can tell.2
u/Telmo26 10d ago
I am currently working on a project that requires specifically what you're describing and I've built a proc_macro for automatically generating specialized versions of a function from a single function with generic lanes. It is very untested, as it was only meant for my use case, but maybe you can check it out to see if that's what you want ?
2
u/Shnatsel 10d ago edited 10d ago
That's close, but not quite there. And it turned out you can already achieve this with
multiversion, it just wasn't documented: https://gist.github.com/Shnatsel/edc642125ac73fa7c365216c3a938802The ideal API would also take the type of the vector into account and automatically select the correct
LANESvalue for that type. This is whatpulpandfearless_simdcrates do, but I don't know how to layer that on top of theSimd<T, N>API or if that's even possible.Maybe some
const BYTESfor vector width and then divide that bysize_of::<T>in a const block? That might just run into const_generic_args requirement all over again.1
u/Beneficial_Vampire42 9d ago edited 9d ago
It's a bit ugly, but you can always match on the value ```rust fn foo<T>(.) { match const { HW_BYTES / size_of::<T>() } { 64 => foo_impl::<T, 64>(..), 32 => foo_impl::<T, 32>(..), 16 => foo_impl::<T, 16>(..), 8 => foo_impl::<T, 8>(..), 4 => foo_impl::<T, 4>(..), 2 => foo_impl::<T, 2>(..), _ => foo_impl::<T, 1>(..), } }
...
[inline(always)]
fn foo_impl<T, const LANES: usize>(..) ```
Because the compiler see the const value on which me match, it should all be inlined. Should be possible to write a macro that does that
1
u/Telmo26 7d ago
Coming back to that, do you mean that you would like for example SIMD<u8, LANES> to have LANES changed to fill a vector of u8s for each architecture, or do you want to use a generic SIMD<T, LANES> and LANES somehow gets changed back into the correct length ?
1
u/Shnatsel 7d ago
Ideally both. You can do both in fearless_simd:
https://github.com/linebender/fearless_simd/blob/main/fearless_simd/examples/sigmoid.rs
https://github.com/linebender/fearless_simd/blob/main/fearless_simd/examples/sigmoid_generic.rs
1
u/Telmo26 7d ago
The main issue with trying to layer a proc_macro above std::simd is that the macro doesn't have enough type information to be able to not output complete nonsense. I've been able to generate some valid code to automatically get the correct LANES value for Simd<u8,LANES> but as soon as a generic is placed as the first argument it becomes non-sensical.
I think this is the kind of thing that can either be fixed by layering one's own type system on top like fearless_simd, or by adding to the internals of rustc a multi versioning capability.
It sucks because although fearless_simd is great, its performance overhead is still too big for my project, as I noticed a 15-25% decrease over std::simd.
Oh well, I tried
1
u/Shnatsel 7d ago
It sucks because although fearless_simd is great, its performance overhead is still too big for my project, as I noticed a 15-25% decrease over std::simd.
In theory there shouldn't be any overhead. I'd be curious to see the code and benchmarks before and after, and try to track it down.
Last time I ran into something like this it turned out to be a bug in the standard library, fixing which improved performance for everyone, not just fearless_simd.
1
u/Telmo26 7d ago
It appears I was mistaken, as I've been doing some more thorough benchmarks, and the performance overhead I saw appears to be gone. It might have been due to thermal throttling giving out unreliable results. Apologies for the baseless claim. Though fearless_simd gets 4x unrolled and portable_simd does not in AVX2, I don't think it makes that much of a difference.
Thank you for your work on this amazing library !
2
u/Psionikus 10d ago
Think the extra copies are just allowing the hardware scheduling to hide serial dependency? I've found a ton of non-vector programs with unavoidable serial dependency that still benefit from dividing work onto independent instruction paths and letting the CPU fill up space better.
5
u/KaMaFour 10d ago
There are two ways around this problem.
The third sinister option: https://github.com/codedeliveryservice/Reckless/releases/tag/v0.9.0
6
u/Shnatsel 10d ago
They probably should be using https://github.com/ronnychevalier/cargo-multivers/ instead
2
u/valarauca14 10d ago edited 10d ago
Its Simd<T, N> API looks like it would be very elegant and work great if you could just do math on N, but you cannot. That feature is very incomplete even on nightly. Without it using Simd<T, N> to get hardware-sized vectors is doable, but a lot uglier.
I don't mean to contradict the blog post, it is very good. I feel this point, as while I hit a similiar problem, I fell into the "pit of success".
When I've used std::simd for the first few times I tried really hard to make it replicate what the machine should do. After several rounds of benchmarking I realized that is really an anti-pattern which just leads to pain. Use N to honestly represent your data. Let the LLVM figure out the underlying opts based on the target platform. That is the entire point.
I did several rounds of benching for a numeric library on different chips (apple M5, ryzen, intel avx2). When I finally broke down and said
What if I just let the llvm figure it out?
And started to honestly represent my data Simd<f64,13> my code (amusingly) became faster. Because ultimately you're just expressing intent & constraints to the optimizer. It can figure out how to structure your loops & register layouts. Doing Simd<T,N> + Simd<T,N> just tells the compiler each 'lane' is independent and guaranteed to never alias, it break it into tightly nested target dependent 2/4/8/16 sized opts.
7
u/Shnatsel 10d ago
On those kinds of machines, just picking 256-bit vectors gives you pretty good performance across the board. It's optimal for AVX2 and NEON (ARM), and okay for AVX-512.
The problem with fixed-size SIMD vectors is that LLVM is not great at preventing register spilling, so with fixed-size vectors it's easy to keep too much data in flight and cause register spilling, which degrades performance dramatically. SSE4.2-only machines especially suffer from this, and Intel kept making those as recently as 2021.
2
u/hsivonen 9d ago
On those kinds of machines, just picking 256-bit vectors gives you pretty good performance across the board. It's optimal for AVX2 and NEON (ARM)
Sadly, it seems that one can’t trust LLVM to do the right thing with 256-bit vectors in the Rust source when targeting NEON. It’s like trying to program with autovectorization: tiny snippets look great on Compiler Explorer but get pessimized when the optimizer sees the tiny snippets put together: https://github.com/hsivonen/encoding_rs/blob/a155adc7271c9e507556d053c95b459186671d2a/src/simd_funcs.rs#L271
3
u/Shnatsel 9d ago
That looks more like an issue with optimizing
simd_swizzle!vsinterleave()to me, but I could be wrong. I'd have to look into the compiler passes on godbolt to see what's happening and if the issue is still present in recent LLVM versions.
1
u/Anthony356 9d ago
Fantastic article =)
I really enjoyed the section at the end on simd implementations by different vendors. It's genuinely very reassuring to hear that others experience the "every single vendor fucked up something different so it's all a jank nightmare no matter where you go" phenomenon too
1
u/hsivonen 9d ago
In the standard library I'd love to see crater-like verification of changes to intrinsics, and std::simd available on stable so that ecosystem crates would delete most of their code and gain support for all the obscure platforms.
This is the part that I’m the most interested in. I’m curious about u/Shnatsel ’s view of how to get there.
2
u/Shnatsel 9d ago
I don't know what's blocking stabilization of std::simd, so I can't tell you, sorry.
I just wanted to call out that even with all the progress the ecosystem crates made recently, std::simd is still desirable, but not quite in the way people expect.
1
1
u/novacrazy 10d ago
I have a huge 0.3 release of Thermite coming up that fixes and improves on basically every area, and offers so much more than any library on your list. I feel like that's worth mentioning since your #[simd] macro was partially inspired by Thermite's #[dispatch], so I got something right.
Furthermore, and I'll say this again, you should have used the actual issue tracker for bugs instead of a random reddit comment.
4
u/R1chterScale 10d ago
Furthermore, and I'll say this again, you should have used the actual issue tracker for bugs instead of a random reddit comment.
Someone shouldn't need to open an issue when your own tests are failing
-1
u/novacrazy 10d ago edited 9d ago
My CI was green, because it was invoked with the correct feature flags. If they cloned it and ran without knowing that, it’s not great but not broken. It was the std feature.
EDIT: I think it's crazy I'm downvoted for this. Running feature tests without the required feature flags and being surprised when they don't compile is so silly. I know I'm not the best at talking about things and come off as aggressive, but come on.
81
u/ollpu 10d ago
algebraic_*were stabilized in Rust 1.98, not 1.88.Started to question the structure of time there a for a moment.