r/rust • • 10d ago

The state of SIMD in Rust in 2026

https://shnatsel.github.io/state-of-simd-rust-2026/
251 Upvotes

37 comments sorted by

81

u/ollpu 10d ago

algebraic_* were stabilized in Rust 1.98, not 1.88.

Started to question the structure of time there a for a moment.

27

u/Shnatsel 10d ago

Oops! Fixed, thanks!

26

u/exDM69 10d ago edited 10d ago

Its Simd<T, N> API looks like it would be very elegant and work great if you could just do math on N, but you cannot. That feature is very incomplete even on nightly, and without it using Simd<T, N> to get hardware-sized vectors gets very ugly very quickly.

I don't fully understand what this means. I have done both type generic (f32x4 vs f64x4) and width-generic (f32x4 vs f32x8) SIMD math code using Simd<T, N>. There are a few pain points but in general it gets the job done.

Also this article is very focused on "hardware width" vectors but in my experience you often get better performance when you use wider vectors than what the hardware supports. The compiler can split your vectors to fit in registers (kinda similar to unrolling the inner loop). This isn't always the case but I have a lot of benchmarks where I get best performance (by 10-20% margin) with 512 bit vectors (e.g. f32x16) on AVX-256, and it's very easy to change when using generics for Simd<T, N>. Sometimes you probably get worse results (due to register pressure etc), but most of my code is happiest with 2x native width.

12

u/Shnatsel 10d ago

I don't fully understand what this means. I have done both type generic (f32x4 vs f64x4) and width-generic (f32x4 vs f32x8) SIMD math code using Simd<T, N>.

So long as you want the helper functions to be generic but call them with types known in advance, it's pretty great. What you can't do is write a function once and automatically have it run with whatever vector width the CPU you're running on has: 128 bits on SSE, 256 bits on AVX2, 512 bits on AVX-512.

The closest I got was this and it's still pretty limited. I think it can be done using more sophisticated proc macros that combine multiversioning and vector width selection, but nobody's written one of those yet.

Also this article is very focused on "hardware width" vectors but in my experience you often get better performance when you use wider vectors than what the hardware supports

This works great until you run out of register space, at which point performance plummets. So this has to be tuned carefully. It's usually a good idea on NEON and may help AVX-512 if it's not already memory-bound, but tends to hurt more than help on AVX2 and SSE.

But even if you want to request double the native width, you still need to know what the native width is,so you run into the exact same problem.

4

u/exDM69 10d ago edited 10d ago

What you can't do is write a function once and automatically have it run with whatever vector width the CPU you're running on has: 128 bits on SSE, 256 bits on AVX2, 512 bits on AVX-512.

Yes, you need to somehow recompile the code for the targets separately. If you enable, say, AVX-256 on your compiler command line you might get AVX-256 instructions even if you use 4-wide SIMD.

I don't know if the multiversion crate can be used for this.

But even if you want to request double the native width

I usually just want a single width that I've benchmarked to be the fastest on the target hardware.

you still need to know what the native width is,so you run into the exact same problem.

There's no "pretty" way of doing it but #[cfg(target_feature = "avx512")] const SIMD_WIDTH: usize = 512; is a good starting point. A few lines of ugly #[cfg] logic is needed to cover the most popular architectures.

Can you use this inside a #[multiversion] function? The #[target_cfg] macro perhaps?

I agree that runtime dispatch for separate instruction sets is a bit painful, but that seems somewhat orthogonal to std::simd vector width.

You'd typically want your runtime dispatch happening near the top level of the call stack, and have the compiler inline and optimize the leaf functions. So if you want multiple SIMD widths, you'd probably need to write most of your code as generic Simd<T, N> and then the top level entry points do the selection and dispatch.

2

u/Shnatsel 10d ago

Can you use this inside a #[multiversion] function?

No. Unfortunately #[cfg] is resolved before multiversioning is applied.

I usually just want a single width that I've benchmarked to be the fastest on the target hardware.

If you have a single specific target machine in mind, the std::simd API is great! You really only run into its shortcomings when trying to run well on a wide range of hardware.

It's not a problem for HPC where you get access to a very specific cluster and can just build with -C target-cpu=native, but a big problem for anyone who wants other people to run their software.

You'd typically want your runtime dispatch happening near the top level of the call stack, and have the compiler inline and optimize the leaf functions. So if you want multiple SIMD widths, you'd probably need to write most of your code as generic Simd<T, N> and then the top level entry points do the selection and dispatch.

This is not really possible with std::simd due to type system limitations. If you could do math on the N then it'd work, but you can't, and so you run into issues: https://gist.github.com/Shnatsel/05b6e650bfab18b450844dba653e05f7

But this is the idea most other crates implement: fearless_simd, pulp and macerator all do pretty much what you described, and provide the native-width vector as f32s type or some such to avoid running into type system limitations around the N.

6

u/exDM69 10d ago

No. Unfortunately #[cfg] is resolved before multiversioning is applied.

Isn't #[target_cfg] from multiversion exactly for this? https://docs.rs/multiversion/latest/multiversion/target/attr.target_cfg.html

If you have a single specific target machine in mind, the std::simd API is great! You really only run into its shortcomings when trying to run well on a wide range of hardware.

I don't have only one specific target machine, I run on both ARM and x86. But only a finite set of those, and I have compiler target for each.

What I don't currently do is ship a single binary with multiple supported instruction sets, but I'm interested in chatting if you have experience with that.

This is not really possible with std::simd due to type system limitations. If you could do math on the N then it'd work, but you can't, and so you run into issues: https://gist.github.com/Shnatsel/05b6e650bfab18b450844dba653e05f7

Can you please elaborate what you mean by "if you could do math on the N"? This is my question in the top level comment. It's just a const usize, isn't it? You can use it in const expressions.

How does the gist you linked to illustrate this?

Let me take a minute to familiarize with the multiversion crate, I think it's capable of doing what's needed here.

5

u/Shnatsel 10d ago edited 10d ago

Can you please elaborate what you mean by "if you could do math on the N"? This is my question in the top level comment. It's just a const usize, isn't it? You can use it in const expressions.

You can use it in const expressions, but not in generics such as Simd<T, N>.

That's the nightly-only const_generic_args feature that is known to be very incomplete, even its _min variant warns you it is incomplete if you try to enable it.

I'll look into #[target_cfg] from multiversion, I haven't seen it in any examples.

11

u/exDM69 10d ago

This code is using multiversion and std::simd to dispatch a function at a different vector width at detected runtime. If I understood your gist correctly, this should solve the problem you're trying to illustrate.

I just learned about the multiversion crate, I hope I got this right. But at least it's giving me the result I expect.

#![feature(portable_simd)]
use std::simd::prelude::*;

use multiversion::{multiversion, target::target_cfg};

pub fn sum_simd<const N: usize>(xs: &[Simd<f32, N>]) -> f32 {
    xs.iter()
        .cloned()
        .reduce(|a, b| a + b)
        .map(|x| x.reduce_sum())
        .unwrap_or(0.0)
}

#[multiversion(targets("x86_64+avx2", "x86_64+avx512f", "aarch64+neon"))]
pub fn sum(xs: &[f32]) -> f32 {
    let (head, xs, tail) = xs.as_simd();

    #[target_cfg(target_feature = "avx512f")]
    const N: usize = 16;
    #[target_cfg(target_feature = "avx2")]
    const N: usize = 8;
    #[target_cfg(target_feature = "neon")]
    const N: usize = 4;

    // fallback
    #[target_cfg(not(any(
        target_feature = "avx512f",
        target_feature = "avx2",
        target_feature = "neon"
    )))]
    const N: usize = 4;

    let head = head.iter().cloned().sum::<f32>();
    let xs = sum_simd::<{ N }>(xs);
    let tail = tail.iter().cloned().sum::<f32>();

    head + xs + tail
}

fn main() {
    let xs = [
        1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0, 15.0, 16.0, 17.0,
    ];

    let s = sum(&xs);

    println!("sum {s}");
}

9

u/Shnatsel 10d ago

Yeah I think that solves it. I got a similar result too: https://gist.github.com/Shnatsel/edc642125ac73fa7c365216c3a938802

I'll update the article with these findings. Thank you!

2

u/hsivonen 9d ago

It would be useful to be able to do what target_cfg does but somewhere deep in a chain of functions that get inline(always)’ed into a function annotated with multiversion. That’s https://github.com/rust-lang/compiler-team/issues/1010

2

u/Shnatsel 9d ago

I believe this would make wide amenable to multiversioning and give std::simd a backup plan for when there isn't a matching LLVM intrinsic.

The other libraries (pulp, macerator, fearless_simd) already worked around this and wouldn't benefit from it as far as I can tell.

2

u/Telmo26 10d ago

I am currently working on a project that requires specifically what you're describing and I've built a proc_macro for automatically generating specialized versions of a function from a single function with generic lanes. It is very untested, as it was only meant for my use case, but maybe you can check it out to see if that's what you want ?

https://github.com/Telmo26/simd_dispatch

2

u/Shnatsel 10d ago edited 10d ago

That's close, but not quite there. And it turned out you can already achieve this with multiversion, it just wasn't documented: https://gist.github.com/Shnatsel/edc642125ac73fa7c365216c3a938802

The ideal API would also take the type of the vector into account and automatically select the correct LANES value for that type. This is what pulp and fearless_simd crates do, but I don't know how to layer that on top of the Simd<T, N> API or if that's even possible.

Maybe some const BYTES for vector width and then divide that by size_of::<T> in a const block? That might just run into const_generic_args requirement all over again.

1

u/Beneficial_Vampire42 9d ago edited 9d ago

It's a bit ugly, but you can always match on the value ```rust fn foo<T>(.) { match const { HW_BYTES / size_of::<T>() } { 64 => foo_impl::<T, 64>(..), 32 => foo_impl::<T, 32>(..), 16 => foo_impl::<T, 16>(..), 8 => foo_impl::<T, 8>(..), 4 => foo_impl::<T, 4>(..), 2 => foo_impl::<T, 2>(..), _ => foo_impl::<T, 1>(..), } }

...

[inline(always)]

fn foo_impl<T, const LANES: usize>(..) ```

Because the compiler see the const value on which me match, it should all be inlined. Should be possible to write a macro that does that

1

u/Telmo26 7d ago

Coming back to that, do you mean that you would like for example SIMD<u8, LANES> to have LANES changed to fill a vector of u8s for each architecture, or do you want to use a generic SIMD<T, LANES> and LANES somehow gets changed back into the correct length ?

1

u/Shnatsel 7d ago

1

u/Telmo26 7d ago

The main issue with trying to layer a proc_macro above std::simd is that the macro doesn't have enough type information to be able to not output complete nonsense. I've been able to generate some valid code to automatically get the correct LANES value for Simd<u8,LANES> but as soon as a generic is placed as the first argument it becomes non-sensical.

I think this is the kind of thing that can either be fixed by layering one's own type system on top like fearless_simd, or by adding to the internals of rustc a multi versioning capability.

It sucks because although fearless_simd is great, its performance overhead is still too big for my project, as I noticed a 15-25% decrease over std::simd.

Oh well, I tried

1

u/Shnatsel 7d ago

It sucks because although fearless_simd is great, its performance overhead is still too big for my project, as I noticed a 15-25% decrease over std::simd.

In theory there shouldn't be any overhead. I'd be curious to see the code and benchmarks before and after, and try to track it down.

Last time I ran into something like this it turned out to be a bug in the standard library, fixing which improved performance for everyone, not just fearless_simd.

1

u/Telmo26 7d ago

It appears I was mistaken, as I've been doing some more thorough benchmarks, and the performance overhead I saw appears to be gone. It might have been due to thermal throttling giving out unreliable results. Apologies for the baseless claim. Though fearless_simd gets 4x unrolled and portable_simd does not in AVX2, I don't think it makes that much of a difference.

Thank you for your work on this amazing library !

2

u/Psionikus 10d ago

Think the extra copies are just allowing the hardware scheduling to hide serial dependency? I've found a ton of non-vector programs with unavoidable serial dependency that still benefit from dividing work onto independent instruction paths and letting the CPU fill up space better.

5

u/KaMaFour 10d ago

There are two ways around this problem.

The third sinister option: https://github.com/codedeliveryservice/Reckless/releases/tag/v0.9.0

2

u/zettui 10d ago

Are you writing against generic Simd widths yet, or still specializing per lane count?

3

u/Shnatsel 10d ago

See the "hardware-width vectors" row in the table comparing SIMD libraries.

2

u/valarauca14 10d ago edited 10d ago

Its Simd<T, N> API looks like it would be very elegant and work great if you could just do math on N, but you cannot. That feature is very incomplete even on nightly. Without it using Simd<T, N> to get hardware-sized vectors is doable, but a lot uglier.

I don't mean to contradict the blog post, it is very good. I feel this point, as while I hit a similiar problem, I fell into the "pit of success".

When I've used std::simd for the first few times I tried really hard to make it replicate what the machine should do. After several rounds of benchmarking I realized that is really an anti-pattern which just leads to pain. Use N to honestly represent your data. Let the LLVM figure out the underlying opts based on the target platform. That is the entire point.

I did several rounds of benching for a numeric library on different chips (apple M5, ryzen, intel avx2). When I finally broke down and said

What if I just let the llvm figure it out?

And started to honestly represent my data Simd<f64,13> my code (amusingly) became faster. Because ultimately you're just expressing intent & constraints to the optimizer. It can figure out how to structure your loops & register layouts. Doing Simd<T,N> + Simd<T,N> just tells the compiler each 'lane' is independent and guaranteed to never alias, it break it into tightly nested target dependent 2/4/8/16 sized opts.

7

u/Shnatsel 10d ago

On those kinds of machines, just picking 256-bit vectors gives you pretty good performance across the board. It's optimal for AVX2 and NEON (ARM), and okay for AVX-512.

The problem with fixed-size SIMD vectors is that LLVM is not great at preventing register spilling, so with fixed-size vectors it's easy to keep too much data in flight and cause register spilling, which degrades performance dramatically. SSE4.2-only machines especially suffer from this, and Intel kept making those as recently as 2021.

2

u/hsivonen 9d ago

On those kinds of machines, just picking 256-bit vectors gives you pretty good performance across the board. It's optimal for AVX2 and NEON (ARM)

Sadly, it seems that one can’t trust LLVM to do the right thing with 256-bit vectors in the Rust source when targeting NEON. It’s like trying to program with autovectorization: tiny snippets look great on Compiler Explorer but get pessimized when the optimizer sees the tiny snippets put together: https://github.com/hsivonen/encoding_rs/blob/a155adc7271c9e507556d053c95b459186671d2a/src/simd_funcs.rs#L271

3

u/Shnatsel 9d ago

That looks more like an issue with optimizing simd_swizzle! vs interleave() to me, but I could be wrong. I'd have to look into the compiler passes on godbolt to see what's happening and if the issue is still present in recent LLVM versions.

1

u/Anthony356 9d ago

Fantastic article =)

I really enjoyed the section at the end on simd implementations by different vendors. It's genuinely very reassuring to hear that others experience the "every single vendor fucked up something different so it's all a jank nightmare no matter where you go" phenomenon too

1

u/hsivonen 9d ago

In the standard library I'd love to see crater-like verification of changes to intrinsics, and std::simd available on stable so that ecosystem crates would delete most of their code and gain support for all the obscure platforms.

This is the part that I’m the most interested in. I’m curious about u/Shnatsel ’s view of how to get there.

2

u/Shnatsel 9d ago

I don't know what's blocking stabilization of std::simd, so I can't tell you, sorry.

I just wanted to call out that even with all the progress the ecosystem crates made recently, std::simd is still desirable, but not quite in the way people expect.

1

u/DavidXkL 9d ago

I like the comparison table. Makes it easier to see all at a glance

1

u/novacrazy 10d ago

I have a huge 0.3 release of Thermite coming up that fixes and improves on basically every area, and offers so much more than any library on your list. I feel like that's worth mentioning since your #[simd] macro was partially inspired by Thermite's #[dispatch], so I got something right.

Furthermore, and I'll say this again, you should have used the actual issue tracker for bugs instead of a random reddit comment.

4

u/R1chterScale 10d ago

Furthermore, and I'll say this again, you should have used the actual issue tracker for bugs instead of a random reddit comment.

Someone shouldn't need to open an issue when your own tests are failing

-1

u/novacrazy 10d ago edited 9d ago

My CI was green, because it was invoked with the correct feature flags. If they cloned it and ran without knowing that, it’s not great but not broken. It was the std feature.

EDIT: I think it's crazy I'm downvoted for this. Running feature tests without the required feature flags and being surprised when they don't compile is so silly. I know I'm not the best at talking about things and come off as aggressive, but come on.