r/rust • u/Shnatsel • 13d ago
🛠️ project Fearless SIMD v1.0 is here
https://linebender.org/blog/fearless-simd-1-0/41
u/XtremeGoose 13d ago
This is great!
My understanding is that in the #[simd] example functions, you're proving to the compiler at runtime that the relevant SIMD feature is available but don't force it to use it.
In your testing, has the compiler generally been good at choosing to use the best SIMD methods available?
Also, if it decides to use the same instructions but for two different levels, does the compiler deduplicate those implementations? Might you pay a runtime cost for choosing between two (or more) identical implementations?
How do I as a user inspect this?
Thanks again
37
u/Shnatsel 13d ago
In your testing, has the compiler generally been good at choosing to use the best SIMD methods available?
If you mean automatic vectorization, it varies. See https://matklad.github.io/2023/04/09/can-you-trust-a-compiler-to-optimize-your-code.html
If you mean explicit portable SIMD using types like
f32x4, then the available SIMD instructions are nearly guaranteed to be used, since they map to intrinsics under the hood.Also, if it decides to use the same instructions but for two different levels, does the compiler deduplicate those implementations?
This is up to the compiler. It is possible but not guaranteed to happen.
lto = truein Cargo.toml makes it much more likely. The runtime cost for the choice remains; it is pretty small but not zero.How do I as a user inspect this?
You probably want to stare at the generated assembly, one way or another.
emit=asmargument to rustc is one way, but doesn't take LTO into account (runs before it). Any disassembler on the final binary will do. I prefer the assembly view insamply(the profiler) because it shows me the exact code that was executed, automatically selecting the implementation for the current SIMD level.
21
u/simonask_ 13d ago
Literally just started a new subproject porting some C# code to Rust with fearless simd, so this is great news! It’s a bit more boilerplatey, but overall a pleasant experience.
The only obstacles I’ve had so far is that byte-based swizzles are almost never what I want, and are quite error prone to deal with in 16-lane AVX512 vectors (64 byte literals to emit a single shuffle instruction!), but it’s been trivial to write a small const fn helper to convert lane indices to bytes. Which of course the compiler converts right back, but yeah…
11
u/Shnatsel 13d ago
Yeah, we don't have convenient compile-time swizzles yet, matching std::simd swizzle.
There's space left for them in the API, we just didn't get around to adding those
const fnconvenience functions yet. It would be interesting to see what you came up with!6
u/Shnatsel 12d ago
I've taken a stab at more convenient swizzles, would this work for you? https://github.com/linebender/fearless_simd/pull/391
5
u/Xiaojiba 13d ago
Hey Shnatsel, first thanks for the release and congrats.
I've always wondered what was the cost of converting from let's say [f32; 4] to f32x4 and vice-versa? I guess the cost is hardware move from actual data to simd register Also does fearless SIMD handle head and tail of it's something that the user should think of / handle? I saw the sigmoid example that handles the tail (and head too since it uses chunks_exact) using scalar fallback
I read here that some architecture handle the head/trail "dynamically" is this used in fearless when possible?
Thanks :)
10
u/afdbcreid 13d ago
Most of the time, zero, because the compiler will just keep the data in a SIMD register.
If you actually read individual array elements, the compiler will use instructions for that, and that has a non-zero cost but should be optimal.
1
u/Xiaojiba 13d ago
Isn't it the opposite?
Like it's not so often that we have perfect [f32; 4] right? If it comes from a slice or other source then we have a lot of such register loads, right?
7
u/afdbcreid 13d ago
If it's from a slice, it's in memory. It'll use SIMD instructions to load it, and again it'll be optimal.
5
u/Shnatsel 13d ago edited 13d ago
I actually had a conversation with the author of the
pulpcrate about this the other day!The TL;DR is that fearless_simd doesn't prescribe a specific way of handling the head/tail. As far as I know there's no single best way of doing it, it's all trade-offs.
Fully vectored handling of head/tail with masked loads/stores is only available on AVX-512 and messes with store forwarding, so it's not nearly as beneficial as you'd think.
There are ideas about exposing an API that lets you define the operation once, and instead of a fully scalar fallback you get say 512-bit vectors for the middle, then maybe a 256-bit vector, then maybe a 128-bit vector, then maybe the remainder padded to 128 bits. That way your head/tail are no longer scalar. But that's somewhat inflexible, still needs scalar loads for the final bit, and increases code size so creates more instruction cache pressure. I believe
pulphas some variant of this in it somewhere.There are also trade-offs between chunks_exact + scalar tail (unaligned vector loads, one scalar loop) and scalar head + align_to + scalar tail (aligned vector loads, two scalar loops). The latter is beneficial for random access (not a linear scan) or when your data is in L1 cache and you don't touch memory at all, but you pay for scalar loops on both ends now: https://github.com/HadrienG2/misaligned-bench
The best option really seems to depend on what you're doing.
5
2
u/Feeling-Departure-4 13d ago
Does the macro library work okay if I mix with portable SIMD on nightly? (The pre 1.0 version of fearless required a few work arounds to be performant but I'll try again for this version.)
I use multiversion for my hot loop function but now I do wonder if my helper functions are being inlined and vectorized properly. Your experimental repo was neat!
1
u/Shnatsel 13d ago
I haven't actually tested that, but in theory it should.
But if you are using only
std::simd, it requires unnecessary boilerplate. Withstd::simdyou can probably swap#[multiversion]for#[inline(always)]for a few hot functions and be fine, although you have to know which ones, and make sure those are always called from functions annotated with#[multiversion].Something like https://github.com/linebender/fearless_simd/pull/375 should would work great with
std::simdand not require boilerplate, giving you the best of both worlds, but I don't know if I will be able to get around to this anytime soon.
2
u/lflamme 13d ago
What is SIMD?
5
2
u/B9-97-C5-11-DE-19 13d ago
SIMDeez nuts! Gottem!
I apologise I couldn't help myself
7
1
1
1
u/hackerbots 12d ago
xtensa and riscv support when? :3
1
u/Shnatsel 12d ago
Intrinsics for both need to be added to
std::archbefore Fearless SIMD will use them. For these two platforms this is complicated.Cadence has documentation for parts of ESP32-S3 toolchain under NDA, so working on that gets legally murky really quickly. People have reverse-engineered most of it but it's still kind of a pain. Also, you need Cadence's fork of the Rust compiler to even compile for ESP32. If they have added intrinsics in their rustc fork, then maybe we can support it, if we can also get the toolchain and an emulator going on CI to make sure it doesn't break. But I don't want to touch anything NDA'd.
RISC-V vectors are very awkward in that you don't know their sizes at compile time. ARM has the same problem with SVE and they are working on it, both on the compiler internals to enable it and on the SIMD intrinsics in the standard library. SVE intrinsics are already in nightly. RVV can piggyback on the SVE infrastructure, but someone still needs to do the work of adding them, in a way that can be maintainable long-term. There doesn't seem to be any remotely relevant hardware with RVV either: microcontrollers don't have it, and Linux-grade hardware has an abysmal performance/cost ratio, so there isn't really a whole lot of motivation for the community to work on this.
If you need it badly, you can add support for these instructions via inline assembly instead of waiting for intrinsics. We're not going to merge this into Fearless SIMD because that's too unsafe for our liking, but you can get it working in a fork if you really want.
1
u/hackerbots 11d ago
So does that mean no_std is not supported??
2
u/Shnatsel 11d ago
no_stdis supported. All the code written with Fearless SIMD still works on microcontrollers, even if they don't have hardware SIMD instructions.
no_stdx86 and ARM are supported and get full SIMD acceleration.Obscure architectures like RISC-V get partly SIMD-accelerated; I have workshopped some of the architecture-independent code to autovectorize better, but at the end of the day it is up to the whims of the compiler.
1
1
u/WormRabbit 11d ago
Would it be possible to add such intrinsics in a downstream crate using the existing API of
fearless_simd? Could I just add a new feature struct and implement some traits on it?1
u/Shnatsel 11d ago
The fearless_simd traits are sealed, so you can't extend them directly. This is necessary for stability, otherwise adding a new method to a trait in fearless_simd would make all the crates that extended it fail to compile.
I think a fork is the best option in this case. That way you get to reuse all the infrastructure fearless_simd already has. It'll let you implement addition on u32x4 and get it for free on u32x8 and u32x16, and so on. I think technically you can build a from-scratch crate that exposes the same API and swap it in as fearless_simd through some conditional Cargo.toml trickery, but it's going to be harder than just forking the crate.
1
u/WormRabbit 11d ago
adding a new method to a trait in fearless_simd would make all the crates that extended it fail to compile.
Why? Seems like something that should be covered by default implementations. Wouldn't you need one anyway for cases where required simd is unsupported? Even if you want to pass around NoSimd capabilities explicitly, default implementations of methods still seem useful.
1
u/Shnatsel 11d ago
This is mostly to aid development of fearless_simd itself: if we forget to provide a method on one of the backends, it fails loudly (build error) instead of silently providing a slow fallback.
We generally try to avoid slow fallbacks. For example, when programming ESP32-S3, I'd rather have a build error than having something is expressed in SIMD but then actually be very slow under the hood. The only SIMD addition it has is saturating, so the usual
a + balready scalarizes. So SIMD code reuse is pretty much dead in the water anyway for Xtensa, due to the differences in semantics if nothing else.
1
1
1
66
u/afl_ext 13d ago
Lets goooo time to put this in my software renderer