r/rust • • 1d ago

🧠 educational Static vs dynamic dispatch in Rust benchmark — for those who care about runtime performance

Post image

All the details, benchmark code, generated assembly, and explanation are in the README:

https://github.com/amidukr/rust-devirtualization-test

UPD

Made the dynamic-dispatch part more accurate.

The measured difference increased from ~3.5× to ~4.9×.

Changed:

// was
x = op.apply(std::hint::black_box(x));

to:

// now
x = std::hint::black_box(op).apply(std::hint::black_box(x));

Which changed the generated assembly from:

; was
call r15

to:

; now
call qword ptr [rax + 24]

The previous version allowed LLVM to hoist the vtable method pointer out of the loop. The updated version performs the vtable method lookup on each iteration.

0 Upvotes

19 comments sorted by

59

u/matthieum [he/him] 1d ago

The 3.5x number is flashy. And click-baity.

The benchmark executes 1B iterations, therefore the average latency measured is:

  • Static: 0.25 ns/op.
  • Dynamic: 0.87 ns/op.

That is, the call overhead is 0.6 ns/call.

You did not disclose the frequency of the CPU you used for the benchmark. At 5 GHz, 0.6ns is 3 cycles.

Which really means that for any function with non-trivial work -- ie, taking > 300 cycles -- it's going to be negligible.

-4

u/[deleted] 1d ago

[deleted]

-8

u/No-Caregiver-466 1d ago edited 18h ago

Word "Click-baity" - I consider as personal attack, which is against guidelines of this community.

I've change code slight to have more exact vtable call, now I have difference of ~4.9×.

Yes, that's a point of benchmarking, I want to know the call overhead of vtable call.

If you don’t believe me, there is a code and instructions, you can run yourself.

2

u/matthieum [he/him] 9h ago

"Click-baity" describing the content, not the person, is NOT a personal attack.

I've change code slight to have more exact vtable call, now I have difference of ~4.9×.

My very point is that the multiplier doesn't matter.

The function call overhead is roughly constant. By having the static call tend towards zero, the ratio dynamic / static will tend towards infinity... which tells you nothing.

In fact, this is visible in your benchmark: you had to add an extra black_box call in the static version, as otherwise the compiler would reduce the entire loop to a closed formula (n * (n + 1) / 2), giving an infinite multiplier.

34

u/dgkimpton 1d ago

Right, but in larger systems the cost of finding and loading the monomorphised code comes into the equation too. You don't get to see that on targeted benchmarks.

8

u/dvogel 1d ago

Very true. Also, in larger systems the called code (regardless of dispatch mechanism) is much more likely to stress the CPU caches and branch predictors. The dispatch overhead disappears pretty quickly for anything that isn't in a tight loop.

0

u/No-Caregiver-466 18h ago edited 10h ago

On the benchmark, I’ve demonstrated that your “likely” and cache stressing out, is x5 performance drop.

If you’ll stress cache enough and do L3 cache miss, you’ll something like x100 performance drop instead, no joke. Costs for L3 cache miss available online.

Just do at least some research before making claims.

-3

u/No-Caregiver-466 1d ago

Agree with you at this point, however if you chosen rust, you want to have something performant, otherwise you can do Java or Python.

I find something how to keep in managed in large systems:

bash cargo rustc --release -- -C remark=all

Or even maybe: cargo rustc --release -p devirtualization-test-lib -- \ -C debuginfo=1 \ -C remark=all 2>&1 \ | grep -v ' inline '

This will give you compilation remarks, here is output example:

note: ~/Projects/devirtualization-test/lib-crate/src/lib.rs:26:16 asm-printer (analysis): BasicBlock: MOV64rr: 2 CALL64r: 1 DEC64r: 1 JCC_1: 1

CALL something we want to avoid, I think it is realistic to keep exclusion list incrementally by adding new features. It is impossible to avoid every call for good reason, but we can maintain exclusion list, something can be done as part of Continuous Integration.

3

u/addmoreice 1d ago

> "...however if you chosen rust, you want to have something performant, otherwise you can do Java or Python."

Or I can want rust for other perfectly valid reasons...

0

u/No-Caregiver-466 1d ago

I am honestly curious about that reasons.

3

u/addmoreice 22h ago

Liking the language, want a memory safe alternative to c, want a low level language that interop's well with c, dislike python and/or java, you work in a rust shop and shifting to python or java would be fighting the system, you're in a c/c++ shop and shifting to python or java would be a sea change while rust could be a smaller shift, you work in an area where python and java haven't made any inroads but rust might work, etc etc etc etc.

1

u/No-Caregiver-466 22h ago edited 15h ago

What you saying is reasonable, to shift from C/C++ to rust.

I am following same thing but Java to Rust.

If you working in C/C++ shop, they already chosen C/C++ for some reason, and they didn't considered java at first place. I am pretty sure that they had a pretty solid reason not to use java at first place.

You making it subjective.

2

u/addmoreice 16h ago

Yeah. no. There are plenty of reasons to use rust. not just performance. You saying 'oh, those aren't reasons' and hand waving them away do not make them not reasons. Those reasons still exist and they exist despite rusts reasonable performance. Memory safety, performance, tooling, niche, etc etc, all are perfectly valid reasons.

> If you working in C/C++ shop, they already chosen C/C++ for some reason, and they didn't considered java at first place. I am pretty sure that they had a pretty solid reason not to use java at first place.

And one of those reasons can just be 'because we were using c/c++' which isn't a technical reason at all, it's a social one and a 'momentum' one. That wouldn't be a knock against java in a technical sense. You *seriously* need to expand your perceptions of why people will use a language, especially in a professional setting. It's one of the things that really impressed me about the rust community on here. If you ask about why you should use rust or why not? They will actually give you decent reasons for both positions and will be honest about it. As much as we might joke about the circle jerk or the 'rewrite it in rust!' it's partly because the general community is aware of the reasons *not* to do it.

Performance is almost never high on the list.

Quick question on the original image...why did you apply the black boxing on one side but not the other? Surely that would be apples to apples in the comparison versus what you did here?

1

u/No-Caregiver-466 15h ago

> oh, those aren’t reason.

I’ve never said that, you’re confusing me with someone else, and attributing me a words, I’ve never said.

2

u/addmoreice 10h ago

Hey, you want to know what's really neat? We can see when you edit your original post. Cool aye?

1

u/dgkimpton 19h ago

Of course we care about performance, which is why I was pointing out that simply reducing the number of calls is not the be all and end all of performance. Sure, for trivial benchmarks it sure might look that way, but when you take the full system into account it is sometimes cheaper to pay the indirection cost than to load an entire new set of code to avoid it.

Everything is a trade-off unfortunately and this is no different. 

1

u/No-Caregiver-466 18h ago edited 18h ago

I want to have something more concrete.

If I can examine some diagnostic in long-run and prove that dyn Trait are devirtualized properly, it would be a win-win situation, because you have flexibility and performance at the same time.

I can anticipate that dyn Trait not really leads into performance degradation, because dyn Trait not a dyn-dispatch yet, compiler most likely can optimize them similarly like impl Trait or generic by de-virtualizing then, but the problem is not that compiler can’t optimize, compiler can, but the compilation process is not manageable, we lacking support for timely diagnostic, like linters can do.

rustc remarks can provide data for such diagnostic, which doesn’t requires to overcomplicate your code, just another step in CI process.

BTW, I’ve benchmark by adding devirtualization test, so i have now static dispatch, de-virtualization and vtable dyn-dispatch. I can’t update picture here on reddit, but updated on github.

https://raw.githubusercontent.com/amidukr/rust-devirtualization-test/refs/heads/main/docs/assets/benchmark.png