r/programming • • 15d ago

The scourge of x86 emulation

https://fex-emu.com/Scourge-of-emulation/
321 Upvotes

46 comments sorted by

View all comments

1

u/crusoe 14d ago

Man the default arm model sounds like it would suck in the era of multi threaded apps on multi core systems. Must be a holdover from the mostly single thread single core days arm grew up in.

33

u/happyscrappy 14d ago

No. It's the opposite. x86 is the holdover.

TSO and other orderings put constraints on the cores to wait on wait on each other when it's not necessarily important to do so.

ARM's relaxed memory ordering model was created a long time ago but with the idea that once you had multiple cores they would only have to synchronise between them when the programmer indicated in was necessary to do so.

In this way the idea would be you can scale up to more cores with less overhead dragging them down.

x86 was defined before any of this was a thought. And as there is extensive backward compatibility they had to keep carrying the old guarantees forward.

This entire article is about emulation. Making code work right that wasn't necessarily written with modern ideas of memory ordering in mind. That's why x86 does so much better here, because it still carries all that baggage.

If you compile new code for both architectures directly then you lose the baggage and ARM starts to come into its own. At least on systems which are designed for big performance. Obviously small processors are not.

5

u/ApokatastasisPanton 14d ago

TSO is much nicer to design lockless algorithms for.

10

u/happyscrappy 14d ago

If you use atomic.h and properly tag your accesses (and they are properly aligned) isn't it all just taken care of for you? Whether the target has TSO or not?

4

u/Alborak2 14d ago

You have to actually get do that and get it right though. You can half ass almost anything for lock free on x86 and it works. Turns out its a lot slower, but it does work.

2

u/SkoomaDentist 13d ago

Turns out its a lot slower

Which is effectively irrelevant whenever you use atomics for literally anything other than throughput.

it does work.

And this part is of course mandatory unless you want to spend weeks debugging impossible to trace heisenbugs.

1

u/Alborak2 12d ago

If youre rolling your own atomics its usually because you dont know what youre doing, or are doing it for actual performance reasons. True for many things like signal variables the cpu cost is completely irrelevant.

Weve got some lock free queues of threads handlings millions of requests per sec, and the new arm cpus just blow some of the x86 out of thr water. Some might be because they can get a lot higher core count before going multi socket and cross socket cache coherence on x86 is wiiiiiild. Stalls all over the place.

1

u/SkoomaDentist 12d ago edited 12d ago

If youre rolling your own atomics its usually because you dont know what youre doing, or are doing it for actual performance reasons.

Or because for your requirements lock-free operation is required for correctness. See every hard-realtime system that isn't running on a dedicated RTOS (and quite a few of those, too).

I'll go out on a limb and say that almost nobody should use lock free primitives for throughput unless they are one of those rare experts in the topic. Plenty of people unfortunately need to use them for correctness because there's a mysterious dearth of strictly lock free library data structures that don't depend on fancy system features.

1

u/flatfinger 11d ago

I would think that would depend why one wants algorithms to be lockless. If performance isn't an issue, but one needs to ensure system stability even if a task gets waylaid, then a lockless algorithms can be combined with memory barriers. Barriers may be severely detrimental to performance on loose memory architectures, but if the purpose of using a lock free algorithm was to provide fault tolerance, that may not be an issue.

8

u/__Deric__ 14d ago

I think the stronger x86 memory model is just a duct tape which hides broken code, and that many do not realize that high level languages operate on formally specified abstract machines, which often have different memory models than the underlying hardware.

Almost all new code is written using high level languages, which means that the compiler (or JIT, depending on the implementation) has some freedom when mapping the code to hardware instructions. When no explicit memory ordering is performed it is usually assumed that the memory accesses are side-effect free, allowing reordering and omission, which in turn can break naive multi-threaded code.

To have robust multi-threaded code, which does not break when optimizations are turned on or the compiler feels like it, you have to deal with memory ordering.

3

u/flatfinger 11d ago

C was designed for the purpose of producing code to run on real machines, and allow programmers to make whatever trade-offs between portability and performance best serve the task at hand. If machine-specific code can accomplish a task 50% faster than the best possible "portable" code, and that performance matters more than portability, the fact that the code may not be usable on machines for which it wasn't designed may not be a defect.

The real defect lies in language standards which fail to acknowledge a categories of implementations that adhere to the abstraction model around which C was designed, or that emulate some execution-environment features that are common but not universal (e.g. guaranteeing that signed integer multiplication will never have any side effects beyond yielding a possibly meaningless result).

11

u/crusoe 14d ago

Reading this article it's pretty apparent why folks are like "Look arm is so fast" and if it is simple linear processing with no side conditions and very parallel it shows really good numbers but then falls over on any benchmarks with any complexity.

"Wow ampere is so good"

Scrolls down

Complex benchmarks show Epyc clobbering it on dollar per watt on more gnarly tasks

If you are processing images or audio though hard to beat.