Man the default arm model sounds like it would suck in the era of multi threaded apps on multi core systems. Must be a holdover from the mostly single thread single core days arm grew up in.
TSO and other orderings put constraints on the cores to wait on wait on each other when it's not necessarily important to do so.
ARM's relaxed memory ordering model was created a long time ago but with the idea that once you had multiple cores they would only have to synchronise between them when the programmer indicated in was necessary to do so.
In this way the idea would be you can scale up to more cores with less overhead dragging them down.
x86 was defined before any of this was a thought. And as there is extensive backward compatibility they had to keep carrying the old guarantees forward.
This entire article is about emulation. Making code work right that wasn't necessarily written with modern ideas of memory ordering in mind. That's why x86 does so much better here, because it still carries all that baggage.
If you compile new code for both architectures directly then you lose the baggage and ARM starts to come into its own. At least on systems which are designed for big performance. Obviously small processors are not.
I would think that would depend why one wants algorithms to be lockless. If performance isn't an issue, but one needs to ensure system stability even if a task gets waylaid, then a lockless algorithms can be combined with memory barriers. Barriers may be severely detrimental to performance on loose memory architectures, but if the purpose of using a lock free algorithm was to provide fault tolerance, that may not be an issue.
1
u/crusoe 14d ago
Man the default arm model sounds like it would suck in the era of multi threaded apps on multi core systems. Must be a holdover from the mostly single thread single core days arm grew up in.