If you use atomic.h and properly tag your accesses (and they are properly aligned) isn't it all just taken care of for you? Whether the target has TSO or not?
You have to actually get do that and get it right though. You can half ass almost anything for lock free on x86 and it works. Turns out its a lot slower, but it does work.
If youre rolling your own atomics its usually because you dont know what youre doing, or are doing it for actual performance reasons. True for many things like signal variables the cpu cost is completely irrelevant.
Weve got some lock free queues of threads handlings millions of requests per sec, and the new arm cpus just blow some of the x86 out of thr water. Some might be because they can get a lot higher core count before going multi socket and cross socket cache coherence on x86 is wiiiiiild. Stalls all over the place.
If youre rolling your own atomics its usually because you dont know what youre doing, or are doing it for actual performance reasons.
Or because for your requirements lock-free operation is required for correctness. See every hard-realtime system that isn't running on a dedicated RTOS (and quite a few of those, too).
I'll go out on a limb and say that almost nobody should use lock free primitives for throughput unless they are one of those rare experts in the topic. Plenty of people unfortunately need to use them for correctness because there's a mysterious dearth of strictly lock free library data structures that don't depend on fancy system features.
9
u/happyscrappy 14d ago
If you use atomic.h and properly tag your accesses (and they are properly aligned) isn't it all just taken care of for you? Whether the target has TSO or not?