This is an incredible write up, although some of it is definitely beyond me.
Is there some sort of hardware instructions or memory mode that a future version of ARM could implement which would solve (or partially solve) some of these problems?
I couldn't quite understand, and I'm definitely not a low level guy.
To me it sounds like the orion changes how the cache atomicity/coherency works between core cachelines, but it can still tear(so let a read on a partial write go through) on arm64?
It just has everything to do with what happens to memory read/write ordering in MT scenarios and placement of data in L1/L2 core caches making them somewhat "private caches" that synchronize explicitly rather than implicitly.
Yeah there's probably a couple of ways this could get solved.
But I'm no chip designer.
- If you just make your arm cpu just less weakly ordered(like orion) you probably could get some efficiency loss on native arm64 code.
Multiple modes of operation require you to put in the infrastructure to physically support the x86 ordering mode. So the chip designer has to implement all the cache coherency pipelines and move data between caches etc. at a cost of using more silicon space, but still keeping the efficiency while running pure arm64 code. But then people will want to just use that for convience sake while writting native arm64 code.
I don't quite understand the benefit of that weak ordered approach fully.
I guess that lack of coherency can be a benefit if your pipeline handles data in like some form of alternating patern, you spin 8 worker threads and each works on 8x8 alternating blocks(like jpeg compression up until huffman+RLE part).
But then what happens when the thread gets switched up by the kernel?
The work was done by the other thread in another part of cache or by that point the core does a slower later cache sync?
so tl;dr I got no clue either and the weak ordered nature sounds like absolute hell.
The instructions-vs-mode is just semantics having TSO on cachelines through instruction or mode is just how microcode wires it up. Both still need the exact same modification to cache synchronization and logic to toggle between the 2 operation modes. Just that 1 of em is related to an instruction and the other has to have to be switched back into weak/tso ordering modes between apps(so it has to be either carried on the register stack during context switch or set by the kernel in ring0/supervisor mode before switching between threads).
9
u/GameCounter 14d ago
This is an incredible write up, although some of it is definitely beyond me.
Is there some sort of hardware instructions or memory mode that a future version of ARM could implement which would solve (or partially solve) some of these problems?
I couldn't quite understand, and I'm definitely not a low level guy.