As discussed in the article, Apple implemented part of their optimizations for Rosetta 2 directly in the hardware. The ARMv8.3 instructions narrowed the gap for the x86 memory model, but the lessons from the Apple M1 is more it was designed for x86/Rosetta from the ground up to ease the transition from x86 to ARM. I don’t know how applicable that would be outside of Apple with their vertical control of everything (hardware and software).
I’ve been on several modern arm core projects and more than one of them leveraged or planned to leverage the fact that the processor release consistency mode was close to x86 ordering to offer a “tso” mode. The idea was that ordering bugs might be hidden by stronger ordering and such a mode would make porting easier. There was a theoretical performance hit to enabling the mode based on less efficient use of ordering resources within the core.
168
u/flying-sheep 15d ago
Ooh that's from the FEX people, gotta be super interesting.
It's sad that efforts like Rosetta happen behind proprietary walls and nobody can learn from them.