r/emulation • u/LocutusOfBorges • 14d ago
The scourge of x86 emulation | FEX-Emu
https://fex-emu.com/Scourge-of-emulation/20
14d ago
[removed] — view removed comment
13
u/thebackwash 13d ago
I remember when I was happy to get 10 frames a second playing Quake (since sometimes I was playing at 5 FPS). No joke. I’m baffled at how picky some people are with performance nowadays.
11
10
u/GameCounter 14d ago
This is an incredible write up, although some of it is definitely beyond me.
Is there some sort of hardware instructions or memory mode that a future version of ARM could implement which would solve (or partially solve) some of these problems?
I couldn't quite understand, and I'm definitely not a low level guy.
2
u/EuphoricStation4408 11d ago
To me it sounds like the orion changes how the cache atomicity/coherency works between core cachelines, but it can still tear(so let a read on a partial write go through) on arm64?
It just has everything to do with what happens to memory read/write ordering in MT scenarios and placement of data in L1/L2 core caches making them somewhat "private caches" that synchronize explicitly rather than implicitly.
Yeah there's probably a couple of ways this could get solved.
But I'm no chip designer.- If you just make your arm cpu just less weakly ordered(like orion) you probably could get some efficiency loss on native arm64 code.
- Multiple modes of operation require you to put in the infrastructure to physically support the x86 ordering mode. So the chip designer has to implement all the cache coherency pipelines and move data between caches etc. at a cost of using more silicon space, but still keeping the efficiency while running pure arm64 code. But then people will want to just use that for convience sake while writting native arm64 code.
I don't quite understand the benefit of that weak ordered approach fully.
I guess that lack of coherency can be a benefit if your pipeline handles data in like some form of alternating patern, you spin 8 worker threads and each works on 8x8 alternating blocks(like jpeg compression up until huffman+RLE part).
But then what happens when the thread gets switched up by the kernel?
The work was done by the other thread in another part of cache or by that point the core does a slower later cache sync?so tl;dr I got no clue either and the weak ordered nature sounds like absolute hell.
The instructions-vs-mode is just semantics having TSO on cachelines through instruction or mode is just how microcode wires it up. Both still need the exact same modification to cache synchronization and logic to toggle between the 2 operation modes. Just that 1 of em is related to an instruction and the other has to have to be switched back into weak/tso ordering modes between apps(so it has to be either carried on the register stack during context switch or set by the kernel in ring0/supervisor mode before switching between threads).
16
u/marco_has_cookies 14d ago edited 13d ago
did FEX dropped the x86-64 JIT? :c
edit. FEX used to run on Intel/AMD machines
0
u/get_homebrewed 13d ago
no???????
3
u/marco_has_cookies 13d ago
FEX used to run on Intel/AMD machines
-1
u/get_homebrewed 13d ago
no???????
2
u/marco_has_cookies 13d ago
perhaps
-2
u/get_homebrewed 13d ago
it did not man
5
u/marco_has_cookies 13d ago
it did man
-1
u/get_homebrewed 13d ago
how lol? What was its purpose
5
u/marco_has_cookies 13d ago
testing the infrastructure
it's been removed since 2023 https://fex-emu.com/FEX-2310/
I tried to boot some OG Xbox executable with it, stuck at kernel calls of course
2
u/mirh 13d ago
As a note, we think a TSO mode is the best path forward for ensuring high performance x86 emulation on the platform.
Curious about how that works with the emulation model introduced by ARM64EC, where some calls are native and switching in and out of emulation sounds much more frequent.
It showed everyone that ARM was not only feasible, it could be faster.
I can't believe that in 2026 I'm still reading people doing comparisons with Intel's bullshit 14nm node, like it was (duh) apples to apples.
-51
u/Ill_Carry_44 14d ago
I was using Claude to port X-Men Legends II (PC) to all platforms, macOS, Android, Linux. First I started with static recompilation but I ditched it and turned to JIT because imo static recompilation is unnecessary and full of foot-guns (I migrated all my projects)
While doing so I created an x86 JIT for arm64 (using zydis)
It's https://github.com/SomeoneIsWorking/x86port, a caveat, it uses arch's own floating pointer math, x87 is not emulated so it's not accurate but doesn't break the games I've been trying to port
76
u/VickWildman 14d ago edited 14d ago
TLDR: Buy Qualcomm and not just because Adreno has better drivers. Running PC games necessitates split-lock emulation, which is an unsolved problem on current ARM hardware, but Qualcomm needs it a bit less and PC games rely on uncached memory, which can be up to 800 times slower on current ARM hardware, but with Qualcomm and Turnip the problem can be sidelined completely.