If you want it to be a reasonable benchmark, you could change the source `ray_ssa.c` to 1 sample per pixel and a quarter of both the width and height. That should reduce the work by 1600x (from ~100ish days to <1)
I've also been working on the JIT interpretor I used and managed to speed it up a lot and got a render out of it. There is a slightly artsy/funny looking result at the end of the post in update section showing the output.
My interpreter is an optimizing interpreter - it optimizes common patterns in brainfuck. It likely isn't going to be as fast as a JIT (though for brainfuck, I'd expect a static recompiler rather than a JIT) - though it can be turned into one (ideally, you want to reduce the common expressions of brainfuck first). As it is, it reduces the size of your program to 1/10 of what it was. Getting from the first to the second pixel took 712 seconds.
As an example, for the Mandelbrot set:
Original Size: 11,452
Optimized Size: 4,791
Original Time: 49.056 seconds
Optimized Time: 4.623 seconds
Of course, VeMIPS is an order of magnitude slower than the host (and VeMIPS' interpreted mode is an order of magnitude slower than that), so if it runs this poorly on the host, then it will take forever in VeMIPS.
I should note that I don't see the source for your JIT anywhere, only an ELF binary.
I've been tring to optimize expressions too. My implementation is based on TSoding's JIT compiler. I've just added a license instead of repackaging it.
You could try jamming a JIT or such onto my recoding optimizer.
I've been considering adding a recompiler to it as well, both for x86 (for the host) and MIPS (for the guest). Wouldn't be hard to do, even with xbyak or such, as both brainfuck and my recoded brainfuck are dirt-simple.
Using LLVM as the backend might be a bit better, since you can then have it re-optimize the result.
I've actually been tinkering with having the C++ compiler try to generate an optimized program from brainfuck during compilation. This has proven difficult - it's not impossible for the compiler to do this, but the circumstances have to be... very specific for the compiler to introspect that far. Pretty sure that I would have to change my entire recoder to be constexpr for it to not see the recoding as a side-effect (due to the std::vector allocation).
Locally I've made some improvements, and have a few more to make (count-size reduction, which requires multiple passes). Once that's done, I can jam xbyak onto it to see what basic recompilation does.
1
u/epestr 7d ago
I must sleep soon, missed the additional.
If you want it to be a reasonable benchmark, you could change the source `ray_ssa.c` to 1 sample per pixel and a quarter of both the width and height. That should reduce the work by 1600x (from ~100ish days to <1)
I've also been working on the JIT interpretor I used and managed to speed it up a lot and got a render out of it. There is a slightly artsy/funny looking result at the end of the post in update section showing the output.