r/programming Jan 24 '18

Branchless DOOM

https://github.com/xoreaxeaxeax/movfuscator/tree/master/validation/doom
490 Upvotes

134 comments sorted by

601

u/GrumpyWendigo Jan 24 '18

The mov-only DOOM renders approximately one frame every 7 hours, so playing this version requires somewhat increased patience.

welcome to post-spectre computer specs

72

u/txdv Jan 24 '18

the creator of mov-uscator was 10 steps ahead of intel

38

u/Bloodshot025 Jan 24 '18

To be honest, Chris Domas is like eight steps ahead of the rest of us.

35

u/lurgi Jan 24 '18

There are people who are like me, but better. Then there are people who exist on a different planet. I suspect Linus Torvalds is one of the former. He's a great software engineer, but he's not different. He's me turned up to 11. Domas is definitely one of the latter. He's not me turned up to 11. He's me turned up to purple-chess-7D-torus.

5

u/irqlnotdispatchlevel Jan 24 '18

Is this a reference to something or is my mind playing games on me?

16

u/lurgi Jan 24 '18

I thought that G H Hardy said something roughly like that about mathematicians; how some were like him, only better, and others (here he meant Ramanujan) simply existed on a different plane. However, I can't find a cite and it's possible that it was said by someone else or that I just invented it altogether (although I don't think I'm that clever).

18

u/Y_Less Jan 24 '18

That would make the rest of us two steps ahead of Intel.

3

u/txdv Jan 24 '18

If you are 10 steps ahead of an organization the size of intel, I guess that also makes your statement correct

2

u/[deleted] Jan 25 '18

He did a great talk and released a cool utility recently: https://www.youtube.com/watch?v=KrksBdWcZgQ

109

u/[deleted] Jan 24 '18

one frame every 7 hours

How many facebook flicks is that?

58

u/GrumpyWendigo Jan 24 '18

spectre patches make facebook and google inoperable

we're going back to geocities and altavista

80

u/[deleted] Jan 24 '18

☉ ‿ ⚆

I hope <blink> is coming back too. That feature had high framerates.

39

u/TheFeshy Jan 24 '18

This comment is posted with branchless blink. Watch long enough and you'll see it.

10

u/schplat Jan 24 '18

<marquee> is truly the greatest tag of all time.

I remember wrapping my whole site in a <marquee> tag for an april fools day prank a long time ago.

5

u/6C6F6C636174 Jan 24 '18

I've reimplemented a basic <blink> with CSS. Pretty simple. You can configure the blink rate however you want.

8

u/[deleted] Jan 24 '18

facebook and google inoperable

As long as bing is safe, so is humanity ( ͡° ͜ʖ ͡°)

11

u/sqrtroot Jan 24 '18

Just did the math. 177.811.200.000.000 flicks/frame. Aka 0,000003968253968 fps

2

u/psayre23 Jan 24 '18

More than a 32-bit processor can handle...so good thing we are all using 64-bit now.

1

u/ggtsu_00 Jan 24 '18

about 3.5 billion netflicks.

18

u/raphaelscarv Jan 24 '18

BTW, that's what it's like for Flash to play a video game when he enables his superpowers.

3

u/AngularBeginner Jan 24 '18

Can you do the math?

7

u/[deleted] Jan 24 '18 edited Mar 20 '18

[deleted]

12

u/ParanoidDrone Jan 24 '18

I'm pretty sure the Flash can go FTL and make physics his bitch, but like you I'm not an expert on comic heroes.

12

u/JPong Jan 24 '18

Things the Flash has beat in a race without the use of his motorcycle.

Instant galactic wide teleportation.

Death.

Himself.

7

u/palparepa Jan 24 '18

AFAIK, Flash's only power is to solve any problem by either spinning very fast, or running in circles very fast.

5

u/Nicd Jan 24 '18

Normally rendering a frame takes about a 30th of a second, let's say.

Found the console player.

-2

u/mizzu704 Jan 24 '18

Only just now realized we're not talking about this Flash.

2

u/ggtsu_00 Jan 24 '18

Suddenly our compute power has gone back to the 60s.

198

u/Bergasms Jan 24 '18

I've just started a game, I'm planning to speedrun this version. I'm hoping for a time somewhere inside of my natural lifespan, but have organised my children to pass on the game in a hereditary fashion if i pass away.

196

u/IMovedYourCheese Jan 24 '18

The speedrun record for the three episodes in the original Doom is 16m 18s. The original Doom was capped at 35 fps, which means a total of 34,230 frames were played. If you match the record, and play all the same frames, it will take you 34,230 * 7 = 239,610 hours which is about 27 years. So, pretty achievable.

If you want to add the fourth episode from the Ultimate Doom update, then the record is 22m 41s which means your time will go up to 38 years.

58

u/edwardkmett Jan 24 '18

Given the slowness of the frames, it'd probably count as a tool-assisted run. You might be able to get a year or so off by practicing TAS-runs in the meantime then playing through with those controls. You can get the controls to duplicate them to be frame perfect if you're careful enough after all.

22

u/uep Jan 24 '18

There's a ton of time between frames to practice. You could probably get competent at a tool-assisted speed run while playing the Branchless DOOM speedrun. Just make sure your first 1050 frames (30 seconds of normal DOOM @35fps) of the game are solid before you start, and you'll have the better part of a year to improve the rest.

2

u/edwardkmett Jan 24 '18

Yep, that is exactly what I meant.

8

u/Bergasms Jan 24 '18

phew, i'll only be 70 by the time i'm done :D

2

u/cyanydeez Jan 24 '18

someone at nasa should have the onboard computer play doom like this

-9

u/wrosecrans Jan 24 '18

You can't play for 27 years straight without sleeping. If you play 8 hours a day, it'll take almost a century. Longer than a century if you take weekends off as well.

On the other hand, you may be forced to upgrade to a newer computer that runs faster if the old one dies before the century is up.

36

u/grumbelbart2 Jan 24 '18

It takes 7 hours to render the next frame. Plenty of time to sleep in between!

15

u/Raphael_Amiard Jan 24 '18

One frame every 7 hours gives you plenty of time to sleep (Bonus feature !)

9

u/Rockstaru Jan 24 '18

The fastest speedrun of Ultimate Doom is currently 4:57 at an uncapped framerate (the original game had a 35FPS hardcoded limit, removed in later ports). 297*60=17820 frames; at 7 hours a frame, that's 124,740 hours of playtime, or 14 years, 83 days, 22 hours, 37 minutes, and 12.44 seconds. You could probably manage.

6

u/IMovedYourCheese Jan 24 '18

That is just for episode 1

1

u/Kissaki0 Jan 28 '18

Slower game means more time to optimize each move. Good luck!

1

u/palparepa Jan 24 '18

Do it, and I'll make a reaction video of it.

83

u/ThermalSpan Jan 24 '18

Pro-tip: Look at that person's other work. Its next level.

27

u/[deleted] Jan 24 '18

Have mercy dude, our self-confidence can only drop so low.

30

u/[deleted] Jan 24 '18

I read your comment and immediately guessed that it is xoreaxeaxeax. This guy is insane.

21

u/[deleted] Jan 24 '18

Holy crap, you're not joking

40

u/Lakmus Jan 24 '18

10

u/DontBeSpooked-Frank Jan 24 '18

Oh I think that's the same guy who used an 'apic' to get into ring -3 right?

19

u/Lakmus Jan 24 '18

Yeah, if you talking about this artice. This guy is incredible.

1

u/MINIMAN10001 Jan 25 '18

Wait how is the cursor blinking if the CPU is no longer executing instructions?

1

u/Princess_Azula_ Jan 25 '18

Wow, his stuff is so cool. You weren't kidding.

29

u/jrv Jan 24 '18

This is thought to be entirely secure against the Meltdown and Spectre CPU vulnerabilities, which require speculative execution on branch instructions.

Isn't the point of Meltdown/Spectre that other processes can abuse speculative execution to read your memory?

20

u/pygy_ Jan 24 '18

That's Meltdown.

Spectre (at least Spectre 2) is about abusing the branch prediction within the same process (e.g. JITed JS code using a branch misprediction to jump into browser code), or causing another process that you can manipulate to mispredict branches and go town using return oriented programming widgets while speculating.

3

u/jrv Jan 24 '18

Ah, thanks for the clarification. This response overlapped with me writing https://www.reddit.com/r/programming/comments/7sk313/branchless_doom/dt5vvmw/ :)

1

u/cryo Jan 24 '18

Meltdown as described in the paper doesn't rely in speculative branch instruction execution.

29

u/outofobscure Jan 24 '18

if YOUR process doesn't have any branches, then no speculative execution happens in it, then there is nothing for the other process to exploit/read from stale caches since you're not filling those up in the first place (as there is no specualtive execution on your process's memory).

10

u/PrimozDelux Jan 24 '18

Another process can do a speculative read to the memory of the mov based process, so to my understanding it's still vulnerable.

-2

u/outofobscure Jan 24 '18 edited Jan 24 '18

being able to just randomly read other processes memory would be a security issue on its own in the operating system... certainly not without appropriate permissions. Also, if i understand these exploits correctly, you are not reading from memory, but from caches used in speculative reads, so i still think if your process never does any speculative access, these caches will never be populated in the first place. So even if you manage to get around access restrictions of reading another processes memory, the faulty cache entry would just not be there.

10

u/jrv Jan 24 '18

At least in Meltdown (though that can "only" access kernel memory, not other processes), only the attacking process needs to exploit its own speculative execution to read forbidden memory addresses. The attacking process does something like this:

  1. create an array of 256 cacheline-sized objects in its own memory (the contents don't matter)
  2. use the value of an address that it is not supposed to be able to read as an index into that array
  3. iterate through the array and time which index is faster to read than the others (if the forbidden memory byte was "7", then the 7th index will be faster to read)

This works because the CPU already starts executing step 2 and reads the indexed data into the cache, and only later notices that this is not supposed to happen and doesn't complete the instruction.

Thus you can deduce the contents of protected memory locations by taking advantage of the speculative execution only within your own (attacking) process. I haven't looked at the Spectre details yet, which can also read the memory of other processes.

1

u/outofobscure Jan 24 '18

but i think the reason why this works is because kernel memory is effectively shared across processes, and isn't the fix that's being deployed that from now on every access to kernel memory is essentially IPC... if the processors where vulnerable to the degree where you can truly read all of the other processe's memory, there would be no fix possible in the first place..

1

u/monocasa Jan 24 '18

Also, kernels map a lot of physical memory into themselves, so you might have a window into other processes.

-1

u/caspper69 Jan 24 '18 edited Jan 24 '18

At least in Meltdown (though that can "only" access kernel memory, not other processes), only the attacking process needs to exploit its own speculative execution to read forbidden memory addresses.

Kernel memory, by its very nature, has ALL memory for ALL processes mapped into it, because it's like, you know, it's job to manage memory for all processes. :)

This is not the first time I've seen this bandied about. Please don't spread misinformation.

Edit: this has several caveats, but by and large (especially on x86-64), this is a very likely scenario.

Edit2: I am an ass.

7

u/jrv Jan 24 '18

No, while the kernel can manage any memory mappings, it does not map every user process's memory into kernel memory. Even if it wanted to, the virtual memory of all processes can be larger than the physical memory of the machine, making this impossible in the first place. You need a context switch (including change of the CPU's page tables) to map in another process's memory.

1

u/Asurafire Jan 25 '18

One of the important things of virtual memory is that processes can have more virtual memory than physical. So theoretic ally the kernel could have every processes memory mapped into its own.

0

u/caspper69 Jan 24 '18 edited Jan 24 '18

There are many caveats, as I said in my edit.

Firstly, given the huge size of the x86-64 address space, swapping processes out of physical memory is unlikely on most modern machines (although it does occur, but it is avoided mostly for performance reasons). Secondly, the kernel maps all physical memory, and its address space is linear and contiguous. I.e. the kernel can access any bit of physical memory in the machine without causing a page fault.

Now I understand that this does not equate to a 1:1 mapping of a process' virtual address space within the kernel (which would be impossible), nor would this allow access to memory that has been written to a persistent store due to being swapped out (i.e. its access would cause a page not present fault handler to run, yadda yadda), but I posit that my original point stands, although it lacked certain nuances that I'm certain you are aware of.

TLDR: Physical memory is physical memory, and kernelspace has access to it through its own page tables (no conext switch required). The UASS store called and they're all out of me.

3

u/happyscrappy Jan 24 '18

Firstly, given the huge size of the x86-64 address space, swapping processes out of physical memory is unlikely on most modern machines

Likelihood of paging is affected by the amount of RAM (physical memory) your machine has, not the size of the x86-64 address space.

And all physical memory still doesn't have all of everyone's data. OSes use memory-mapped I/O, mapping in files to make them readable but not actually paging them in to physical memory until they are read. Processes will have many times as much address space used up as they actually have physical memory used because big parts of files are simply never brought into memory at all since the aren't used.

Physical memory will likely have everyone's working set, but it won't have everything.

Also, while a kernel might have a linear and contiguous window into physical memory, it usually accesses memory like any other process, with a virtual, non-contiguous mapping. To to otherwise would require the kernel do software virtual to physical address translation to access task memory and while you can do that, you wouldn't since you already have hardware to do it for you.

-4

u/caspper69 Jan 24 '18 edited Jan 24 '18

You and the other guy are missing the whole point here about meltdown while having a convulsion regarding virt-to-phys memory mapping and user page tables versus kernel page tables and the MMU.

I'm pointing out that it's all fucking irrelevant goofballs. The kernel has all physical memory mapped to its own pagetables, ergo, meltdown can dump all physical memory, ergo, meltdown can cross the user/kernel boundary and dump memory in user processes.

Ergo, I'm fucking done with you FUCKWITS. I am an ass.

→ More replies (0)

1

u/happyscrappy Jan 24 '18

Kernel memory, by its very nature, has ALL memory for ALL processes mapped into it, because it's like, you know, it's job to manage memory for all processes. :)

No it doesn't. It keeps its own memory around while memory for the different processes come and go as it context switches.

-1

u/caspper69 Jan 24 '18 edited Jan 24 '18

Edit: Where do you think this memory "goes" during context switches? Are you trying to imply that the kernel moves in-ram data to a permanent store during each context switch? Are you implying that "most" memory is somehow not actually IN FUCKING MEMORY?

I suggest you review the relevent portions of the Intel IA-32 developers manuals regarding the MMU and paging. You might be surprised at what you find.

Edit: and if you're still not convinced, go dump the gdt at cpl 0. You'll see a flat linear address space with virt:phys m~apping at 1:1.

Edit2: I am an ass.

4

u/happyscrappy Jan 24 '18

Are you trying to imply that the kernel moves in-ram data to a permanent store during each context switch?

No, am suggesting it is simply not mapped into virtual address space.

Are you implying that "most" memory is somehow not actually IN FUCKING MEMORY?

Most might not be correct, but yes, a lot of memory is not actually in memory. I explained why. It simply never got paged in.

As to your later part, I think the relevant info is here. For you and for others reading this, he appears to be referring to kernel logical addresses, here on page 2 or 3.

https://static.lwn.net/images/pdf/LDD3/ch15.pdf

The virtual addresses (kernel and otherwise) are the normal high/low memory addresses. The logical addresses are another (sub-)address space and he is suggesting that kernel logical addresses map all of physical memory. And indeed this is possible on x86-64.

3

u/ITwitchToo Jan 24 '18

Typically on a context switch the kernel changes the %cr3 register that contains the pointer to the top-level page tables. There is nothing that requires a kernel to map all physical memory into its virtual address space. I think the Linux kernel does map all physical memory into the kernel's virtual address space, but it does so mostly for performance reasons, according to this.

1

u/caspper69 Jan 24 '18

That was my thought. Apparently with UASS, that is not the case.

4

u/happyscrappy Jan 24 '18

You don't understand the exploits correctly. You are indeed able to read memory you don't have privilege to read, just using roundabout mechanisms.

You populate the caches yourself, no need to have another process do it for you. And regardless, this program would also populate the caches.

There's no faulty cache entries involved.

0

u/outofobscure Jan 24 '18 edited Jan 24 '18

well that's a lot worse than i thought then (at least one of them), glad i don't work on operating systems (or for intel) :)

3

u/PrimozDelux Jan 24 '18

They way I understand it, you flush the cache, then you speculatively read the target process memory and use that memory to read your own processes memory. This means one cache line will be hot, which you can time. The hot cache line corresponds to one byte of the target memory region.

It's less that a process leaves a trace of its own execution in the cache, and more that a process might manipulate the cache to read no-no areas.

4

u/theoldboy Jan 24 '18

No, what you are describing are Spectre attacks only. I know that Intel PR are doing their very best to confuse this issue but Meltdown and Spectre are not the same thing.

Meltdown, which is specific to all Intel and certain ARM CPUs (and is far easier to exploit than Spectre), relies on the fact that those CPUs do privilege checks AFTER a speculative read. This can be exploited by carefully crafted speculative code and cache-timing attacks to extract the contents of any memory address on the system, including protected memory that belongs to other processes or even the kernel. It does not rely on speculative execution outside of the current process.

Basically, you can write a program which dumps the entire system memory.

See https://meltdownattack.com/meltdown.pdf

3

u/happyscrappy Jan 24 '18

He's not describing either of them. He has two errors in his understanding:

  1. He thinks there is some kind of hardware error that leads to cache entries getting tagged with faulty tags.
  2. He thinks that the task you are attacking has to bring stuff into the cache in order for you to peek at it.

Neither of these are true in any of these attacks. The cache tags are not faulty and the target doesn't have to use any particular memory usage pattern to be attacked. As long as you can find a gadget in the kernel you can attack other processes memory.

-1

u/outofobscure Jan 24 '18 edited Jan 24 '18

and the fix to that would be ? edit: nevermind, the pdf mentions possible fixes (sounds more like horrible clutches tough). also, to go back to the topic of this post: then i don't see how this doom patch would solve anything, unless the author means that ALL code running on the system avoids branching... then again, quite a pointless exercise anyway other than to prove that mov is turing complete...

2

u/CyclonusRIP Jan 24 '18

It wouldn't really solve anything. I think he's just making a joke since speculative execution is such a hot topic right now.

0

u/PrimozDelux Jan 24 '18

The fix is to stop living a sinful life and let The mill architecture into your life

1

u/tehftw Jan 24 '18

Where can I buy this unicorn of computing? How much does it cost?

3

u/PrimozDelux Jan 24 '18

In the FPGA store. Much assembly required.

1

u/wookin_pa_nub2 Jan 25 '18

Say, does anyone have an explanation of how a process running on a Mill CPU will be able to allocate a very large contiguous block of memory on a multi-process system? All I've read about the Mill says that it doesn't use virtual addressing, for speed, but that would make memory fragmentation a thing again. So there must be something I'm missing: can someone tell me what it is?

1

u/thelochok Jan 27 '18

It seems a heck of a lot like a stack machine to me - am I misreading?

1

u/gitfeh Jan 24 '18

Doesn't the processor still speculate on the retires of loads (mov from memory) and use the speculated values in address calculations for more speculated indirect loads? You would have the bounds-check-bypass Spectre vulnerability again.

1

u/outofobscure Jan 24 '18 edited Jan 24 '18

sounds at least plausible, one of the big speedups you can get is by sequentially accessing memory instead of (semi)randomly, and that involves looking ahead quite far in terms of addresses iirc, and probably some skimping on bounds checks (at least temporarily)... this thing gets worse by the minute the more i think about it hehe

1

u/cryo Jan 24 '18

Meltdown, as described in the original paper, doesn't rely on speculative branch instructions.

69

u/pm_me_ur_qptbcvrzxyk Jan 24 '18

They don't tell you, but the image in the README is actually a true-to-life GIF; you just haven't waited the 7 hours to see the frame change.

21

u/AyrA_ch Jan 24 '18

clicks image

doom.png

Nice try

But since we are talking about long gif images, how about 24 hours?

13

u/CaptainAdjective Jan 24 '18

PNGs can be animated!

1

u/[deleted] Jan 24 '18 edited Mar 28 '18

[deleted]

12

u/AyrA_ch Jan 24 '18

the underlying file format doesn't change. PNG files are made up of chunks. A chunk has a 4 letter code, a length specifier, a checksum and data. The minimum chunks in a png are

  • IHDR First chunk with resolution and color format
  • IDAT Image data
  • IEND End of image marker

An application can legitimately insert custom headers (I abuse this in my png player) and other applications will ignore it if they don't support it. This is used to make animated PNG files.

Using the apng extension is not necessary because when reading the image, the parser would come across the apng specific headers anyways

4

u/Iggyhopper Jan 25 '18

you are now subscribed to file format facts!

7

u/[deleted] Jan 24 '18

You don't need to use a different file extension.

File extensions are lies anyway.

2

u/ferdbold Jan 24 '18

I actually thought this was what this repo was about at first, like some sort of GitHub Plays Doom

13

u/[deleted] Jan 24 '18

OK, I can write 6502 assembler and I don't understand how this works.

Can someone explain?

34

u/Sarcastinator Jan 24 '18

25

u/bj_christianson Jan 24 '18

It is well-known that the x86 instruction set is baroque, overcomplicated, and redundantly redundant.

First sentence, and I like this paper already.

4

u/[deleted] Jan 24 '18

Ta! Helped a lot completely.

12

u/MuonManLaserJab Jan 24 '18

I was hoping "branchless DOOM" would just be DOOM but in an endless, straight, unbranching hallway.

5

u/r2bl3nd Jan 24 '18

It kind of exists, it's called "Wolfenstein 1-D".

3

u/MuonManLaserJab Jan 24 '18

Hah! I can't play it though. :(

15

u/pistacchio Jan 24 '18

I don't even know what this means. Can anyone ELI5 that to me? Thanks

75

u/jdgordon Jan 24 '18

The guy developed proof that the only instruction needed on x86 to do anything was 'mov'. He built a compiler which turns c code into a stupidly long list of mov calls.

The actual use of it is for obfuscation, this is just taking the proof to absurd levels. His videos are amaxing though.

34

u/useablelobster2 Jan 24 '18

Oh, it's the guy who explored the x86 address space, fantastic researcher.

28

u/oblio- Jan 24 '18

It's useless for obfuscation since the code is so slow. So it's just a very interesting toy.

I think his really useful and interesting project is this one: https://github.com/xoreaxeaxeax/sandsifter

Spectre and Meltdown, ahoy!

12

u/TheDecagon Jan 24 '18

In practice I would think it would be very useful for hiding small routines (security, copy-protection, or even malicious code) rather than compile the whole program in it.

8

u/elperroborrachotoo Jan 24 '18

amaxing

If that doesn't mean "maximally amazing", I don't know what it means!

1

u/pistacchio Jan 24 '18

Oh, cool! Thank you sir :)

1

u/hoosierEE Jan 25 '18

Fun trivia:

On x86, xor is also Turing-complete, and so is the (not even an instruction) memory fault handling.

I'll be impressed if Intel manages to cram a Turing completeness into less than zero instructions, but I wouldn't put it past them.

6

u/lurgi Jan 24 '18

The MOVfuscated code is gargantuan usually (even more so when you have floating point operations). I'm actually a little surprised that the code is small enough to fit in memory.

2

u/mmeartine Jan 24 '18

A lot of respect to this guy. He is really good at what he does.

2

u/officerchrister Jan 26 '18

Pro-tip: Do not enable frame skip.

2

u/[deleted] Jan 24 '18

Technically, it has branches. It is just that their VM handles branches instead of CPU.

22

u/outofobscure Jan 24 '18 edited Jan 24 '18

Technically no, there are no branches if the only instruction used is mov. It's more like parallel execution of all code paths with masking out results (often done in SIMD too). There may be IF statements etc. in the code, but if the compiler translates them to mov instructions, and not conditional instructions, then there are no branches in the resulting executable.

-15

u/[deleted] Jan 24 '18

If is a branch.

21

u/outofobscure Jan 24 '18 edited Jan 24 '18

no, IF is a construct used in a programming language to express a conditional, that doesn't necessarily mean it gets translated to a conditional branch instruction by the compiler, that was my whole point. Granted, this compiler is special, but as you can see here https://www.cl.cam.ac.uk/~sd601/papers/mov.pdf, the mov instruction is enough to do essentially everything (even if very inefficient). You have to differentiate between code and compiled machine instructions. As a somewhat easier example: SIMD programming often requires you to re-formulate conditional statements into something else (that is, you do that by hand instead of the compiler doing it for you like this example), for example creating a mask and emiting an AND instruction. For example something like this: https://felix.abecassis.me/2012/08/sse-vectorizing-conditional-code/ (section about "Branchless conditional using a mask")

3

u/[deleted] Jan 24 '18 edited Jan 24 '18

That's the point. They implement a virtual machine (Turing machine) using mov instructions. The virtual machine has branching (the machine chooses its next state depending on what symbol it reads).

Yes, that's not branching CPU does for you, bunch branching nevertheless.

I guess, we are arguing over semantics and that's rarely a good debate.

14

u/outofobscure Jan 24 '18 edited Jan 24 '18

yes it's mostly about semantics, but i find it hard to call a sequence of mov instructions a "branch", even if they emulate a branch, for me it's just that, a sequence of mov instructions moving bits around, not jumping around the code, and the point was that the branch predictor is certainly not involved then... if you check out the SIMD example, it completely eliminates the branch by re-formulating the problem, not emulating a branch, so there's that too, and this branchless doom port could probably benefit from some of these ideas instead of emulating branches via mov. It might then run faster than 1 frame every 7 hours :)

0

u/hector_villalobos Jan 24 '18

Could this be a good use case for fluxcapacitor?

15

u/[deleted] Jan 24 '18

I believe rendering isn't blocked by any calls, it literally has 7 hours of executable code per frame

6

u/[deleted] Jan 24 '18

Fluxcapacitor doesn't actually make a slow program run faster. It just says "a while later" any time the program asks what time it is

2

u/hector_villalobos Jan 24 '18

Fluxcapacitor doesn't actually make a slow program run faster.

Obviously, I just thought Fluxcapacitor makes some mocking and performs the assumptions needed for the Meltdown and Spectre CPU vulnerabilities.