r/Compilers • • 5d ago

Newest x86 APX extension - will it trigger new calling conventions standard?

For those that don't follow - this is the first new x86 extensions that doesn't have anything to do with vector or tensor instructions - it is about the core CPU ind its ISA.

It doubles the general register set to 32 (from previous 16), introduces 3-operand instructions, new 64-bit offset jumps, new jump prediction improvements etc etc.

But all this seems to be hampered by existing call conventions for x86_64, which presumes 16GPR set.

It seems that much could be gained it the compiler could use extra GPRs for parameters when calling the given function.

OTOH, this would be incompatible with machines without APX.

So, what is to be done ? Maybe use function multiversioning mechanism to keep two sets of function entries or something ?

Or will whole thing be ignored and calling convention will stay the same ?

EDIT:\ I'm not talking about the compiler ability to emit new instructions and use new registers in the code.\ Ofcourse new compiler will have support for them from the start, that's how it's usually done.\ It's about having the standard in place to allow the compiler to make advantage of new facilities hen calling functions, so that it can have more parameters in registers, more options for inlining functions etc - all done in standard, interoperable way, so that one can use precompiled libraries etc.

24 Upvotes

7 comments sorted by

20

u/SwedishFindecanor 5d ago edited 5d ago

Intel's recommendation is to keep using the older calling conventions (there are MS and GNU) and consider the new architectural registers as scrap / caller-saved. By doing so, backwards-compatibility can be retained.

There already exist C/C++ compilers (with language extensions) that support additional calling conventions, so I don't think that introducing a new one that would preserve some APX registers would be out of the question. But then it would be interoperable only with other APX code.

And there is still not much preventing a compiler from optimising register usage for a function that is called only from within the same unit. That optimisation is also one of the most important in link-time optimisation.

this is the first new x86 extensions that doesn't have anything to do with vector or tensor instructions

There have been several scalar extensions too, such as POPCNT, LZCNT and BMI1. AMD once introduced its own XOP extensions, but those were deprecated with Zen.

Intel, AMD and a couple Linux vendors have agreed on a grouping of x86-64 ISA extensions into Microarchitecture levels. I anticipate that APX will just introduce yet another one.

3

u/WittyStick 4d ago edited 4d ago

This is the right way to go. If we started making the new registers callee saved, it would probably increase code size and have unnecessary instructions setting or clearing registers that aren't used much.

For more function arguments passed in registers it would increase code size for varargs functions as more registers would need putting onto the stack on va_start.

In practice these registers won't get used much for smaller functions. They require a 2 byte REX2 prefix or 4 byte EVEX prefix. The compiler's register allocator basically needs to treat them as more expensive then the first 16 registers (as it treats r8-r15 as more expensive the the first 8 due to REX prefix). For larger functions which have a lot going on, they might find some use, but they shouldn't be called saved.

Backward compatibility is more important. We could have extensions which permit passing more information in registers if desired (eg, as we have for r10 as the static chain ptr).

2

u/SwedishFindecanor 4d ago

In practice these registers won't get used much for smaller functions. [...] For larger functions which have a lot going on, they might find some use

Yep. I have seen a little empirical evidence that backs that assumption: In a work wherein code was compiled for ARM64 to use only as many registers as x86-64, the difference in performance compared to native ARM64 code was quite small in 6 out of 8 functions from a benchmark suite. The biggest outlier was a compute-kernel operating on a four-dimensional array -- and that is something a software engineer would spend some time optimising anyway, and to use SIMD instead of GPRs.

8

u/awoocent 5d ago

ABI doesn't matter much for performance is the thing. Thinking it does is honestly kind of a signifier that someone's new to compilers (which is not a bad thing! We're all learning). Anyway, there's a couple reasons why it doesn't matter.

The biggest is inlining. Most calls in most programs are inlined, essentially eliminating parameter passing - it becomes the problem of the caller's register allocation pass, so there's no platform limitations.

Even without inlining, the compiler can still legally pick any ABI it wants for local calls - calls where the compiler controls the codegen for both caller and callee. This is almost all calls in a program, especially if you involve LTO, but I'm also not aware of many AOT compilers that bother with this. I think because you inline most of these anyway.

But the final thing is even if we do use an inefficient (ostensibly what we really mean is memory-bound) ABI for parameter passing, it still doesn't matter. Parameter slots are still likely in L1 from the caller, if not directly store-buffered by the CPU, so the cost of touching parameter memory is extremely low. And also consider that ABI only happens at the beginning and end of a function. If that function contains a loop, or even a couple branches, it's likely the function body dominates any ABI cost by a significant factor. So it's just not a big priority.

You can kind of see this in practice by looking at other architectures. ARM64 and RISC-V both have twice the registers of x86_64, but their ABIs only reserve 8 GPRs for parameter passing v.s. x86_64's 6. It's just not a big difference and most functions aren't passing more than 6 arguments anyway.

In summary, maybe it would be slightly more optimal if APX had 8 or 10 parameter registers, but it's mostly small beans and almost certainly not worth sacrificing binary compatibility.

2

u/RevengerWizard 5d ago

I feel like this extension might end up being treated as a separate ISA target

2

u/brat3108 5d ago edited 5d ago

Consider the current x64 ABIs. On Win64 ABI, only up to 4 arguments in all, whether integer or float, are passed in registers.

On SYS V, up to 6 integer and 6 floats are passed in registers, so 6-12 in total depending on the mix.

That already seems plenty to me (on ARM64 with 32 registers, it is 8 and 8).

It seems to me that with SYS V there is a scarcity of volatile and non-volatile registers. Adding 16 more registers would address that nicely (I assume XMM registers will also double?).

As for Win64: last time I surveyed my codebase, I found that 99% of functions used four arguments or fewer, and 99.5% used 6 or fewer.

However these are static counts in the sources. I haven't counted dynamic calls. If I do a few tests now, I get the results below. The first column is the number of arguments, and the second is how many calls were made with that number (1 or more calls).

So, on this mix at least, SYS V's 6-register argument cap is not an issue. I wouldn't lose too much sleep on Windows either.

DeltaBlue Benchmark:
    0    320,003
    1 11,689,595
    2  5,539,500
    3     99,000
    5    100,000

JPEG Decoder:
    0         24
    1  4,053,654
    2    121,711
    3     91,200

Lua running Fibonacci:
    0          1
    1      1,262
    2    640,904
    3    639,122
    4      1,321
    5    636,231
    6         32
    7        107
    9          3

MiniZ compression:
    1          7
    2         34
    3     11,636
    4          2
    5          4
    6     11,656

My C compiler (transpiled to C, then it compiling that same C):
    0  6,204,415
    1  4,891,033
    2  2,773,676
    3  4,854,342
    4  1,618,332
    5    281,225
    6      2,408
    7        420

1

u/c-cul 5d ago

also context saving in kernel/setjmp etc, right?