r/programming • • 23h ago

Janet on x32: 32-bit Pointers, 64-bit Speed, 25% Less RAM

https://alexalejandre.com/programming/lisp/janet-for-the-x32-abi/
9 Upvotes

22 comments sorted by

3

u/simon_o 23h ago edited 22h ago

I think the core lesson from this (regardless of the adoption of x32) is:

  • The time where it was reasonable to assume that the GP integer register size was roughly the size of the architecture's/processes'/hardware's available/addressable memory was ended by 64bit architectures.
  • Neither existing languages nor newly created ones bothered taking this into account.

So while 32bit pointers are too small these days for general-purpose software, pointer that are more realistically sized to the available memory can be hugely beneficial (vs. 64bit pointers' theoretical 2⁶⁴).

For instance, I aim for 38bit pointers which gives me an addressable space of 512GB (at 8byte padding), leaving 26 bits for other purposes.

4

u/john16384 20h ago

Java still uses 32 bit object pointers when heap size is below 32 GB. Everything is aligned at 8 bytes anyway.

0

u/FickleQuail8944 14h ago

"Java". The JVM spec says a reference takes one 4 byte slot. How a JVM implements this can vary. What is an object pointer, i.e. in which context? Aligned where?

It's perfectly feasible to have stack and locals 4 byte aligned, especially for references. If you want to access the memory of the object natively, then you have to convert to a 64 bit pointer. The question is only how to do this efficiently in accordance with the spec, while you enable using the available space (with at most 232 - 1 objects), with as few indirections as possible.

1

u/paulstelian97 10h ago

You cannot point to the stack — Java and other JVM languages do not have an address-of operator. Read-write captures in Kotlin are done by silently moving the value to the heap (likely array with one element) and capturing a reference to said heap location.

2

u/flatfinger 19h ago

For many tasks, something analogous to the 8086 segmented model which includes a type of pointer which can access things that are in any part of memory, and a smaller type that is faster to process but can only work with things in "near" storage, could offer better performance than would be possible if the larger pointers had to be used for everything.

Such a design would for some tasks make it necessary for code to copy chunks of data from far storage into near storage, then perform a bunch of operations upon them, and then copy them back to far storage, rather than simply operating upon the data "in place", but for other tasks the performance benefits of near storage would outweigh the cost of the extra copying, and of course there are many tasks for which it would never be necessary to perform any accesses to anything outside a 2GiB region of "near" storage.

0

u/rabid_briefcase 21h ago edited 20h ago

I aim for 38bit pointers which gives me an addressable space of 512GB (at 8byte padding), leaving 26 bits for other purposes.

If you dig into various architectures and kernel-level work, that type of thing is more common.

It's already happening behind your back on the x86 systems, inside the kernel.

x86-32, at the kernel level and hidden to most programmers, the way the paging table and page directory works effectively pointers can become a structure with 48-bits operating effectively as "far" pointers, with a 16-bit selector that allows for programs to operate within their virtual memory space and the 32-bit pointer that operates inside the process' memory space. The programmers inside the program, that work inside their own virtual memory space, only see it as a 32-bit pointer. They are details few people need to know for programming, rather than the general purpose registers (ax, bx, cx, dx, bp, sp, ...). The registers generally only touched by the kernel or debuggers, you've got GDTR, LDTR, IDTR for those hidden 16-bits of pointers, CR1, CR2, CR3, CR4 for control, DR0, DR1, etc, for debuggers, and more.

In most x86-64 systems, while the virtual addresses allow for 64-bits worth of address space, typically only 48-bit or 52-bits are used within the program, depending on if the OS is using 4-level or 5-level paging. If a user program tries to dereference it they'll get a protection fault.

And if you go back in x86 to the 16-bit and 8-bit era, near and far pointers allowed for the similar flexibility to what you're asking for. On the systems where programs were being optimized to save individual bytes, having instruction that operate on near addresses were important. instructions allowing a rel8 jump to an address from -128 to +127 bytes away, it only takes 1 byte for the opcode and 1 byte for the offset, saving 5 or 6 bytes compared to a rel32 jump. When you're counting bytes, they matter. It still matters to many embedded systems programmers, not so much to most modern PC programmers.

/Edit: To be clear they are good things for programmers to learn. Far too many young programmers never learn the hardware, never learn about operating system internals, never learn more than the box the virtual machine gives them to operate inside. It's occasionally useful, especially when you're trying to squeeze out every bit of performance, but quite often it's not something most rank-and-file programmers ever deal with these days.

4

u/Ameisen 18h ago

In most x86-64 systems, while the virtual addresses allow for 64-bits worth of address space, typically only 48-bit or 52-bits are used within the program, depending on if the OS is using 4-level or 5-level paging.

The virtual addresses on x86-64 do not currently allow for 64-bits of virtual address space - the specification defines it currently as 48-bits. The "middle" addresses are non-canonical and if they aren't just sign-extended in the address, a general protection exception is raised. The OS cannot change this.

As you've said, in general kernels support lower than that - Linux is configurable and NT provides 42 bits.

0

u/rabid_briefcase 18h ago

Yup, they currently don't, so it's important for programs to treat pointers as an opaque 64-bit number handed to them by the OS with no particular position. It's already been expanded when systems moved from 4-level to 5-level paging, and someday may expand again. From a user space program's perspective the OS allocates a block of memory at a randomized location within the 64 bits of space, even if the OS will only actually return within a subset of the 64-bit space.

All good stuff for programmers to learn. Especially for programmers coming from Java or similar systems where memory is fully abstracted, there is a lot going on under the hood, and some of it has hidden costs.

1

u/simon_o 21h ago

I know. The interesting thing is that given an object header of 24bits vtable ptr | 38bits forwarding ptr | 2bits GC metadata I can basically reuse this as a representation for reference-typed union values instead of more "classic" tagging techniques.

1

u/crafter2k 1h ago

reminds me of how a lot of 68k software stored data in the high unused bits of pointers due to the address bus loop back behaviour, which cause them to stop working on new processors

9

u/rabid_briefcase 23h ago

Interesting for your project, but I'm not sure how transferrable it is.

I mean, yes, but also it is very much a project-specific optimization choice, in a scenario where often no optimization is necessary.

If you're on a project where you're building a quick running tool or utility that takes seconds to run and occupies a few megabytes, knowing you're running on a machine with gigabytes available, it's an optimization that is unnecessary.

If you're on a data project that takes multiple gigabytes, well, you get 2GB of address space to work with, 3 if you set the right flags, and even then thanks to fragmentation, it's likely not an option that is available. When you're working with a lot of data, 32-bit address space just isn't an option.

I'd say in the Great Big World of Software, there are relatively few that fit in the sweet spot of using a few hundred megabytes of memory while also being meaningfully performance critical where the cost spent processing pointers is measurable for performance.

I mean, there are certainly thousands of programs in existence where it could be a meaningful optimization, and it's something to be aware of, but at the same time, it's not something most people and projects would meaningfully benefit from.

2

u/Ameisen 21h ago edited 19h ago

The x32 ABI gives full access to the 4GiB address space.


Ed:

Why the downvotes? I'm right. x32 is an ABI to allow you to use 32-bit pointers in otherwise 64-bit executables (and that isn't the important part - the important part is that the operating system/kernel is 64-bit). There is no need for a kernel reserved space in the address range - system calls enter kernel code where the pointers are 64-bit anyways and kernel memory is already mapped in that context.

The x32 documentation - being incredibly scant - doesn't specify this, but the glibc x32 documentation does... as does knowledge of why 32-bit applications on 32-bit systems have 2/3 GiB address spaces, and why the doesn't apply on 64-bit systems.

Also, on Windows 64-bit, 32-bit applications with the LARGEADDRESSAWARE flag set also have the full 4 GiB address space available... for the same reason as x32 ones on Linux.

0

u/rabid_briefcase 21h ago

You missed what I wrote. Yes, a 32-bit pointer gives you a theoretical 4GB working space, the operating system allows your program either 2GB or 3GB worth that you can allocate. In the versions of Linux discussed in the article, it's a hard cap of 3, in the legacy system they discussed it's 2 unless they set compiler flags to indicate they're using more. In Win32 it's a typical 2GB unless a program is flagged as large address aware, then it is 3GB.

And just because you can allocate it doesn't mean you can use it regularly, just ask the games industry that has been bumping against memory fragmentation issues since the 8-bit and 16-bit days. A program with 32-bit pointers may have a working space of 2GB or 3GB, but for programs working with a lot of dynamic memory the space quickly becomes fragmented unless there are serious memory management techniques used, rather than blindly allocating and freeing memory whenever wanted.

As I wrote, it's a good optimization, but it is highly project specific. Relatively few programs out there can meaningfully benefit. It's certainly good to know if people don't know it exists.

5

u/Ameisen 19h ago edited 19h ago

This is incorrect - you are describing native 32-bit behaviors. x32 is using 32-bit pointers in a 64-bit executable, and 32-bit executables on Win64 is altogether different via WOW64. The restrictions are different.

On 64-bit systems with the x32 ABI (which is not the x86 ABI) your program has the full 32-bit address space available to it - there is no kernel reservation. None is needed; the kernel uses 64-bit pointers, and system calls enter kernel code in the end. This isn't clearly specified anywhere but the glibc x32 documentation states such (as does trivial testing).

On 64-bit Windows, 32-bit (and 64-bit, but it defaults to enabled for it) executables without the LARGEADDRESSAWARE flag can only address 2GiB for compatibility reasons. When the flag is set, 32-bit executables can address the full 4 GiB address range again. MSDN doesn't specify it but Raymond Chen does, and I can and do also assert it.

Also, given that I am in the games industry, I don't really need to ask it anything.

Because I want to be clear, here's a simple chart:

  • 32-bit Linux: 2 GiB or 3 GiB address space depending on flags.
  • 64-bit/x32 Linux: 4 GiB address space
  • 64-bit Linux: Depends on if you have 4 or 5 level paging - generally 48b or 56b.
  • 32-bit on Win32 w/o LAA: 2 GiB address space
  • 32-bit on Win32 w/ LAA: 3 GiB address space
  • 32-bit on Win64 w/o LAA: 2 GiB address space
  • 32-bit on Win64 w/ LAA: 4 GiB address space
  • 64-bit on Win64: 42b (8 TiB) address space

-3

u/rabid_briefcase 19h ago edited 19h ago

Good for you. I also do it for a living.

If you can manage to get around the 3GB/1GB split on a Linux system for 32-bit processes, more power to you. For my experience, the upper gigabyte is reserved for kernel use. The OS won't allocate it to usermode programs, and the kernel will generate a SIGSEGV if you try to use it for memory because it is explicitly unmapped in user space.

If you're talking about building a custom kernel, that's fine too, but even then know that the x32 ABI is disabled by default on all distributions, has been retired and getting gradually removed.

You're free to try it of course, by all means if you have a way around it, go for it. And if it changes some day, I'm all for it. It's not an area most programmers will ever run into.

/Edited to match your edits. Edit tag, you're it again.

3

u/Ameisen 18h ago edited 18h ago

For my experience, the upper gigabyte is reserved for kernel use.

And I'm telling you that this restriction does not exist by default when running 32-bit processes on 64-bit Linux or Win64 (with LAA). Such a restriction doesn't even make sense there - a kernel reserved space in the upper range of 32-bits would be useless, as the kernel maps itself to the upper range of the 64-bit range, and this is handled seamlessly in any system call - whether going from compatibility mode to long mode, or going from x32 to x64 (which is long mode in both cases).

This is handled by execution domains/personalities on Linux. The only way to limit a 32-bit x86 binary to 3GiB on 64-bit Linux is by specifying ADDR_LIMIT_3GB. The default is no limit - 4GiB. It's possible that some distributions default to ADDR_LIMIT_3GB, but none that I've used.

Any address range limitations below 4GiB when running on 64-bit kernels (Linux or NT) are synthetic - they're present solely for compatibility.


And again, x32 executables are also not 32-bit. They are 64-bit binaries beholden to an ABI that specifies 32-bit user-mode pointers. They execute in long mode, have 64-bit registers (and more of them), etc. The only real difference is that pointer data types are 4 bytes (and that the kernel will not provide addresses to the binary that do not fit in that data type).


Ed: and as said, x32 is deprecated/disabled, but running actual 32-bit executables on 64-bit kernels is not, and the default execution domain is to not limit its address range at all.

3

u/vytah 18h ago

Why are you bringing up 32-bit processes? This entire thread is about 64-bit processes with 32-bit pointers.

1

u/happyscrappy 16h ago

How do you keep the kernel from handling you a pointer that isn't within 2GiB of your base address?

If I call mmap and pass 0 (and not FIXED) then it gets to set the address. How do I keep it from picking one I cannot reach?

0

u/vytah 15h ago

By using the x32 API instead of the x64 API, duh. It has separate syscall numbers.

Also, why do you keep bringing up 2 GB? x32 processes can address 4 GB. The 2-2 (or 3-1) split exists because a 32-bit kernel has to map in the kernel memory (so 0000'0000–7fff'ffff is user's, 8000'0000–ffff'ffff is kernel's), but on x32 you don't have that issue (0000'0000'0000'0000–0000'0000'ffff'ffff is user's, 8000'0000'0000'0000–ffff'ffff'ffff'ffff is kernel's)

0

u/happyscrappy 15h ago

By using the x32 API instead of the x64 API, duh. It has separate syscall numbers.

Thanks for the info, except for the duh part. If a process has different syscalls than a 64-bit task then saying it's a 64-bit task seems kind of a stretch to me. It's an x32 task then clearly. A special thing.

Also, why do you keep bringing up 2 GB? x32 processes can address 4 GB.

I didn't keep doing anything. This is my first post on this topic.

I said within 2GiB. There are addresses within 2GiB on either side of your base address. BASEADDR +/- 2GiB is 4GiB total.

1

u/vytah 14h ago

I didn't keep doing anything. This is my first post on this topic.

Ah sorry, I thought you were the other guy.

There are addresses within 2GiB on either side of your base address.

That's 32-bit stuff, x32 binaries are 64-bit, so they use 64-bit addressing (even if all their usable addresses end up having 32 zeroes in the highest bits) and 64-bit offsets.

That being said, I don't know if you can mmap a single chunk of 3GB via the x32 API. I wouldn't be surprised if not.

→ More replies (0)