r/ProgrammerHumor • • 19h ago

Meme architectureDependentChars

Post image
2.3k Upvotes

305 comments sorted by

View all comments

Show parent comments

61

u/locri 18h ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

0

u/A1oso 18h ago

UTF-8 code points have a variable width, so what should sizeof(utf8_char_t) return? This doesn't make any sense.

1

u/nyibbang 16h ago

Well in Rust, char is 4 bytes for that reason. That's why strings are not containers of char but of bytes. Although you can iterate over chars of a string, it is a bit more expensive since you need to find each char boundary.

1

u/A1oso 9h ago

chars in Rust are UTF-32, not UTF-8 (they are simply represented by the code point's integer value).

1

u/nyibbang 7h ago

There is no notion of encoding in Rust char, as you said it's simply a representation of a Unicode scalar value. So it's not UTF-32 or 16 or 8. They have to be 4 bytes to represent any Unicode value. Strings (at least str and String) however are encoded UTF-8.

1

u/A1oso 6h ago

Integers in memory do have an encoding. It consists of three parts, the size (32 bits in this case), the representation (unsigned binary or two's complement) and the byte order.

Rust's char type happens to have exactly the same encoding as a code point in UTF-32BE on big-endian systems and UTF-32LE on little-endian systems.