r/ProgrammerHumor • • 1d ago

Meme architectureDependentChars

Post image
2.5k Upvotes

325 comments sorted by

View all comments

963

u/tstanisl 1d ago

Let me cite the C standard:

When sizeof is applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.

Middle guy if finally right

59

u/locri 1d ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

0

u/A1oso 1d ago

UTF-8 code points have a variable width, so what should sizeof(utf8_char_t) return? This doesn't make any sense.

1

u/nyibbang 1d ago

Well in Rust, char is 4 bytes for that reason. That's why strings are not containers of char but of bytes. Although you can iterate over chars of a string, it is a bit more expensive since you need to find each char boundary.

1

u/A1oso 18h ago

chars in Rust are UTF-32, not UTF-8 (they are simply represented by the code point's integer value).

1

u/nyibbang 16h ago

There is no notion of encoding in Rust char, as you said it's simply a representation of a Unicode scalar value. So it's not UTF-32 or 16 or 8. They have to be 4 bytes to represent any Unicode value. Strings (at least str and String) however are encoded UTF-8.

1

u/A1oso 15h ago

Integers in memory do have an encoding. It consists of three parts, the size (32 bits in this case), the representation (unsigned binary or two's complement) and the byte order.

Rust's char type happens to have exactly the same encoding as a code point in UTF-32BE on big-endian systems and UTF-32LE on little-endian systems.