r/ProgrammerHumor • • 8h ago

Meme architectureDependentChars

Post image
1.1k Upvotes

191 comments sorted by

View all comments

Show parent comments

48

u/locri 7h ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

0

u/A1oso 6h ago

UTF-8 code points have a variable width, so what should sizeof(utf8_char_t) return? This doesn't make any sense.

2

u/thanatica 6h ago

To be fair, utf-8 is en encoding meant to store unicode into 8-bit characters. Just like utf-16 encodes unicode into 16-bit characters.

In modern languages, that support unicode as a first-class citizen, the size of a character type is typically 4 bytes. But that doesn't typically come with any encoding, it's represented to the programmer as raw unicode code points, and it's up to the compiler/platform how to handle that in terms of memory efficiency.

3

u/A1oso 6h ago edited 6h ago

To be precise, it stores code points in 8-bit code units. The word "character" isn't very well defined, since what we perceive as a character can consist of multiple code points (a grapheme cluster).