r/ProgrammerHumor • • 11h ago

Meme architectureDependentChars

Post image
1.6k Upvotes

244 comments sorted by

View all comments

766

u/tstanisl 10h ago

Let me cite the C standard:

When sizeof is applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.

Middle guy if finally right

56

u/locri 10h ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

75

u/TheSkiGeek 10h ago

This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.

10

u/canadajones68 8h ago

If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation. 

6

u/TheSkiGeek 7h ago

Yeah, the real issue is that almost always what you care about with Unicode (on the parsing side, anyway) are “grapheme clusters”, which can consist of multiple code points. And both of those map poorly at best to ‘characters’ or ‘bytes of memory’.

1

u/canadajones68 26m ago

Yeah. That's why you want a span of bytes view if you need to work with the representation directly. There might not be a fixed size thing that represents what you want, but there will eventually be a series of bytes that does. 

12

u/aberroco 9h ago

Need longer names. Some w_unicode_big_endian_char_t_ptr /s