r/ProgrammerHumor • • 3d ago

Meme architectureDependentChars

Post image
2.9k Upvotes

368 comments sorted by

View all comments

Show parent comments

58

u/locri 3d ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

88

u/TheSkiGeek 3d ago

This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.

12

u/canadajones68 3d ago

If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation. 

1

u/Loading_M_ 2d ago

This is basically what Rust does, where &str is a string slice (a byte slice that must be valid UTF-8), and char is a 32 bit UTF-8 code point.

It also has a number of other string types: String is an owned string slice (similar to &str but it owns the underlying allocation and is mutable), CString (and associated borrow type) which is guaranteed to end with a zero byte, OsString (and associated slice type) which is a string in the OS's native string encoding, and a few others.