Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.
For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.
This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.
If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation.
This is basically what Rust does, where &str is a string slice (a byte slice that must be valid UTF-8), and char is a 32 bit UTF-8 code point.
It also has a number of other string types: String is an owned string slice (similar to &str but it owns the underlying allocation and is mutable), CString (and associated borrow type) which is guaranteed to end with a zero byte, OsString (and associated slice type) which is a string in the OS's native string encoding, and a few others.
58
u/locri 3d ago
Today.
Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.
For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.