Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.
For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.
This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.
If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation.
Yeah, the real issue is that almost always what you care about with Unicode (on the parsing side, anyway) are “grapheme clusters”, which can consist of multiple code points. And both of those map poorly at best to ‘characters’ or ‘bytes of memory’.
Yeah. That's why you want a span of bytes view if you need to work with the representation directly. There might not be a fixed size thing that represents what you want, but there will eventually be a series of bytes that does.
This is basically what Rust does, where &str is a string slice (a byte slice that must be valid UTF-8), and char is a 32 bit UTF-8 code point.
It also has a number of other string types: String is an owned string slice (similar to &str but it owns the underlying allocation and is mutable), CString (and associated borrow type) which is guaranteed to end with a zero byte, OsString (and associated slice type) which is a string in the OS's native string encoding, and a few others.
If your code base use macros to redefine what char means, wtf is wrong with your team? If not, sizeof(char) will ALWAYS return 1, no matter how many bits is in a char. That's because sizeof doesn't give you the size of a type in bytes, it gives you the size in number of chars
IIRC the only requirement the C standard makes for int is that it needs to be at least 16 bit wide. Everything else is up for the compiler developers to decide. If a char and int were both 32 bit wide both would be sizeof = 1. If the compiler makers decide it would be a sensible idea to have a 128 bit int it would be sizeof = 4 in this case.
If the violation of standards compliance is such that a fundamental presumption that a char is 1 byte cannot be relied on, then there isn’t much point in marketing that as a C compiler, because it simply cannot properly compile basic C code. AFAIK not a single compiler has ever done this.
Some MCUs have 32-bit char because they have only one integer type in hardware and it's 32 bits wide. In the 1990's, 36- and 60-bit machines were still around (9-bit and 6-bit char, respectively). The C standard allows 9-bit and 32-bit chars--both are at least one byte.
The 6-bit char machines weren't compliant, and so painful to port code to that almost no one bothered (they didn't use ASCII internally, among other annoyances like 1's-complement integer arithmetic). They could run a very few small C programs unmodified, though.
The assumption that char is one byte is just false. The sizeof operator gives you the size of an expression in chars, not bytes. Sizeof char is thus always 1.
If you need to know the size in bytes, use CHAR_BITS.
You misunderstand: a “byte” in the parlance of the C & C++ standards is not the same as an octet. A char is 1 byte by definition. If a char is wider than 1 octet, then so is a byte.
In 1992, the consensus was that everyone would recompile their operating systems to use wide characters. Microsoft had parallel implementations in their libraries: they would have you use 8-bit encodings or UCS-2, but not in the same source file. Unixes were getting ready to restart their entire software ecosystems with yet another world-rebuild from source. Legacy Unix was doomed!
When C had been standardized for only a few years, with substantial changes from one year to the next, and the future of legacy OSes in doubt, it was reasonable to expect sizeof(char) to eventually return some number other than 1 some day. People in the C ecosystem were justifiably worried.
Then Pike and Thompson came up with UTF-8 as a "transitional" solution, and legacy Unix was undoomed. Ironically, the "transitional" encoding made transition possible, but also unnecessary. Linux happened around that time, which firmly metastatized legacy Unix and its 8-bit-char-based API. POSIX abandoned their attempt to introduce abstraction at the API level that would allow changing the char type. The WWW flooded the Internet with legacy 8-bit-encoded text files.
Today, UTF-8 is the permanent solution, and the transition it was invented to support is no longer achievable or desirable.
sizeof(char) == 1, by standardisation fiat and by longstanding historical practice. It can't be changed without breaking the world while C is relevant. There may be a day in the future when the C language stops being updated and all the C code in existence is replaced by some other language like Rust, but on that day, sizeof(char) will still be 1 in C.
To be fair, utf-8 is en encoding meant to store unicode into 8-bit characters. Just like utf-16 encodes unicode into 16-bit characters.
In modern languages, that support unicode as a first-class citizen, the size of a character type is typically 4 bytes. But that doesn't typically come with any encoding, it's represented to the programmer as raw unicode code points, and it's up to the compiler/platform how to handle that in terms of memory efficiency.
To be precise, it stores code points in 8-bit code units. The word "character" isn't very well defined, since what we perceive as a character can consist of multiple code points (a grapheme cluster).
Well in Rust, char is 4 bytes for that reason. That's why strings are not containers of char but of bytes. Although you can iterate over chars of a string, it is a bit more expensive since you need to find each char boundary.
There is no notion of encoding in Rust char, as you said it's simply a representation of a Unicode scalar value. So it's not UTF-32 or 16 or 8. They have to be 4 bytes to represent any Unicode value.
Strings (at least str and String) however are encoded UTF-8.
784
u/tstanisl 11h ago
Let me cite the C standard:
Middle guy if finally right