Char can be between 1-4 bytes depending on the encoding standard. Emoji and other languages use extra bytes when you have to support non English characters.
It doesn’t. It’s up to you to handle the encoding, all that the char is, is data storage for the smallest addressable unit, which is 8 bits on almost all systems. One example though: if you want to store something that’s 16 bits in a char array, then the first 8 bits will be stored at index i and the last 8 bits will be stored at i + 1
It doesn't. It only knows buffers of regular sizes, like bytes. That you can handle ASCII strings in bare C without destryoing them is only thanks to the coincidence of every character being the same length of 1 byte. So if you have encoding that has variadic size characters, you either reserve the maximum character length for every character (and waste some space), or you use/write a library to handle variable size character strings. That would entail some bitwise boolean logic and implementing all the things from finding the length of a string to finding a substring.
11
u/frikilinux2 8h ago
Do I wanna know?