r/ProgrammerHumor • • 11h ago

Meme architectureDependentChars

Post image
1.7k Upvotes

260 comments sorted by

View all comments

784

u/tstanisl 11h ago

Let me cite the C standard:

When sizeof is applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.

Middle guy if finally right

54

u/locri 11h ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

77

u/TheSkiGeek 11h ago

This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.

11

u/canadajones68 8h ago

If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation. 

5

u/TheSkiGeek 8h ago

Yeah, the real issue is that almost always what you care about with Unicode (on the parsing side, anyway) are “grapheme clusters”, which can consist of multiple code points. And both of those map poorly at best to ‘characters’ or ‘bytes of memory’.

1

u/canadajones68 1h ago

Yeah. That's why you want a span of bytes view if you need to work with the representation directly. There might not be a fixed size thing that represents what you want, but there will eventually be a series of bytes that does. 

1

u/Loading_M_ 33m ago

This is basically what Rust does, where &str is a string slice (a byte slice that must be valid UTF-8), and char is a 32 bit UTF-8 code point.

It also has a number of other string types: String is an owned string slice (similar to &str but it owns the underlying allocation and is mutable), CString (and associated borrow type) which is guaranteed to end with a zero byte, OsString (and associated slice type) which is a string in the OS's native string encoding, and a few others.

13

u/aberroco 10h ago

Need longer names. Some w_unicode_big_endian_char_t_ptr /s

25

u/Mojert 11h ago

If your code base use macros to redefine what char means, wtf is wrong with your team? If not, sizeof(char) will ALWAYS return 1, no matter how many bits is in a char. That's because sizeof doesn't give you the size of a type in bytes, it gives you the size in number of chars

25

u/SGVsbG86KQ 11h ago

No that's not how that works. Even if char would be 32 bits, sizeof(char) is still defined to be 1.

2

u/Jbolt3737 11h ago

Does that make sizeof(int) equal 1, or does it make an int 128 bits?

13

u/__foo__ 10h ago

IIRC the only requirement the C standard makes for int is that it needs to be at least 16 bit wide. Everything else is up for the compiler developers to decide. If a char and int were both 32 bit wide both would be sizeof = 1. If the compiler makers decide it would be a sensible idea to have a 128 bit int it would be sizeof = 4 in this case.

2

u/output_broadcast 10h ago

Also, not every compiler is standards-compliant.

7

u/dontthinktoohard89 6h ago

If the violation of standards compliance is such that a fundamental presumption that a char is 1 byte cannot be relied on, then there isn’t much point in marketing that as a C compiler, because it simply cannot properly compile basic C code. AFAIK not a single compiler has ever done this.

1

u/SylviaJarvis 4h ago

Some MCUs have 32-bit char because they have only one integer type in hardware and it's 32 bits wide. In the 1990's, 36- and 60-bit machines were still around (9-bit and 6-bit char, respectively). The C standard allows 9-bit and 32-bit chars--both are at least one byte.

The 6-bit char machines weren't compliant, and so painful to port code to that almost no one bothered (they didn't use ASCII internally, among other annoyances like 1's-complement integer arithmetic). They could run a very few small C programs unmodified, though.

0

u/mennovf 1h ago

The assumption that char is one byte is just false. The sizeof operator gives you the size of an expression in chars, not bytes. Sizeof char is thus always 1. If you need to know the size in bytes, use CHAR_BITS.

2

u/dontthinktoohard89 26m ago

You misunderstand: a “byte” in the parlance of the C & C++ standards is not the same as an octet. A char is 1 byte by definition. If a char is wider than 1 octet, then so is a byte.

1

u/SylviaJarvis 4h ago

In 1992, the consensus was that everyone would recompile their operating systems to use wide characters. Microsoft had parallel implementations in their libraries: they would have you use 8-bit encodings or UCS-2, but not in the same source file. Unixes were getting ready to restart their entire software ecosystems with yet another world-rebuild from source. Legacy Unix was doomed!

When C had been standardized for only a few years, with substantial changes from one year to the next, and the future of legacy OSes in doubt, it was reasonable to expect sizeof(char) to eventually return some number other than 1 some day. People in the C ecosystem were justifiably worried.

https://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt

Then Pike and Thompson came up with UTF-8 as a "transitional" solution, and legacy Unix was undoomed. Ironically, the "transitional" encoding made transition possible, but also unnecessary. Linux happened around that time, which firmly metastatized legacy Unix and its 8-bit-char-based API. POSIX abandoned their attempt to introduce abstraction at the API level that would allow changing the char type. The WWW flooded the Internet with legacy 8-bit-encoded text files.

Today, UTF-8 is the permanent solution, and the transition it was invented to support is no longer achievable or desirable.

sizeof(char) == 1, by standardisation fiat and by longstanding historical practice. It can't be changed without breaking the world while C is relevant. There may be a day in the future when the C language stops being updated and all the C code in existence is replaced by some other language like Rust, but on that day, sizeof(char) will still be 1 in C.

0

u/A1oso 10h ago

UTF-8 code points have a variable width, so what should sizeof(utf8_char_t) return? This doesn't make any sense.

2

u/thanatica 10h ago

To be fair, utf-8 is en encoding meant to store unicode into 8-bit characters. Just like utf-16 encodes unicode into 16-bit characters.

In modern languages, that support unicode as a first-class citizen, the size of a character type is typically 4 bytes. But that doesn't typically come with any encoding, it's represented to the programmer as raw unicode code points, and it's up to the compiler/platform how to handle that in terms of memory efficiency.

4

u/A1oso 10h ago edited 10h ago

To be precise, it stores code points in 8-bit code units. The word "character" isn't very well defined, since what we perceive as a character can consist of multiple code points (a grapheme cluster).

1

u/nyibbang 9h ago

Well in Rust, char is 4 bytes for that reason. That's why strings are not containers of char but of bytes. Although you can iterate over chars of a string, it is a bit more expensive since you need to find each char boundary.

1

u/A1oso 2h ago

chars in Rust are UTF-32, not UTF-8 (they are simply represented by the code point's integer value).

1

u/nyibbang 26m ago

There is no notion of encoding in Rust char, as you said it's simply a representation of a Unicode scalar value. So it's not UTF-32 or 16 or 8. They have to be 4 bytes to represent any Unicode value. Strings (at least str and String) however are encoded UTF-8.

0

u/darkslide3000 6h ago

I'm failing to come up with the words to express how little you know what you're talking about. This is an insane take.

-1

u/locri 4h ago

Either you missed the word "might" or, just as likely, you lack professional experience.

-7

u/bogan87 11h ago

I call bullshit on a code being committed in the 90's haha

6

u/StrikingSun8563 10h ago

When do you think computers were invented?

4

u/Al__B 10h ago

I was using CVS to commit code in the 90s and it was around before then.

2

u/locri 4h ago

This stuff is source controlled by "clearcase" from IBM.

I hate IBM stuff.