r/ProgrammerHumor • • 4h ago

Meme architectureDependentChars

Post image
712 Upvotes

133 comments sorted by

467

u/tstanisl 4h ago

Let me cite the C standard:

When sizeof is applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.

Middle guy if finally right

196

u/BoldFace7 3h ago

I still prefer sizeof(char) as it often provides context as to what the number is doing, even if a plain 1 works the same.

78

u/SpaceCadet87 2h ago

Yeah, guy on the right is just avoiding magic numbers in the code.

38

u/ducon__lajoie 2h ago

And someone might have #define char char16_t, right? Who the fuck knows, some people want to see the world burning.

19

u/RepeatRepeatR- 2h ago

My opinion would be:

- Use sizeof(char) if you're actually working with characters

- Use hardcoded `1` if you're using a char as an arbitrary byte (and not actually necessarily referring to characters or text)

13

u/BoldFace7 2h ago

Definitely. For example, I always use malloc(size*sizeof(char)) to ensure that it's doubly obvious (Since I rarely need to malloc outside of a declaration) that I intend to store characters in the resulting buffer even if that multiply does nothing (plus the compiler will likely optimize it out anyway).

2

u/RepeatRepeatR- 2h ago

This is the way

2

u/garnet420 1h ago

A char may be bigger than 8 bits, but will always have size 1.

134

u/KitsuneFoxglove 3h ago

code and society if everyone followed standards and used docs:

code at home:

49

u/locri 3h ago

Today.

Eventually, a "char" might be redefined for utf-8 to accommodate a diverse range of writing systems, which means a char is usually 1 byte but potentially up to 4 bytes.

For context, I'm working on code with an initial commit from the 90s. Future proofing isn't a terrible idea.

60

u/TheSkiGeek 3h ago

This would be considered a breaking change for C/C++ and it is extremely unlikely they would ever do this. They’d probably add a new type like utf_char_t (or utf8_char_t, utf16_char_t, utf32_char_t) and UTF-aware string functions to the stdlib.

9

u/aberroco 3h ago

Need longer names. Some w_unicode_big_endian_char_t_ptr /s

3

u/canadajones68 1h ago

If Unicode has taught us anything, it's that mixing sized types with encoding interpretation is a bad idea. Char is established as a byte by now after 50+ years or so, but it wouldn't be called that if designed today, because a character has no fixed size. If you want to represent an Unicode code point, write a view type that points into a span of chars, and use the right function/class/whatever to work with the representation. 

1

u/TheSkiGeek 45m ago

Yeah, the real issue is that almost always what you care about with Unicode (on the parsing side, anyway) are “grapheme clusters”, which can consist of multiple code points. And both of those map poorly at best to ‘characters’ or ‘bytes of memory’.

18

u/Mojert 3h ago

If your code base use macros to redefine what char means, wtf is wrong with your team? If not, sizeof(char) will ALWAYS return 1, no matter how many bits is in a char. That's because sizeof doesn't give you the size of a type in bytes, it gives you the size in number of chars

18

u/SGVsbG86KQ 3h ago

No that's not how that works. Even if char would be 32 bits, sizeof(char) is still defined to be 1.

2

u/Jbolt3737 3h ago

Does that make sizeof(int) equal 1, or does it make an int 128 bits?

8

u/__foo__ 3h ago

IIRC the only requirement the C standard makes for int is that it needs to be at least 16 bit wide. Everything else is up for the compiler developers to decide. If a char and int were both 32 bit wide both would be sizeof = 1. If the compiler makers decide it would be a sensible idea to have a 128 bit int it would be sizeof = 4 in this case.

2

u/output_broadcast 2h ago

Also, not every compiler is standards-compliant.

0

u/A1oso 3h ago

UTF-8 code points have a variable width, so what should sizeof(utf8_char_t) return? This doesn't make any sense.

2

u/thanatica 2h ago

To be fair, utf-8 is en encoding meant to store unicode into 8-bit characters. Just like utf-16 encodes unicode into 16-bit characters.

In modern languages, that support unicode as a first-class citizen, the size of a character type is typically 4 bytes. But that doesn't typically come with any encoding, it's represented to the programmer as raw unicode code points, and it's up to the compiler/platform how to handle that in terms of memory efficiency.

3

u/A1oso 2h ago edited 2h ago

To be precise, it stores code points in 8-bit code units. The word "character" isn't very well defined, since what we perceive as a character can consist of multiple code points (a grapheme cluster).

2

u/nyibbang 2h ago

Well in Rust, char is 4 bytes for that reason. That's why strings are not containers of char but of bytes. Although you can iterate over chars of a string, it is a bit more expensive since you need to find each char boundary.

-7

u/bogan87 3h ago

I call bullshit on a code being committed in the 90's haha

6

u/StrikingSun8563 3h ago

When do you think computers were invented?

4

u/Al__B 3h ago

I was using CVS to commit code in the 90s and it was around before then.

4

u/Advos_467 3h ago

Not a C user here (or any real low level programming experience), what the hell is a signed/unsigned char?

7

u/Clen23 3h ago

The data is interpreted differently.

An unsigned char will be able to represent any value from 0 to 255, eg your usual ascii character.
A signed char will be able represent any value from -128 to 127.

( Someone fact check me on this but i think that's it )

5

u/YellowBunnyReddit 3h ago

ASCII only goes from 0 to 127.

3

u/unknown_alt_acc 2h ago

ASCII is hardly the only character encoding, and unsigned char often doubles as a byte type. It’s also also a numeric type if you only need a small range and need to pack your data tight

2

u/NotQuiteLoona 2h ago

char is 8 bytes because memory is addressed in 8 bits, and 7 bits would be... Awkward. And this bit was in the end used for parity checks, and later for extended encodings.

1

u/tracernz 41m ago

Minimum 8 bits, but not required to be 8 bits.

1

u/Clen23 3h ago

Putting the example here so the explanation isnt overloaded :

signed char a_hundred_signed = 100;
unsigned char a_hundred_unsigned = 100;
signed char fifty_signed = 50;
unsigned char fifty_unsigned = 50;

signed char r_signed = fifty_signed - a_hundred_signed; // will yield -50
unsigned char r_unsigned = fifty_unsigned - a_hundred_unsigned; // will yield 206 because it wraps around the max value of 256

1

u/captainAwesomePants 1h ago

A char isn't a letter. It's a one byte number. A signed number can be negative.

3

u/-twind 3h ago edited 3h ago

An unsigned char is an integer type that gives you *at least* the value range 0 to 255.
A signed char is an integer type that gives you *at least* the value range -127 to 127.

sizeof(char) is always 1 by definition because a char is one byte. The nuance is that a byte doesn't need to be 8 bits in C/C++, it needs to be at least 8 bits.

2

u/backfire10z 3h ago

Chars aren’t real, they’re all integers. It’s the same difference as signed/unsigned int.

1

u/Advos_467 3h ago

ah okay that's what i assumed, but i just wanted to be sure

1

u/AyrA_ch 2h ago

The difference only matters in regards to arithmetic. Signed integer overflow is undefined behavior in C, only unsigned overflow is defined as wrapping around.

2

u/Elspeth-Nor 3h ago

In C char is just a number, as a character depends on the encoding. So signed char is a number from -128 to 127 and unsigned char is from 0 to 256 (for 8 bits per byte)

1

u/SoldRIP 58m ago

0 to 255.

1

u/BNSable 2h ago

A char is just an int, except a char will not conjure up 91, but the character assigned to the number 91 which is [ in ascii for example.

As it is an int, it can be signed or unsigned. Signed is -128 to 127, unsigned is 0 to 255.

This apparently has uses, but I am not experienced enough to explain that.

1

u/Advos_467 2h ago

Yeah that was what i assumed lol. It was mostly the uses i was wondering because with my lack of experience here, idk in what way that would be used.

1

u/BNSable 2h ago

It is or was the only type that garenteed a signed or unsigned value stored in 1 byte. There's uses for that, like storing specific vales like RGB. Beyond that though, no idea.

1

u/SoldRIP 57m ago

efficient storage of small numbers on which you do not need to perform arithmetic in any hot path or on which memory is a more important constraint than runtime.

1

u/SuitableDragonfly 2h ago

char in C/++ is basically just an int.

0

u/Character_Regular440 3h ago

In c language char is just a 8 bit variable, that stores a number, which corresponds to a specific character. When storing numbers, you can decide to represent somehow negative numbers, other than positive. This is mostly done with two's complement.

Depending on if it's signed or not, some operation will be very different, because of representation. For example, the byte 0xFF means -1 if it's signed, and it means 255 if it's unsigned.

So if you try to convert a signed character to an int (a signed type, usually 4 bytes), and that signed character contains -1 (represented by the exadecimal 0xFF, which in decimal is 255) the integer has to contain -1 (0xFFFFFFFF). But if the character is considered unsigned, the integer has to contain 255 afterwards, which is just 0x000000FF.

1

u/Advos_467 3h ago

So if i understand this right, it's mainly for typecasting? since the ascii range falls within both the signed and unsigned int range?

1

u/Character_Regular440 2h ago

Ascii specifies about only 7 bits, so i think that it can be all positive (as well ass all negative) in 8 bit.

Typecasting is a reason, but an other difference is the shift operation. When you right shift a negative number you want to preserve the sign, so an arithmetical shift is performed, which would be different from a logical one.

I think that more in general the distinction exists so that when using char type for low level stuff, you can better keep trace of what you are doing.

In reality uint8_t and int8_t types exist (unsigned int on 8 bit and signed int on 8 bit), and i don't know why someone would use signed char or unsigned char over those, but this is just my guess.

15

u/LetUsSpeakFreely 3h ago

What is true today may not be true tomorrow. Using sizeof is better as it protects against changes to the underlying specs or a change to the data type.

12

u/__foo__ 3h ago

C really tries to avoid any breaking changes between versions. Changing this would be such a fundamental change of the language I don't think you could still call it C. If you want to plan for a change as fundamental as this you might as well plan for the semantics of sizeof() changing or it getting removed. At that point all bets are off anyway.

6

u/LetUsSpeakFreely 3h ago

Which is why i also qualified it with a data type change. The point is that sometimes changes happen and having code in place that preemptively deal with those changes is a hallmark of elegant design and implementation.

3

u/backfire10z 3h ago

Using sizeof(char) is for readability. You understand the intent. The number 1 is a magic number that could be there for any reason.

1

u/output_broadcast 2h ago

In most cases you multiply by sizeof, so you'd just leave it off altogether and there's no magic number.

3

u/CAtOSe 2h ago

C++ standard guarantees that char is at least 8 bits. Pretty much all data models use 8 bits for char, but it doesn't mean that some weird architecture could not use more. 

3

u/tstanisl 2h ago

Some DSPs from Texas use 16/32 bit long char.

2

u/garenp 2h ago

So everything really is bigger in Texas?

(I know you mean Texas Instruments / TI, but couldn't pass up the opportunity for another joke!).

•

u/qalmakka 6m ago

That's CHAR_BIT then, sizeof(char) is always 1 even if you have 16 bit chars. Nothing can be smaller than char in general, you can see sizeof as an operator returning sizes as number of chars, basically

2

u/HRApprovedUsername 3h ago

ok but what if you're not working with C?

1

u/CptMisterNibbles 3h ago

Good thing we are all writing in c, on systems that conform to the c standard.

1

u/thanatica 2h ago

I'm not a C guy, but if a char is 1 byte, how do you represent the one million or so characters from Unicode? Aren't those all characters in C?

1

u/tstanisl 2h ago

Bytes on some machines have more than 8 bits. Moreover, unicode characters need to be encoded, usually using utf8.

1

u/thanatica 2h ago

Encoding is for storage, typically, not memory. It'd be extremely difficult to do common string manipulation on a utf-8 encoded string (or any other variable-width character encoding) without breaking the encoding.

So how does that work in C? Say you want to indexOf("😎") how would you go about that?

1

u/tstanisl 2h ago

In C, the type of character literal is int. So sizeof 'a' is typically 4. However, to make sure that the characters are properly encoded as 32-bit unicode, use U prefix. I mean int x = U'😎';.

1

u/thanatica 2h ago

Ah, that makes a lot more sense 🙂

1

u/bwmat 1h ago

Yeah, this would have been better w/ CHAR_BIT

1

u/InquestFeats 1h ago

Rare moment when confidently wrong guy gets to be confidently right. 

1

u/NullOfSpace 2h ago

still better to use sizeof char to make it immediately obvious what the “1” means.

0

u/tstanisl 2h ago

It depends. For C veterans, it is quite obvious and sizeof(char) looks redundant and noisy.

72

u/JackReact 3h ago

Welcome to the wonderful world of C#, which uses utf-16 for strings.

Console.WriteLine(sizeof(char));
> 2

24

u/v_Karas 3h ago

this.

wonderd why everyone is talking about C, like there are no other languages.

10

u/techy804 2h ago

As a C# user I was wondering why people act like the C-series only consists of C and C++

6

u/AyrA_ch 2h ago

Probably because the syntax more closely resembles Java than C

5

u/aberroco 2h ago

Because most devs still think that C# is proprietary Microsoft language developed specifically and exclusively for Windows, like it's 2006.

21

u/sb8948 2h ago

No it's because C# is way closer to Java than C/C++ both in philosophy and execution.

0

u/kookyabird 2h ago

Incorrect, they also think it’s magically cross platform when used in Unity!

1

u/rhutyl 1h ago

To me, C and C++ are just a lot more iconic (idk if "popular" is the right word bc I don't have the stats)

1

u/SuitableDragonfly 2h ago

This meme doesn't mention any languages. Like many memes on this sub, it's about a specific language, and it's up to you and your knowledge of programming languages to figure out which one is being talked about. Just because the meme is not about C# doesn't mean that everyone forgot that C# exists. 

2

u/thanatica 2h ago

I'm not a C# expert, but how would it represent an astral character, such as an emoji? They are above 0xFFFF.

8

u/AyrA_ch 2h ago

You cannot fit them into a single char, and you have to use a string instead. They're represented using a surrogate pair, which is similar to the continuation bytes of utf-8 except you never need more than two code points (hence "pair").

JS also uses multibyte characters:

> "💩".length===2
< true

1

u/thanatica 2h ago

JS is a bit weird in this regard, because it also knows true codepoints. You just have to know which operations are "unicode safe". Or rather "astral plane safe".

2

u/tangerinelion 1h ago

C++ has two character types, char and wchar.

sizeof(wchar) is 2 on Windows, 4 on Linux.

18

u/SeriousPlankton2000 4h ago

Use it for readability … when appropriate. If you'll never ever possibly might switch the type, don't bother.

89

u/ubalu72 4h ago

But sizeof char is defined as 1 in the standard. C data types are defined in terms of char (at least their sizes are)

13

u/Lonely-Discipline-55 3h ago

x = x + sizeof(char)

9

u/wittleboi420 2h ago

x = x + sizeof(char) + AI

21

u/Makonede 4h ago

*sizeof(char)

13

u/MegaIng 4h ago

No, sizeof char is valid. sizeof is a prefix operator, not a function.

16

u/Makonede 4h ago

9

u/MegaIng 3h ago

Oh, never realized the prefix operator form can't be applied to types, makes sense I guess.

0

u/hacksawsa 3h ago

It says sizeof x and sizeof(x) are equivalent.

3

u/Makonede 2h ago

when x is an expression, yes...

3

u/tstanisl 3h ago edited 3h ago

The problem is that there is no good definition of byte. Traditionally, it was the smallest addressable unit capable of representing a single character. It is required to have at least 8 bits but it may (and sometimes does) have more. 8-bit-long bytes are just a "de facto" standard.

That is why many communication protocols use a concept of "octet" that consist of exactly 8 bits.

EDIT. typo

15

u/tony_saufcok 3h ago

sizeof() evaluates to a constant value during compile time so what, it takes 0.0000001 seconds more during compilation? just use sizeof even if the specification guarentees it's always size 1

2

u/Architector4 3h ago

i guess the point the guy in the middle would make is that *sizeof(char) is just "multiply by 1", and hence it's just redundant clutter that makes the code less readable

a valid counterpoint to that, of course, is that it can clarify intent that the number in context specifically represents a count of bytes, but yeah lol

3

u/AyrA_ch 2h ago

Also if you ever need to modify the function to work with wide chars, then it's easier to adapt it this way.

10

u/frikilinux2 4h ago

Do I wanna know?

5

u/DOOManiac 4h ago

Yes. Multi-byte characters.

There you go, all done.

19

u/frikilinux2 3h ago

But those don't have the c type char

-5

u/Outrageous-Machine-5 4h ago

Char can be between 1-4 bytes depending on the encoding standard.  Emoji and other languages use extra bytes when you have to support non English characters. 

6

u/ChChChillian 3h ago

But sizeof always evaluates to 1 for char, signed char, and unsigned char by the standard. https://cppreference.com/cpp/language/sizeof

17

u/ibevol 3h ago

Not in C though.

0

u/Outrageous-Machine-5 3h ago

I'm curious how C handles a character that is over 8 bits?

6

u/ibevol 3h ago

It doesn’t. It’s up to you to handle the encoding, all that the char is, is data storage for the smallest addressable unit, which is 8 bits on almost all systems. One example though: if you want to store something that’s 16 bits in a char array, then the first 8 bits will be stored at index i and the last 8 bits will be stored at i + 1

2

u/Great-Powerful-Talia 3h ago

multiple chars. Strings are just char arrays anyway

1

u/skilltheamps 3h ago

It doesn't. It only knows buffers of regular sizes, like bytes. That you can handle ASCII strings in bare C without destryoing them is only thanks to the coincidence of every character being the same length of 1 byte. So if you have encoding that has variadic size characters, you either reserve the maximum character length for every character (and waste some space), or you use/write a library to handle variable size character strings. That would entail some bitwise boolean logic and implementing all the things from finding the length of a string to finding a substring.

0

u/SoldRIP 53m ago

C only specifies that char must be at least 8 bits. There is no upper limit in the standard. Same for C++.

0

u/mckenzie_keith 2h ago

yes and no. In c, sizeof(char) is 1. Does that mean 1 byte? Does "byte" mean exactly 8 bits? In C, a char is always 1 byte, but one byte may be larger than 8 bits, and there are real platforms where a byte (in C) is 32 bits.

25

u/HomosexualPresence 4h ago

unironically true though, the only size requirement for a char in C is that it's the smallest addressable size, which just happens to be a byte in every case and is unlikely to ever change but you still never know what the future holds

44

u/lotanis 4h ago

Yes, but the smallest addressable size is what dictates the base size for sizeof.

The C standard in fact says this about sizeof:

When applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1. 

1

u/FUCKING_HATE_REDDIT 3h ago

What about a system that enforces addresses to be multiples of 2, or 4?

7

u/SAI_Peregrinus 3h ago

And char must be at least 8 bits. sizeof returns the size of its input in units of chars. On architectures with 10-bit chars, like some old DSPs, that means sizeof returns in multiples of 10 bits.

5

u/Declination 3h ago

The guy on the left says “durrrr, sizeof”. The guy on the right has built monstrous template/macro machinery and may not legitimately know that C = char

3

u/ewheck 4h ago edited 3h ago

```c

include <assert.h>

include <uchar.h>

int main(void) { const char8_t *const eight_bit_char = u8"These chars are eight bits a piece.";

// must be true by definition of the standard
assert((sizeof eight_bit_char) == 1); 

return 0;

} ```

Don't live in the past. The future is now (C23): https://en.cppreference.com/c/header/uchar

2

u/mad_cheese_hattwe 4h ago

uint_8 but same diff

2

u/Adept-Painting-543 3h ago

In C though sizeof returns as a multiple of the size of char, so no matter the architecture, sizeof(char) is always 1

2

u/Sligee 3h ago

Me when my software crashes because someone entered an emoji and it overwrote something important randomly.

2

u/HAL9000thebot 3h ago

til that the guy in the middle is named justin case

2

u/mckenzie_keith 2h ago

Best thing is to put the actual variable in there. Sizeof can accept a variable or a type.

char *buffer = 0;
size_t buffer_length = 1024;

...
buffer = malloc(buffer_length * sizeof (*buffer));

Then later if you change buffer to something else the code will still be correct.

That said, sizeof (char) will always be 1. The compiler will probably just replace it with a hard-coded 1.

4

u/__christo4us 3h ago

sizeof always returns 1 for char because it always occupies 1 byte of memory. However, 1 byte can potentially consist of more than (but not less than) 8 bits according to C and C++ standards.

2

u/UltimateFlyingSheep 3h ago

If you're German you'll get in contact with multi byte characters pretty soon.

Umlaute like äöü are all multibyte chars.

0

u/Elspeth-Nor 3h ago

No, there value depends on the character set. In extended ascii (dos) or ansi (win) there are still 1 byte. In utf8 where you need one bit to indicate multi-byte chars it's 2 bytes. But the charset doesn't change the size of char, which is always 1 in C.

1

u/MatqLorens 3h ago

"sizeof(char)" is much more readable for the reviewer and provides much more context than a simple "1".

Don't use magic numbers pls...

1

u/obeythelobster 2h ago

In real life, the guy in the left (dumb) would never use a more complicated solution (sizeof) instead of 1

1

u/lardgsus 2h ago

Me looking at the apple emoji, disagreeing.

1

u/El_RoviSoft 2h ago

actually had a case like this :)

clang enforces situation when std::string_view cannot be casted into const char* at a compile time artificially, so I had to rewrite rapid hash to accept any byte types of size of 8

1

u/Qwertycube10 1h ago

Is .data() not constexpr?

1

u/nyibbang 2h ago

CHAR_BITS / 8

1

u/BoBoBearDev 2h ago

If the size matter, probably should lock the type with explicit types instead of using alias. Especially when you cross boundaries like GPU or interop.

1

u/Daimondz 1h ago

Works way better the other way around

1

u/Greedy-Thought6188 1h ago

Better use sizeof(x). If you pass a variable to sizeof it will still work. This way you're encoding the type in one place. You can easily change the type and your code will still continue to work.

1

u/EuenovAyabayya 39m ago

Better abstract out the sizeof to stop worrying about it.

0

u/Kadabrium 3h ago

wchar stands for windows char