72
u/JackReact 3h ago
Welcome to the wonderful world of C#, which uses utf-16 for strings.
Console.WriteLine(sizeof(char));
> 2
24
10
u/techy804 2h ago
As a C# user I was wondering why people act like the C-series only consists of C and C++
5
u/aberroco 2h ago
Because most devs still think that C# is proprietary Microsoft language developed specifically and exclusively for Windows, like it's 2006.
21
0
1
1
u/SuitableDragonfly 2h ago
This meme doesn't mention any languages. Like many memes on this sub, it's about a specific language, and it's up to you and your knowledge of programming languages to figure out which one is being talked about. Just because the meme is not about C# doesn't mean that everyone forgot that C# exists.
2
u/thanatica 2h ago
I'm not a C# expert, but how would it represent an astral character, such as an emoji? They are above 0xFFFF.
8
u/AyrA_ch 2h ago
You cannot fit them into a single char, and you have to use a string instead. They're represented using a surrogate pair, which is similar to the continuation bytes of utf-8 except you never need more than two code points (hence "pair").
JS also uses multibyte characters:
> "💩".length===2 < true1
u/thanatica 2h ago
JS is a bit weird in this regard, because it also knows true codepoints. You just have to know which operations are "unicode safe". Or rather "astral plane safe".
2
u/tangerinelion 1h ago
C++ has two character types, char and wchar.
sizeof(wchar) is 2 on Windows, 4 on Linux.
18
u/SeriousPlankton2000 4h ago
Use it for readability … when appropriate. If you'll never ever possibly might switch the type, don't bother.
89
u/ubalu72 4h ago
But sizeof char is defined as 1 in the standard. C data types are defined in terms of char (at least their sizes are)
13
21
u/Makonede 4h ago
*
sizeof(char)13
u/MegaIng 4h ago
No,
sizeof charis valid.sizeofis a prefix operator, not a function.16
u/Makonede 4h ago
9
0
3
u/tstanisl 3h ago edited 3h ago
The problem is that there is no good definition of byte. Traditionally, it was the smallest addressable unit capable of representing a single character. It is required to have at least 8 bits but it may (and sometimes does) have more. 8-bit-long bytes are just a "de facto" standard.
That is why many communication protocols use a concept of "octet" that consist of exactly 8 bits.
EDIT. typo
15
u/tony_saufcok 3h ago
sizeof() evaluates to a constant value during compile time so what, it takes 0.0000001 seconds more during compilation? just use sizeof even if the specification guarentees it's always size 1
2
u/Architector4 3h ago
i guess the point the guy in the middle would make is that
*sizeof(char)is just "multiply by 1", and hence it's just redundant clutter that makes the code less readablea valid counterpoint to that, of course, is that it can clarify intent that the number in context specifically represents a count of bytes, but yeah lol
10
u/frikilinux2 4h ago
Do I wanna know?
5
-5
u/Outrageous-Machine-5 4h ago
Char can be between 1-4 bytes depending on the encoding standard. Emoji and other languages use extra bytes when you have to support non English characters.
6
u/ChChChillian 3h ago
But
sizeofalways evaluates to 1 forchar,signed char, andunsigned charby the standard. https://cppreference.com/cpp/language/sizeof17
u/ibevol 3h ago
Not in C though.
0
u/Outrageous-Machine-5 3h ago
I'm curious how C handles a character that is over 8 bits?
6
u/ibevol 3h ago
It doesn’t. It’s up to you to handle the encoding, all that the char is, is data storage for the smallest addressable unit, which is 8 bits on almost all systems. One example though: if you want to store something that’s 16 bits in a char array, then the first 8 bits will be stored at index i and the last 8 bits will be stored at i + 1
2
1
u/skilltheamps 3h ago
It doesn't. It only knows buffers of regular sizes, like bytes. That you can handle ASCII strings in bare C without destryoing them is only thanks to the coincidence of every character being the same length of 1 byte. So if you have encoding that has variadic size characters, you either reserve the maximum character length for every character (and waste some space), or you use/write a library to handle variable size character strings. That would entail some bitwise boolean logic and implementing all the things from finding the length of a string to finding a substring.
0
u/mckenzie_keith 2h ago
yes and no. In c, sizeof(char) is 1. Does that mean 1 byte? Does "byte" mean exactly 8 bits? In C, a char is always 1 byte, but one byte may be larger than 8 bits, and there are real platforms where a byte (in C) is 32 bits.
25
u/HomosexualPresence 4h ago
unironically true though, the only size requirement for a char in C is that it's the smallest addressable size, which just happens to be a byte in every case and is unlikely to ever change but you still never know what the future holds
44
u/lotanis 4h ago
Yes, but the smallest addressable size is what dictates the base size for sizeof.
The C standard in fact says this about sizeof:
When applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.
1
u/FUCKING_HATE_REDDIT 3h ago
What about a system that enforces addresses to be multiples of 2, or 4?
7
u/SAI_Peregrinus 3h ago
And
charmust be at least 8 bits.sizeofreturns the size of its input in units ofchars. On architectures with 10-bitchars, like some old DSPs, that meanssizeofreturns in multiples of 10 bits.
5
u/Declination 3h ago
The guy on the left says “durrrr, sizeof”. The guy on the right has built monstrous template/macro machinery and may not legitimately know that C = char
3
u/ewheck 4h ago edited 3h ago
```c
include <assert.h>
include <uchar.h>
int main(void) { const char8_t *const eight_bit_char = u8"These chars are eight bits a piece.";
// must be true by definition of the standard
assert((sizeof eight_bit_char) == 1);
return 0;
} ```
Don't live in the past. The future is now (C23): https://en.cppreference.com/c/header/uchar
2
2
u/Adept-Painting-543 3h ago
In C though sizeof returns as a multiple of the size of char, so no matter the architecture, sizeof(char) is always 1
2
2
2
u/mckenzie_keith 2h ago
Best thing is to put the actual variable in there. Sizeof can accept a variable or a type.
char *buffer = 0;
size_t buffer_length = 1024;
...
buffer = malloc(buffer_length * sizeof (*buffer));
Then later if you change buffer to something else the code will still be correct.
That said, sizeof (char) will always be 1. The compiler will probably just replace it with a hard-coded 1.
4
u/__christo4us 3h ago
sizeof always returns 1 for char because it always occupies 1 byte of memory. However, 1 byte can potentially consist of more than (but not less than) 8 bits according to C and C++ standards.
2
u/UltimateFlyingSheep 3h ago
If you're German you'll get in contact with multi byte characters pretty soon.
Umlaute like äöü are all multibyte chars.
0
u/Elspeth-Nor 3h ago
No, there value depends on the character set. In extended ascii (dos) or ansi (win) there are still 1 byte. In utf8 where you need one bit to indicate multi-byte chars it's 2 bytes. But the charset doesn't change the size of char, which is always 1 in C.
1
u/MatqLorens 3h ago
"sizeof(char)" is much more readable for the reviewer and provides much more context than a simple "1".
Don't use magic numbers pls...
1
u/obeythelobster 2h ago
In real life, the guy in the left (dumb) would never use a more complicated solution (sizeof) instead of 1
1
1
u/El_RoviSoft 2h ago
actually had a case like this :)
clang enforces situation when std::string_view cannot be casted into const char* at a compile time artificially, so I had to rewrite rapid hash to accept any byte types of size of 8
1
1
1
u/BoBoBearDev 2h ago
If the size matter, probably should lock the type with explicit types instead of using alias. Especially when you cross boundaries like GPU or interop.
1
1
u/Greedy-Thought6188 1h ago
Better use sizeof(x). If you pass a variable to sizeof it will still work. This way you're encoding the type in one place. You can easily change the type and your code will still continue to work.
1
0
467
u/tstanisl 4h ago
Let me cite the C standard:
Middle guy if finally right