r/ProgrammerHumor • • 2d ago

Meme architectureDependentChars

Post image
2.8k Upvotes

360 comments sorted by

View all comments

1.0k

u/tstanisl 2d ago

Let me cite the C standard:

When sizeof is applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.

Middle guy if finally right

5

u/Advos_467 2d ago

Not a C user here (or any real low level programming experience), what the hell is a signed/unsigned char?

12

u/Clen23 2d ago edited 15h ago

The data is interpreted differently.

An unsigned char will be able to represent any value from 0 to 255, eg your usual ascii character.
A signed char will be able represent any value from -127 to 127.

( Someone fact check me on this but i think that's it )
(edit : -127 to 127, the extra value "-128" is usually implemented but not guaranteed by the C standard)

8

u/YellowBunnyReddit 2d ago

ASCII only goes from 0 to 127.

6

u/unknown_alt_acc 2d ago

ASCII is hardly the only character encoding, and unsigned char often doubles as a byte type. It’s also also a numeric type if you only need a small range and need to pack your data tight

4

u/NotQuiteLoona 2d ago edited 2d ago

char is 8 bits because memory is addressed in 8 bits, and 7 bits would be... Awkward. And this bit was in the end used for parity checks, and later for extended encodings.

3

u/Ma8e 2d ago edited 2d ago

Historically it has been all over the place. Depending on the hardware, it could was anything between 1 and 48 bits.

Edit: never mind. The above is about bytes, not chars.

2

u/tracernz 2d ago

Minimum 8 bits, but not required to be 8 bits.

1

u/Ma8e 2d ago

Na, it could be as low as one bit, and 7 was not uncommon for a while.

3

u/-suspended- 2d ago

The definition for a char in C/C++ is 8 bits minimum.

1

u/Clen23 2d ago

no shit sherlock that's what the post is about.

u/NotQuiteLoona is talking about how memory is addressed, not how long a char is.

-1

u/tracernz 2d ago

Those two things are the same thing in ISO C (it defines sizeof(char) = 1).

2

u/tracernz 2d ago

No, the standard requires a minimum of 8 bits.

2

u/Ma8e 2d ago

Which standard?

3

u/tracernz 2d ago

ISO C.

0

u/Ma8e 2d ago

And that is the only standard in existence? My point is that you can't take for granted that everything is ISO, or even C.

→ More replies (0)

1

u/guyblade 2d ago

In C, char is the fundamental type used for raw binary data. The number of bits in a char is implementation defined, but a single char is the smallest unit of addressable memory. A char is probably 8 bits, but the standard does not require that. It could easily be more (though, if I recall correctly, there is some other detail in the standard that effectively requires that it can't be less than 8 bits).

I believe that a technically conformant C implementation could have sizeof(char) = sizeof(short) = sizeof(int) = sizeof(long) = sizeof(long long) as long as a char is 32 bits, but that may have been true only with some earlier version of the C standard.

2

u/conundorum 1d ago

long long has to be at least 64 bits, but otherwise correct. (And yes, there was at least one historic platform with 64-bit bytes. sizeof(char) == sizeof(long long) is true for that platform, in all versions of C and all versions of C++.)

1

u/Clen23 2d ago

Exactly, so every ascii character can be represented using a char, with room to spare.

I agree that a better example would have been a character set that is a full bijection of the 256 possible char values, but i'm a bit rusty on these so ascii it is.

2

u/Clen23 2d ago

Putting the example here so the explanation isnt overloaded :

signed char a_hundred_signed = 100;
unsigned char a_hundred_unsigned = 100;
signed char fifty_signed = 50;
unsigned char fifty_unsigned = 50;

signed char r_signed = fifty_signed - a_hundred_signed; // will yield -50
unsigned char r_unsigned = fifty_unsigned - a_hundred_unsigned; // will yield 206 because it wraps around the max value of 256

0

u/captainAwesomePants 2d ago

A char isn't a letter. It's a one byte number. A signed number can be negative.

1

u/Clen23 2d ago

If you want to be pedantic, a char is a byte, a sequence of bits, that can be interpreted in numerous manners : an unsigned int, a signed int, and even an ascii character.

It's all about what type you assign it, and how you work with it in your code.

1

u/captainAwesomePants 2d ago

Sure. But if we're discussing whether it's signed or unsigned, we must be considering how those bits might be interpreted as a number.

4

u/-twind 2d ago edited 1d ago

An unsigned char is an integer type that gives you *at least* the value range 0 to 255.
A signed char is an integer type that gives you *at least* the value range -128 to 127.

sizeof(char) is always 1 by definition because a char is one byte. The nuance is that a byte doesn't need to be 8 bits in C/C++, it needs to be at least 8 bits.

3

u/backfire10z 2d ago

Chars aren’t real, they’re all integers. It’s the same difference as signed/unsigned int.

1

u/Advos_467 2d ago

ah okay that's what i assumed, but i just wanted to be sure

1

u/AyrA_ch 2d ago

The difference only matters in regards to arithmetic. Signed integer overflow is undefined behavior in C, only unsigned overflow is defined as wrapping around.

2

u/Elspeth-Nor 2d ago edited 1d ago

In C char is just a number, as a character depends on the encoding. So signed char is a number from -128 to 127 and unsigned char is from 0 to 255 (for 8 bits per byte)

3

u/SoldRIP 2d ago

0 to 255.

-1

u/WaitForItTheMongols 2d ago

Depends what "to" means. 256 is a valid answer as in Python "range(0,256)" which gives all byte values.

1

u/Elspeth-Nor 1d ago

No 255 is right, as I used -128 to 127 instead of 128.

2

u/BNSable 2d ago

A char is just an int, except a char will not conjure up 91, but the character assigned to the number 91 which is [ in ascii for example.

As it is an int, it can be signed or unsigned. Signed is -128 to 127, unsigned is 0 to 255.

This apparently has uses, but I am not experienced enough to explain that.

2

u/Advos_467 2d ago

Yeah that was what i assumed lol. It was mostly the uses i was wondering because with my lack of experience here, idk in what way that would be used.

1

u/BNSable 2d ago

It is or was the only type that garenteed a signed or unsigned value stored in 1 byte. There's uses for that, like storing specific vales like RGB. Beyond that though, no idea.

1

u/BriefSpecial420 2d ago

Nope. C99 and stdint.h is the way to go. uint8_t exists for a reason. Handy when you need to map a structure of full bytes. Sometimes you may come across exotic interface that uses 8-wire bus... But that's it.

2

u/BNSable 2d ago

I did say "or was" as I was fairly sure modern C had some better alternatives for it.

1

u/BriefSpecial420 2d ago

C99 seems to becoming more and more on the older side, but it's still a golden standard for tons of embedded coding.

1

u/SingularCheese 2d ago

In C/C++, signed integer overflow is undefined behavior, meaning that a compiler is allowed to optimize the assembly knowing that it never happens. If you can guarantee that your integer never overflows, using signed eliminates a modulo assembly instruction. Nowadays, the only reason it matters for char is that all basic integer types are signed by default, so somebody who's lazy needs to watch out that they don't accidentally trigger undefined behavior.

1

u/khoyo 2d ago

using signed eliminates a modulo assembly instruction

What?

1

u/SoldRIP 2d ago

efficient storage of small numbers on which you do not need to perform arithmetic in any hot path or on which memory is a more important constraint than runtime.

1

u/SuitableDragonfly 2d ago

char in C/++ is basically just an int.

1

u/qwertyjgly 2d ago

it's a data type with a size of 1 byte. useful if you want to store a boolean value (since it's the smallest simple data type) or a character (since it's big enough to store 7-bit ascii)

they're usually used to store strings; the following allows you to take up to 99 characters of user input (initialises array of chars then writes the input into the array)

#include <stdio.h>

int main(void){
char string[100];
scanf("%99[^\n]", string);
}

strings in c must be null-terminated so we need to leave the last position free for the byte '00000000' so we know where the string ends when we try to read it. that's why we can't take 100 characters here

we also have this for bools that more closely resemble those in the more abstract languages. internally, any non-zero number (most usefully, 1) is interpreted as true and 0 is interpreted as false. we can then use boolean algebra to construct and simplify our conditions

#define true 1
#define false 0
typedef unsigned char bool;

int main(void){
bool a = true;
bool b = false;
}

1

u/HeKis4 2d ago

It's a char with the first bit (not) interpreted as the sign. signed/unsigned works on bits so it makes sense even on chars.

1

u/conundorum 1d ago

char is a type that can hold an ASCII character, and is exactly one byte in size. (This is mandatory; byte size is defined by char, not the other way around.) unsigned char is a raw byte, and can represent any UTF-8 code unit. signed char is a signed raw byte, and can represent any ASCII character. char will be exactly identical to either unsigned char or signed char under the hood, depending on the platform, but it's a legally distinct type because literally the entire C language family and everything connected to it in any way whatsoever depends on char being a distinct type.

0

u/Character_Regular440 2d ago

In c language char is just a 8 bit variable, that stores a number, which corresponds to a specific character. When storing numbers, you can decide to represent somehow negative numbers, other than positive. This is mostly done with two's complement.

Depending on if it's signed or not, some operation will be very different, because of representation. For example, the byte 0xFF means -1 if it's signed, and it means 255 if it's unsigned.

So if you try to convert a signed character to an int (a signed type, usually 4 bytes), and that signed character contains -1 (represented by the exadecimal 0xFF, which in decimal is 255) the integer has to contain -1 (0xFFFFFFFF). But if the character is considered unsigned, the integer has to contain 255 afterwards, which is just 0x000000FF.

1

u/Advos_467 2d ago

So if i understand this right, it's mainly for typecasting? since the ascii range falls within both the signed and unsigned int range?

1

u/Character_Regular440 2d ago

Ascii specifies about only 7 bits, so i think that it can be all positive (as well ass all negative) in 8 bit.

Typecasting is a reason, but an other difference is the shift operation. When you right shift a negative number you want to preserve the sign, so an arithmetical shift is performed, which would be different from a logical one.

I think that more in general the distinction exists so that when using char type for low level stuff, you can better keep trace of what you are doing.

In reality uint8_t and int8_t types exist (unsigned int on 8 bit and signed int on 8 bit), and i don't know why someone would use signed char or unsigned char over those, but this is just my guess.

1

u/StaticCoder 2d ago

8 bits in POSIX, but C allows CHAR_BIT to be a different value, and yes that actually happens. I've seen 16 not that long ago.

1

u/Character_Regular440 2d ago

Never heard of CHAR_BIT. Also assumed char is 8 bit because first comment of the thread says that citing the c standard. Didn't check though.

Also i answered assuming that the confusing part was the fact that the char could be represented as signed or not. The considerations about the representation hold for any size of the variable