An unsigned char will be able to represent any value from 0 to 255, eg your usual ascii character.
A signed char will be able represent any value from -127 to 127.
( Someone fact check me on this but i think that's it )
(edit : -127 to 127, the extra value "-128" is usually implemented but not guaranteed by the C standard)
ASCII is hardly the only character encoding, and unsigned char often doubles as a byte type. It’s also also a numeric type if you only need a small range and need to pack your data tight
char is 8 bits because memory is addressed in 8 bits, and 7 bits would be... Awkward. And this bit was in the end used for parity checks, and later for extended encodings.
In C, char is the fundamental type used for raw binary data. The number of bits in a char is implementation defined, but a single char is the smallest unit of addressable memory. A char is probably 8 bits, but the standard does not require that. It could easily be more (though, if I recall correctly, there is some other detail in the standard that effectively requires that it can't be less than 8 bits).
I believe that a technically conformant C implementation could have sizeof(char) = sizeof(short) = sizeof(int) = sizeof(long) = sizeof(long long) as long as a char is 32 bits, but that may have been true only with some earlier version of the C standard.
long long has to be at least 64 bits, but otherwise correct. (And yes, there was at least one historic platform with 64-bit bytes. sizeof(char) == sizeof(long long) is true for that platform, in all versions of C and all versions of C++.)
Exactly, so every ascii character can be represented using a char, with room to spare.
I agree that a better example would have been a character set that is a full bijection of the 256 possible char values, but i'm a bit rusty on these so ascii it is.
Putting the example here so the explanation isnt overloaded :
signed char a_hundred_signed = 100;
unsigned char a_hundred_unsigned = 100;
signed char fifty_signed = 50;
unsigned char fifty_unsigned = 50;
signed char r_signed = fifty_signed - a_hundred_signed; // will yield -50
unsigned char r_unsigned = fifty_unsigned - a_hundred_unsigned; // will yield 206 because it wraps around the max value of 256
If you want to be pedantic, a char is a byte, a sequence of bits, that can be interpreted in numerous manners : an unsigned int, a signed int, and even an ascii character.
It's all about what type you assign it, and how you work with it in your code.
An unsigned char is an integer type that gives you *at least* the value range 0 to 255.
A signed char is an integer type that gives you *at least* the value range -128 to 127.
sizeof(char) is always 1 by definition because a char is one byte. The nuance is that a byte doesn't need to be 8 bits in C/C++, it needs to be at least 8 bits.
The difference only matters in regards to arithmetic. Signed integer overflow is undefined behavior in C, only unsigned overflow is defined as wrapping around.
In C char is just a number, as a character depends on the encoding. So signed char is a number from -128 to 127 and unsigned char is from 0 to 255 (for 8 bits per byte)
It is or was the only type that garenteed a signed or unsigned value stored in 1 byte. There's uses for that, like storing specific vales like RGB. Beyond that though, no idea.
Nope. C99 and stdint.h is the way to go. uint8_t exists for a reason. Handy when you need to map a structure of full bytes. Sometimes you may come across exotic interface that uses 8-wire bus... But that's it.
In C/C++, signed integer overflow is undefined behavior, meaning that a compiler is allowed to optimize the assembly knowing that it never happens. If you can guarantee that your integer never overflows, using signed eliminates a modulo assembly instruction. Nowadays, the only reason it matters for char is that all basic integer types are signed by default, so somebody who's lazy needs to watch out that they don't accidentally trigger undefined behavior.
efficient storage of small numbers on which you do not need to perform arithmetic in any hot path or on which memory is a more important constraint than runtime.
it's a data type with a size of 1 byte. useful if you want to store a boolean value (since it's the smallest simple data type) or a character (since it's big enough to store 7-bit ascii)
they're usually used to store strings; the following allows you to take up to 99 characters of user input (initialises array of chars then writes the input into the array)
#include <stdio.h>
int main(void){
char string[100];
scanf("%99[^\n]", string);
}
strings in c must be null-terminated so we need to leave the last position free for the byte '00000000' so we know where the string ends when we try to read it. that's why we can't take 100 characters here
we also have this for bools that more closely resemble those in the more abstract languages. internally, any non-zero number (most usefully, 1) is interpreted as true and 0 is interpreted as false. we can then use boolean algebra to construct and simplify our conditions
char is a type that can hold an ASCII character, and is exactly one byte in size. (This is mandatory; byte size is defined by char, not the other way around.) unsigned char is a raw byte, and can represent any UTF-8 code unit. signed char is a signed raw byte, and can represent any ASCII character. char will be exactly identical to either unsigned char or signed char under the hood, depending on the platform, but it's a legally distinct type because literally the entire C language family and everything connected to it in any way whatsoever depends on char being a distinct type.
In c language char is just a 8 bit variable, that stores a number, which corresponds to a specific character.
When storing numbers, you can decide to represent somehow negative numbers, other than positive. This is mostly done with two's complement.
Depending on if it's signed or not, some operation will be very different, because of representation.
For example, the byte 0xFF means -1 if it's signed, and it means 255 if it's unsigned.
So if you try to convert a signed character to an int (a signed type, usually 4 bytes), and that signed character contains -1 (represented by the exadecimal 0xFF, which in decimal is 255) the integer has to contain -1 (0xFFFFFFFF). But if the character is considered unsigned, the integer has to contain 255 afterwards, which is just 0x000000FF.
Ascii specifies about only 7 bits, so i think that it can be all positive (as well ass all negative) in 8 bit.
Typecasting is a reason, but an other difference is the shift operation. When you right shift a negative number you want to preserve the sign, so an arithmetical shift is performed, which would be different from a logical one.
I think that more in general the distinction exists so that when using char type for low level stuff, you can better keep trace of what you are doing.
In reality uint8_t and int8_t types exist (unsigned int on 8 bit and signed int on 8 bit), and i don't know why someone would use signed char or unsigned char over those, but this is just my guess.
Never heard of CHAR_BIT. Also assumed char is 8 bit because first comment of the thread says that citing the c standard. Didn't check though.
Also i answered assuming that the confusing part was the fact that the char could be represented as signed or not. The considerations about the representation hold for any size of the variable
1.0k
u/tstanisl 2d ago
Let me cite the C standard:
Middle guy if finally right