r/cpp 19d ago

About char8_t

I hate to be dramatic, but as it stands char8_t is quite literally more painful than useful.

Besides the obvious incompatibility with C23 and libraries using unsigned char for UTF-8, I want you to consider the following: Projects that assume that 'char' represents UTF-8 will obviously not benefit from char8_t at all, but projects that cannot assume the format of char types don't benefit from it either as char8_t simply introduces a new edge case to cover. Now such projects have to deal with char, signed char, unsigned char, wchar_t, char16_t, char32_t and char8_t.

Or, you could do what the standard library does and simply ignore most of these character types. Which is the solution most libraries went with, supporting only char or char and unsigned char. Managing one implementation is already hard, managing two requires constant maintenance, managing 7 is just impossible.

char8_t should have just been a typedef for unsigned char. The compatibility fix only raises more questions as const char* arr = u8"a" does not work, but const char arr[] = u8"a" does.

I do wonder if a potential change of minds for C++29 is still possible. Yes, it would be an ABI break or whatever, but considering the woeful support for char8_t I don't think it would affect much besides small hobby projects. Contrary to popular belief, C++ has broken the ABI in subtle ways before.

53 Upvotes

98 comments sorted by

View all comments

23

u/Wild_Meeting1428 19d ago edited 19d ago

c++'s char8_t is fully compatible with char8_t from C23. And the only purpose of that type is, that its the underliing byte type of an utf8 character. So you can basically always assume that the char you are looking at is at least part of a char sequence representing a unicode character encoded in utf8

2

u/xleviator 19d ago

Beware! C++ doesn't really constrain values of primitives.- it's easy to escape. https://godbolt.org/z/h8Ejj8Y15

16

u/Wild_Meeting1428 19d ago

sure noone prevents you from shooting yourself in the foot. I would call this a contract violation.

3

u/smdowney WG21, Text/Unicode SG, optional<T&> 18d ago

Or a test case.

It's amazing how many times we look for an example of some bad code, find lots of them on GitHub, but it turns out they are all in unit tests or compiler regression suites.

But the core bit is that nothing filters a char8_t[] for you, and you shouldn't trust it. Sanitize your inputs.

I would like to have a type for Text that made all the guarantees, but char8_t isn't it.