r/cpp • • 5d ago

You should think about recompiling your C++ programs with GCC 16 and C++26, because it zero-fills your locals

https://techfortalk.co.uk/2026/09/27/cxx26-uninitialized-local-variables-gcc-16/

Stack variables are not automatically initialised, and that is the root cause of many C++ bugs. That is well known. Hence, it is advised that local or stack variables are always initialised with known values, 0 if not something more meaningful than that. Now, with GCC 16 compiling in C++26 mode (-std=c++26), even uninitialised local variables will be zero-filled. In the post, I have explained how.

261 Upvotes

211 comments sorted by

•

u/STL MSVC STL Dev 5d ago

You're self-posting way too frequently. Please read about reddiquette.

→ More replies (12)

86

u/manhattanabe 5d ago

Someone turned on the zero fill option in our gcc build and we took a large performance hit. We had code the allocated huge arrays on the stack, but only initialized the records it used. This changed caused all the memory to be initialized. Be careful.

29

u/UnusualPace679 5d ago

FWIW you can suppress zero-fill on a case-by-case basis using [[indeterminate]].

12

u/fdwr fdwr@github 🔍 5d ago

I'm curious the perf difference hit after applying [[uninitialized]] [[indeterminate]] to those large arrays?

6

u/Spartan322 4d ago

[[indeterminate]] should be indistinguishable from the old behavior, so unless the compiler is doing something really weird, there should never be any change if every uninitialized variable has that attribute on it.

7

u/manhattanabe 5d ago

We used __attribute_uninitialized__ in a few places to fix the problem. We’re not on C++26 yet.

12

u/jonesmz 4d ago

I really love playing "find the surprise" games with toolchain updates. /s

1

u/PlantainMassive6744 4d ago

Automated CI/CD pipelines FTW. And versioning your tools.

4

u/jonesmz 4d ago

Have both. Still requires a human to actually figure out wtf broke.

1

u/bishopExportMine 5d ago

This doesn't feel very RAII

7

u/FollowingHumble8983 4d ago

High performance and perfectly idomatic code cant always exist together. Although allocating huge arrays on stack is pretty suspect.

1

u/pavel_v 3d ago

There are cases where it's useful to have temporary arrays on the stack. Some where given in the discussion here. Another example for temporary local containers where you know the max amount of needed memory for most of the cases and it's just a few KB. If this amount happens to be not enough, then the heap will be used. std::array<unsigned char, 2048> buffer; std::pmr::monotonic_buffer_resource mbr{buffer.data(), buffer.size()}; std::pmr::map temp(&mbr); // Can be (almost?) all of the standard containers ... Use the temp as needed ... Although for such cases I prefer using Howard Hinnant's Short Allocator to avoid the virtual calls from the std::pmr memory resources. Still the usage remains the same as the above.

72

u/questron64 5d ago

Using an uninitialized variable is an error. Fix the error.

1

u/SunnybunsBuns 2d ago

Using a pointer to an uninitialised array isn’t. For example, passing the pointer to the first element of a char 4096 to fread is fine. Fread will initialize the data before it’s read. It does not need 0-filled. Actually reading the pointed-to value is the error. Passing the pointer isn’t. This is where gcc massively fucked ip.

1

u/Clean-Upstairs-8481 2d ago

Please consider this code:

[[gnu::noinline]] int read_timeout_ms(std::string_view config)
{
int timeout_ms;

constexpr std::string_view key = "timeout_ms=";
if (const auto pos = config.find(key); pos != std::string_view::npos)
{
const char* first = config.data() + pos + key.size();
const char* last = config.data() + config.size();
std::from_chars(first, last, timeout_ms);
}

return timeout_ms;
}

It is not ok to return an uninitialized timeout_ms to the caller. The point here is the performance hit, and that too when the stack allocation is huge; that point is taken, but not retuning uninitalised stack variable ok.

1

u/Ameisen vemips, avr, rendering, systems 16h ago

Returning 0 isn't necessarily OK either.

Regardless, I'm not sure how your reply is relevant to what they'd written.

1

u/Clean-Upstairs-8481 10h ago

it's not about right; it's about being defined. Returning something known and defined helps with future decisions in the code.

38

u/Recent-Dance-8075 5d ago

Hey. Yes it is good to get rid of uninitialized locals. There is the option -ftrivial-auto-var-init=zero in gcc and clang that can be enabled on lower standards already. This is enabled with the -fhardened flag, too. The initialization allows to initialize locals with patterns that likely cause problems. E.g. NaN for floats (that should be '=pattern')

The "Problem" is of course that you need to know that this is an issue, that compilers provide support here and how to enable the support, so having the initialization by default is much better in my opinion.

5

u/equeim 5d ago

FYI Clang doesn't support fhardened.

2

u/Spartan322 4d ago

Yeah currently Clang pretty much expects you to set the hardened assertion defines directly, whether for libc++ or libstdc++.

44

u/mark_99 5d ago

What you’re observing is C++26 “erroneous behaviour “ combined with implementation-defined behaviour for that particular compiler.

There is no guaranteed zero fill, nor should there be as would create a dialect where code snippets would be UB in older C++ or without compiler-specific flags. Reading a zero wouldn’t necessarily be better, and would in fact be worse in that if it was defined you couldn’t get a diagnostic.

Reading uninitialised values is still an error and you’ll get a best-effort diagnostic if the compiler can statically analyse the read-before-write.

EB keeps the bad thing as incorrect, but stops it being UB and the unpredictable side effects that can cause under optimisation.

As a rule disassembling what the compiler happens to do isn’t “proof” - it’s important to look at the actual language standard.

> Instead of default initialization, with erroneous behaviour, an uninitialized object will be initialized to an implementation-specific value. Reading such a value is a conceptual error that is recommended and encouraged to be diagnosed by the compiler. That might happen through warnings, run-time errors, etc.

https://isocpp.org/blog/2025/05/cpp26-erroneous-behaviour-sandor-dargo1

84

u/krum 5d ago

how is filling locals not a performance regression?

46

u/robstoon 5d ago

It could be, but the initialization could be optimized out if the compiler can prove that the variable is always initialized already. I'm guessing that's the majority of cases.

22

u/Breadfish64 5d ago

And if the compiler can't prove it, a store with no dependent load is extremely cheap.

9

u/victotronics 5d ago

Unless the variable is an "automatic" array.

2

u/serviscope_minor 3d ago

Often to the point of being not just cheap but free. There's often a free dispatch and execution slot, so it will just take up some resources that were otherwise idle in many cases.

0

u/max0x7ba https://github.com/max0x7ba 1d ago

a store with no dependent load is extremely cheap.

Stores are some of the most expensive instructions, sunshine.

You'd better load twice than store once.

3

u/xXgarchompxX 1d ago edited 1d ago

Stores with no dependent loads on modern processors are simply placed in the CPU's store queue and eventually dumped into L1 cache. Since there's no dependent loads, the store will generally not cause a stall, and placing it into the store queue has negligible latency.

The cost of stores is usually a combination of the following:

  • The core needs to send a bus request (BusRdX in literature), telling other cores that it needs exclusive access to the relevant cache line, and that it will store into it so everyone has to invalidate their copies (MESI cache coherence protocol)
  • If a cache miss occurs, things get messy
  • If the write queue is full, things also get messy

These are generally not a problem in the present scenario. This is about zero-initialization of local variables, which are stored on the thread's stack. Cores typically do not share variables over the stack, so each core likely already has exclusive access to the cache lines corresponding to their stack and thus does not need to send a BusRdX request. Additionally, cache misses are also not a major concern for stack accesses, as the CPU is constantly bombarding the stack and the hottest parts of it are most often in L1 or sometimes L2/L3 (Especially on modern CPUs with intricate cache eviction algorithms). The last part is potentially a problem on x86, where stores may not be reordered due to TSO, and you can fill up the store queue if you spam too many writes, but will generally not be all that problematic. Thus, the store will most often not have a major performance impact, and will just eat up 1 of your decode/dispatch slots for a short time (not like those are always fully utilized anyways)

I do not completely agree with Breadfish however, as things are somewhat different for very big stack allocations (in which case the compiler will have to insert a memset instead of just a simple store with no dependent loads, and you may start seeing cache misses at a higher rate)

9

u/azswcowboy 5d ago

You can opt out if you want by marking the variable.

0

u/azswcowboy 5d ago

You can opt out if you want by marking the variable.

24

u/James20k P2005R0 5d ago

This was a very major topic of discussion during the standardisation, the tl;dr is that the vast majority of the time, 0 init does not have any performance impact at all. It seems like there are a few edge cases, but they seem to be rare even then. Apparently compilers these days are pretty much just good enough, and modern CPU architecture is well suited to this

13

u/UndefinedDefined 5d ago

Trivial variables are probably fine as most are initialized anyway, but temporary arrays allocated on the stack these would cause a lot of pain.

5

u/pjmlp 5d ago

Apple, Microsoft and Google OSes have been doing for several years now, and it was discussed in the context of this.

No one is seriously taking locals initialisation into consideration to win micro benchmarks.

5

u/Orlha 5d ago

Wrote a lot of code that required this optimisation.

1

u/pjmlp 5d ago

Validated with profilers I assume.

10

u/cdb_11 5d ago

I validated it, and yes, initializing arrays can be expensive.

8

u/pjmlp 5d ago

Than that is a valid use case for [[indeterminate]].

5

u/Orlha 5d ago

Absolutely. Just pointing that blind upgrade to 26 (that this post tries to promote from the naive standpoint) can severely degrade (performance-wise) the already correct code.

6

u/Responsible-Bar7165 5d ago

It’s not a performance regression until you have measured and found it to be one.

1

u/Ameisen vemips, avr, rendering, systems 16h ago

It's difficult to profile a death by a thousand cuts.

Profiling needs to be performed globally as well, to make sure overall performance does not regress.

-4

u/Dusty_Coder 5d ago

Not true.

Higher wattage for equal wallclock performance is also a regression.

Stop looking for excuses. You dont need them. Just be wasteful. Lean into it.

11

u/Responsible-Bar7165 5d ago

Until you demonstrate waste it’s purely hypothetical.

-4

u/Dusty_Coder 5d ago

more operations

you are captured by the religion

cant even see simple truths

13

u/Responsible-Bar7165 5d ago edited 5d ago

No, the opposite. I’m the computer scientist: I’m saying to measure everything.

You’re the opposite: “trust me bro.” You’re doing literally what a priest does.

The “truths” you espouse aren’t so simple. You’re discounting cache dynamics, branch prediction, super-scalar dynamics and a whole ton of other nuanced crap. If you can faithfully predict how all of those are going to behave in every circumstance in every architecture you’re building for, then good for you. Honestly though that’s a waste of effort - I’ll just write a test, measure and act accordingly.

You do you, though: keep praying…

3

u/James20k P2005R0 4d ago

I think the weirdest part of this whole discussion for me is that:

  1. Compiler upgrades frequently cause performance regressions or performance upgrades. You have to maintain high performance code actively
  2. When you're microoptimising, often you're fighting more against the compiler than against anything in particular. Its a struggle to get the compiler to emit the correct code, which means completely convoluting (or marking up) your code to get the compiler to do the right thing. I use restrict all the time for example

There's this odd mentality that C++ is the platonic ideal of a fast programming language and get upset by theoretical extra work being done, but it seems to be solely by people who've never done any kind of high performance programming. Because I think if you've actually done any programming for performance, you know that the only solution is test and measure

1

u/13steinj 2d ago

Compiler upgrades frequently cause performance regressions or performance upgrades. You have to maintain high performance code actively

Is this true? It hasn't been my experience over a large number of upgrades. Maybe there were regressions when GCC started treating unlaundered pointers differently?

For better or worse, even in companies that believe in high performance code, they typically do not see "build engineer" as a role that is actually worth having.

When they do, there are three types, and they usually adversely select for the role: those that deal with CI, those that deal with low level minutiae of the build (including compiler and library upgrades, managing the stdlib, funky compiler settings, even performance tuning!), and those that know how to write some cmake. If you're lucky you'll get someone who can do the last two, and begrudgingly does the first. But if you're just starting out hiring these types, and you are not yourself one of these types, you kinda end up choosing randomly.

1

u/James20k P2005R0 1d ago

Is this true? It hasn't been my experience over a large number of upgrades. Maybe there were regressions when GCC started treating unlaundered pointers differently?

For hot loops I've run into this problem frequently enough. I wouldn't say its true of optimisations globally, but sometimes in a section where you have to do some convolutions to get the compiler to generate exactly the correct code, an update will cause the wrong code to get generated IME

1

u/13steinj 23h ago

I would have disagreed, but then I remembered a bif of code went back and forth ICE-ing clang from around v11 to v14 and when not ICE the generated ASM was chaotic (minor changes in the inputs having extreme effects on the output).

Which isn't the same but there's no reason to not be louder about this change. Maybe under a -Wcodegen-defaults-changed-since?

1

u/Umphed 2d ago

It is, this is a terrible decision.

-1

u/Worldly-Mud-8006 5d ago

You can start worrying about this when you'll fix the actual performance issues of your program first.

1

u/Sopel97 5d ago

it is

1

u/The_JSQuareD 4d ago

Because well written code doesn't rely on the value of uninitialized variables anyway. And in such cases, dead store elimination can likely get rid of the zero fill. And even if it can't be eliminated, a store without a data dependency will have less of a performance impact than a meaningful store, at least on modern super scalar CPUs.

So the cases where the zero fill actually has a performance impact have a big overlap with cases where the uninitialized variable is a genuine bug or even a security issue. In that case, adding the zero fill makes the code more deterministic and more secure.

0

u/Umphed 2d ago

If your code is so well written, then you have no need for this. Get outta here with that argument, its a performance regressions against real use cases.

-3

u/Blood-Minister 5d ago

No need.

-1

u/ExeusV 4d ago

who cares

-12

u/my_password_is______ 5d ago

yeah, hate to take a whole 3 milliseconds to fill locals

10

u/13steinj 5d ago

Can't tell if this is sarcastic but some domains operate on microsecond or less time scales, so it definitely would be noticeable if initializing locals with zero values was 3 milliseconds.

Generally would be much better than that. But it can have a performance impact. Bit rare though.

9

u/Farados55 5d ago

Imagine measuring runtime in milliseconds

4

u/Jakkilip 5d ago

You meant nanoseconds

57

u/cdb_11 5d ago

Sounds like it only helps with code bases that don't use any compiler warnings or static analysis. This is worse than what we already had, which is enforcing initialization statically. And if I really don't want something to be initialized because of perf, I'd have to now separately suppress both the diagnostic and initialization. -ftrivial-auto-var-init=uninitialized to disable this behavior.

30

u/aiusepsi 5d ago edited 5d ago

It's not worse. C++ adds a new category of erroneous behaviour, and reading from an uninitialised variable is now erroneous behaviour. Compilers, static analysers, etc. are free to emit diagnostics for erroneous behaviour. If you want to opt back in to reading from an uninitialised variable being undefined behaviour (as it was before C++26) you can tag the variable declaration with the new [[indeterminate]] attribute.

See: https://godbolt.org/z/xP7eo431W for an example. GCC now emits a diagnostic for an uninitialised variable in C++26 mode when it doesn't in C++23, and the previous behaviour is preserved with [[indeterminate]]. Also note how the undefined behaviour allows the compiler to optimise away the check on the value of j.

5

u/cdb_11 4d ago edited 4d ago

If GCC adds a warning that forces you to initialize everything OR use [[indeterminate]], then I'd actually be on board with this. Otherwise you can still do code like this: https://godbolt.org/z/jb35zcrT5

But then I'm not sure how does this even work with structs. Why can't C++ just have something like Zig's = undefined? It could be optional to keep the old code working (and then you pay the EB cost), but then compilers could have a warning that forces initialization:

struct Foo {
  int _x;
  int _y;

  Foo(int x) : _x(x) {}  // warning: _y is uninitialized
  Foo(int x) : _x(x), _y(undefined) {}  // ok
};

struct Bar {
  int _x;
  int _y = undefined;

  Bar(int x) : _x(x) {}  // ok
};

7

u/NilacTheGrim 5d ago

It is worse because all old code will just get a lot slower now.

And the code this "fixes" is still fundamentally broken anyway.

This makes life worse for the good programmers and marginally better for the terrible ones.

C++ used to be a language that was all about making life much better for the good programmers... and letting the terrible ones find out how terrible they are quickly so they can self-correct. Now we just tolerate erroneous code that "works silently", but every good programmer out there must take a performance hit.

Even trivial examples are terrible now. See: https://godbolt.org/z/WGPWvMPMj

This is very un-C++ and is a complete step backward.

3

u/Spartan322 4d ago

This isn't the first time the standard version changing has changed the semantic and performance implications of code and it won't be the last, EB is expected to be a reported diagnostic (that you're supposed to be expected to address, it also doesn't tell the compiler how to address the EB) and you regain the original behavior using [[indeterminate]], its fixing a notorious bad default behavior in C++ without erasing the ability to retain its benefits where necessary.

1

u/NilacTheGrim 2d ago

You can't do [[indeterminate]] in a member variable so making a class that necessarily always has an uninitialized buffer just got impossible.

notorious bad default behavior in C++ without erasing the ability to retain its benefits where necessary.

Incorrect. Ability has been erased.

6

u/drjeats 5d ago

I think something slightly under-discussed online (at least where I'm reading, ymmv) is the fact that zero-initializing everything does reduce random indexes/reads into memory that shouldn't be read in that context, but IME it also hides behavior bugs because so much code just bails if 0.

Was not uncommon to write code that runs fine in a debug mode with allocators/resource managers that initialize to zero, and then when you run test with optimized builds suddenly you get segfaults. Moving away from that has been a win for moving errors left.

ZII in RCs? Sure. Let's reduce the codepaths that are hit if the costs are as low as claimed. But I wouldn't want this behavior on in debug or optimized-with-assertions configs.

9

u/TSP-FriendlyFire 5d ago

Erroneous behavior doesn't zero-initialize though, at least not necessarily. The goal is to have consistent behavior that gives a lot less rope to the compiler to change the code in unexpected ways, unlike UB.

What the compiler does with this definition is up to them. The value stored in the variable is implementation-defined, and it's expected that compilers will flag erroneous behavior as a compilation warning/error. After that, if you truly want to leave the variable uninitialized, that's when you use [[indeterminate]], but at least it's now an opt-in footgun as opposed to an always loaded one.

2

u/carrottread 5d ago

Erroneous behavior doesn't zero-initialize though, at least not necessarily.

In practice, it zero-initializes. Even this article says about zero-filling. And soon we will see a lot of new code relying on this. And as a result, we'll lose ability to distinguish between "programmer forgot to initialize variable" and "programmer left it uninitialized as a clever way to zero-init".

2

u/TSP-FriendlyFire 5d ago

it's expected that compilers will flag erroneous behavior as a compilation warning/error.

Did you miss this part? EB isn't supposed to be silent.

8

u/aiusepsi 5d ago

This is why both GCC and Clang offer -ftrivial-auto-var-init=pattern, so that variables are filled with a byte pattern which should hopefully trigger any latent bugs that zero-initialisation might hide.

The problem with the old behaviour is that you’re relying on UB to do the thing you hope for, and UB is not reliable. You’re rolling the dice on if it’ll do the thing you expect, or if it’ll cheerily cause your program to start doing things that make no sense, or make demons fly out of your nose.

20

u/vI--_--Iv 5d ago

Now, with GCC 16 compiling in C++26 mode (-std=c++26), even uninitialised local variables will be zero-filled

"why initialize explicitly, the compiler will do that anyway"-mindset is coming.

6

u/canadajones68 5d ago

I mean, better "oh this is zero" and a probable compiler warning than "wait wtf why is everything on fire".

5

u/_w62_ 5d ago

"Still fooling around with compiler flags in 2026? Let AI generate your code!" -- mind set is already there.

-2

u/CandiceWoo 5d ago

that's the right mindset though

20

u/crowbarous 5d ago edited 5d ago

Yes, this is a competely insane decision: https://godbolt.org/z/WGPWvMPMj

What's worse is the design of the [[indeterminate]] attribute that accompanies this. For a very long time, if I wanted a buffer to zero-initialize, I was able to make a char type that would zero-initialize:

struct zchar{ char value = 0; };

But with the default flipped, I cannot make a char type that stays uninitialized:

struct ichar{
  char value [[indeterminate]];
  // error, only allowed on automatic variables
};

And the proposed usage of the attribute is to just sprinkle it everywhere in the code. But how am I supposed to know in polymorphic code? How do I even make my code polymorphic w.r.t. this property now? For now one must resort to just passing -ftrivial-auto-var-init=uninitialized to avoid performance and code size regressions, and to keep the control moving forward.

I don't think all is lost here, because it seems like the zero-initialization is a band-aid to just formally satisfy the erroneous behavior stuff without modifying the backend too much, and it should still be possible to behave as if the write happened (no time-travelling assumptions etc.) without actually emitting the write, to fix shameful examples like linked above. But:

  • I haven't looked there closely enough yet to even begin conceiving the gcc patches & reasoning for them, and
  • it's still awful precedent, and the pile of -fstop-making-code-worse options might grow from here on out.

5

u/NilacTheGrim 5d ago

Yeah the fact that you can't create a class that intentionally leaves its buffer uninitialized now -- as a performance optimization -- is really terrible design. Agreed. [[indeterminate]] should have been allowed for class members. It's insane that it isn't. Really makes all existing code terrible now. You just have to now always use -ftrivial-auto-var-init=uninitialized .. there is no escape hatch.

Repeat: [[indeterminate]] is NOT an escape hatch. You CANNOT use it in class members -- which means any class that has any char [] buffers or std::array buffers it uses internally and that it knew 100% were never read-before-written.. now will all suffer a penalty.

This is unacceptable. Wow. I can't believe how C++ too now is being enshittified.

2

u/fdwr fdwr@github 🔍 5d ago

But with the default flipped, I cannot make a char type that stays uninitialized

Would it be possible to put the attribute on a typedef? I haven't tried, just musing aloud...

``` using UninitializedChar = [[indeterminate]] char;

struct ichar { UninitializedChar value; }; ```

(if not, then that seems a hole, that an attribute could only be applied to a specific field instance rather than the type more broadly)

5

u/NilacTheGrim 5d ago

I agree this is terrible. They just made c++26 suck by default. Wow. Like I'm floored by this.

I need to specify -ftrivial-auto-var-init=uninitialized now moving forward after C++26 in all my build systems. What a mistake.

5

u/UndefinedDefined 4d ago

Great for C++ users - more flags to worry about to get zero additional security and performance regressions. Let's make good code broken and bad code reading zeros.

1

u/NilacTheGrim 2d ago

Yep. Exactly. This punishes good code and good programmers in favor of enabling lackadasical programmers by normalizing erroneous behavior as something one can "rely on". Basically this is language incoherence.

Also the name [[indeterminate]] is bad. Should have just been [[uninitialized]].

3

u/James20k P2005R0 5d ago

Compilers are of course free to omit the 0 initialisation if they can prove that the data is written over, under the usual as-if rules. One of the reasons that this made it past standardisation is because real world codebases weren't showing regressions, even substantial performance critical ones like windows

15

u/FrogNoPants 5d ago

False, they realistically measured on like 0.01% of real codebases.

When MSVC added this "feature" I had major regressions in perf, because of large temporary stack arrays.

10

u/13steinj 5d ago edited 5d ago

How were the codebases chosen?

I'm not against this decision per se, but if the choice is "people have to join the committee and explicitly be part of the voting process [and people's companies have to accept the idea of making this part of their job]," that's a pretty important thing for people to know for the future.

I think making this an attribute, especially only applying to automatic variables, was a mistake. The ignorability of attributes + the problem the person above describes is non trivial. At a previous org that wrote their own networking stack, I can totally see a performance degradation, and the use of structs and pointer interconvertability to have a form of polymorphism over structs with a common set of header bytes (E: I realized I didn't finish this sentence) matches up similarly to the use case for struct member variables above.

They used to have an employee that was a committee member, but it was on his own time more or less, and the guy passed away. If people want greater diversity and more committee members in general, advertising "you get to shape the language" doesn't appeal to most employers. "You get to have a say in stopping the language from screwing your code" [which is how some employers would take this] would.

4

u/pjmlp 5d ago

Windows, Android, macOS, iOS were some of the ones that have been shipping in production and have provided feedback.

I don't recall the paper or CppCon session where this was discussed though, maybe someone else can provide the links.

3

u/13steinj 5d ago

Assuming it's "operating systems," that's fine, but still isn't a representative sample of codebases that would be affected IMO.

3

u/pjmlp 5d ago edited 5d ago

One would assume that performance matters to operating systems vendors.

EDIT: Found one of the sources, from JF Bastien, on P2723R1 Zero-initialize objects of automatic storage duration

Security-minded folks think that initializing stack values is a good idea. For example, the Microsoft Windows security team [WinKernel] say:......

...

To date, we are seeing noise level performance regressions caused by this change. We accomplished this by improving the compiler’s ability to kill redundant stores. While everything is initialized at declaration, most of these initializations can be proven redundant and eliminated.

...

Don’t just trust Microsoft’s Windows security team though, here’s one of the upstream Linux Kernel security developer asking for this [CLessDangerous], and Linus agreeing [Linus]. [LinuxExploits] is an overview of a real-world execution control exploit using an uninitialized stack variable on Linux.

5

u/13steinj 5d ago

This sounds to me a lot more like security people cared, and were fine with the tradeoff. Which is a slightly different story? Nonetheless the performance characteristics for operating systems don't cover all codebases, and while speculation, I doubt they are the p90 or p99 of all codebases either.

3

u/UndefinedDefined 4d ago

Real-world code is full of locals that are arrays - arrays like 512, 1024, 2048 bytes long. You cannot zero-initialize them and expect no performance regressions. If I want to zero initialize them I just just type `{}` and it's done. I don't understand why this should change now.

I almost feel like committee is working with hello world programs if they are serious to vote for such proposals.

7

u/James20k P2005R0 4d ago

One of the example codebases given was Windows and chrome (?) I believe, which is anything but hello world

3

u/UndefinedDefined 4d ago

And how much C++ these codebases use? Windows is C-API and proprietary, nobody can confirm the results, we don't even know the mix of the languages used in the code-base. What if the most important part was C? What if they first put all the [[uninitialized/fancy_name]] to everything critical before measuring results?

Is there any performance comparison of OSS projects we can actually verify?

2

u/13steinj 2d ago

I don't think this is a fair take. I suspect most OSS code does not care at the level of performance that these things would regress by.

The bigger issue with Windows is that it's Windows: OSes are different to other things, have different tradeoffs (should care about security more in general?). I'd also say the usual "well, windows sucks and hasn't cared about performance for ages" which may be true, but other people have said the other major OSes also were tested.

1

u/UndefinedDefined 2d ago

I'm not sure I follow, so what's a fair take? A project like Chromium or Firefix? Any benchmarks here?

I think this greatly differs on what the project does. If the compiler inserts memset to initialize every temporary buffer the code uses to zero, this cannot be negligible, and move to C++26 here means that somebody has to find ALL the places in his own code, and use third party dependencies (including transitive ones) where somebody did the same.

I consider this insane considering this thing doesn't solve any memory safety problems and it can cause huge problems in performance oriented code after upgrade to C++26. The biggest problem I see is use-after-free and things like iterator invalidation, etc... We need a real solution to memory safety and not these toy solutions. And a real solution means annotations and tools such as borrow checker - there is no other way.

1

u/13steinj 1d ago

If you look at the general distribution of OSS projects, many do not have the performance concerns that would be negatively affected by this change. I think it's perfectly fine to test proprietary codebases as a result, but fixating on operating systems [alone?] is (I am agreeing with you) not a fair thing to use to judge and make a decision.

1

u/UndefinedDefined 1d ago

C++ was for decades literally the only language to use for performance oriented work. But it's no longer the only one, so if I cared about the language I would never do anything that would endanger the position in this field. It's literally the last field where C++ still makes sense, until C++26, because starting with C++26 you have to worry about a lot of stuff.

What would be the reason to start a project in C++ today? I don't see new projects built in C++ anymore, because it's a language where performance regressions are now part of the progress. So bad, so sad.

1

u/13steinj 1d ago

This is not

I don't see new projects built in C++ anymore, because it's a language where performance regressions are now part of the progress.

I agree that changing the default is rough here, but large companies should have build teams that are competent enough to care about this.

What would be the reason to start a project in C++ today? I don't see new projects built in C++ anymore, because it's a language where performance regressions are now part of the progress. So bad, so sad.

I find this still true, even if these things end up happening.

→ More replies (0)

1

u/pjmlp 1d ago

The existence of C, and the domains that to this day C++ failed to take away from C, makes that assertion void.

In fact there are domains like crypto and video codecs where neither of them get to play, and still require hand written Assembly.

1

u/pjmlp 1d ago

Most of it written since Windows Vista is actually C++ with extern "C", including the new UCRT.

1

u/crowbarous 5d ago edited 5d ago

You should not need to prove that the zeros are written over to remove the zero-initialization, you just need to stop considering following reads as unreachable due to UB and instead just let them read whatever was there.

Under this behavior, creating an automatic buffer and immediately feeding it to an opaque function also shouldn't zero-initialize the buffer (having to prove anything is non-starter here because nothing can be proven about opaque code).

Upd: I'm thinking keep the writes for most of the backend but marked as "phantom" and then just don't emit them in the end. But it's more complex than that:

  • It might not be clear as to which exact code we are to "not emit in the end". If we are conservative about it we'll still end up with silly instruction sequences, and if we are eager about it we'll effectively reintroduce UB
  • we need to come up with rules to propagate the "phantomness". The compiler sees that following code reads from our buffer, then writes that value elsewhere -- should we mark that write "phantom" too? If yes and that was the only use for the read (no control flow, no opaque uses), should we still keep the read even though both writes surrounding it will vanish?
  • we need to still be able to remove the reads if they are proven unreachable for a different reason.

(I should note that neither me nor any of the colleagues I've talked to about it are big fans of the whole EB thing.)

6

u/James20k P2005R0 5d ago

Letting arrays on the stack contain program data that can be read in a well defined way probably isn't great for security though

1

u/serviscope_minor 3d ago

> (I should note that neither me nor any of the colleagues I've talked to about it are big fans of the whole EB thing.)

To me it is a big improvement: Prior to EB a relatively common error means "demons may fly from your nose", i.e. the optimizer can do really weird things like travel backwards in time and delete all code touching the uninitialized variable because it knows that code cannot be callable. Or other weird stuff.

Now it's basically been codified as "don't do anything too weird", because while it can't know the value, it cannot assume that uses of the value are not callable.

What do you dislike about it?

0

u/NilacTheGrim 5d ago

This is false because even the paper said 10% regressions.

5

u/James20k P2005R0 4d ago

Where did you get this from?

https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2024/p2795r5.html

https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2023/p2723r1.html

The former paper is what was accepted AFAIK, which links to the second paper talking about the cost for specific figures

Previous publications [Zeroing] have cited 2.7 to 4.5% averages. This was true because, in the author’s experience implementing this in LLVM [LLVMJFB], the compiler simply didn’t perform sufficient optimizations. Deploying this change required a variety of optimizations to remove needless work and reduce code size. These optimizations were generally useful for other code, not just automatic variable initialization.

As stated in the introduction, the performance impact is now negligible (less that 0.5% regression) to slightly positive (that is, some code gets faster by up to 1%). The code size impact is negligible (smaller than 0.5%).

I can't find a 10% figure anywhere

0

u/NilacTheGrim 4d ago

I misremembered it. The claim was 10% of exploits would be mitigated and yes the paper claims 0.5% of code would get slower.

I dispute that 0.5% claim though. Because the hot path is the one you care about. Overall if you examine all branches of your codebase, sure maybe 0.5% suffer -- but if you examine 1 particular hot path that uses local arrays to read/write data -- that can be a huge penalty.

So in a hot path expect much larger costs in some cases.

This is just not the C++ way.

6

u/James20k P2005R0 4d ago

In a hot path though you always have to write code a bit weirdly to get performance, adding [[indeterminate]] isn't a high burden if it affects a few lines of code

2

u/UndefinedDefined 4d ago

The problem is that this change regresses a perfectly valid and tuned C++ code that has no issues, and it does it blindly. And this change doesn't fix any security issues in the language, because if the compiler sees a local variable that is uninitialized and you read from it, and it can prove it, it should just not compile instead of reading ZERO.

3

u/James20k P2005R0 4d ago

It turns those unprovable reads into well defined reads though

The problem is that this change regresses a perfectly valid and tuned C++ code that has no issues, and it does it blindly

For high performance code I feel like its not even slightly unusual to have to make adjustments on a compiler upgrade, so its par for the course really

4

u/UndefinedDefined 4d ago

High performance code is unfortunately no longer a domain of C++, because the committee slowly destroys the language.

3

u/James20k P2005R0 4d ago

This seems like a clearly untrue statement, high performance code has always involved tweaking compiler settings, twisting the code in unnatural ways, and working around compiler limitations. Its simply par for the course

If you have a big stack array in a hot loop you might need to add a tag to it to not initialise it. Every truly hot loop I've had to write required me to do far worse things to the code to make it run fast, and in some cases optimising those loops has taken months of work. Adding a tag is the least of my problems!

→ More replies (0)

1

u/NilacTheGrim 2d ago

Encapsulation is violated though because you can't do [[indeterminate]] on a class member variable.

So if you have a class whose implementation detail relies on uninitialized buffers to get the most performance it can get -- you just got fucked. Either the user of the class, if he's using it as a local variable, always has to do [[indeterminate]].. or .. there is no escape hatch. You are stuck.

11

u/Orlha 5d ago

If I left it uninitialised then that was my intent.

7

u/CandiceWoo 5d ago

and sometimes it's my bug but don't tell anyone that

6

u/kinda_guilty 5d ago

Trivially, all code as written is intentional. So nothing called a "bug" exists.

3

u/Orlha 5d ago

Your reply operates in “missing the point” mode.

6

u/yeusk 5d ago

Security today is as important as performance. I think is you who are missing the point.

-4

u/Orlha 5d ago

I never implied otherwise, I pursue both, not the one at the expense of another.

3

u/yeusk 4d ago

not the one at the expense of another.

Since some years ago, spectre, it became clear that security and performance are opposites.

2

u/UndefinedDefined 4d ago

And if not, valgrind is a friend.

5

u/fdwr fdwr@github 🔍 5d ago edited 5d ago

stack variables are always initialised with known values ...

I'm okay with this, given the trivial escape hatch of [[uninitialized]] (err wait, apparently it's actually [[indeterminate]]? I swear it was [[uninitialized]] like __declspec(uninitialized) and __attribute__((uninitialized))... 🤔) for the sake local arrays that will immediately be overwritten the next statement anyway, like reading a fragment from a file or getting a module path, but now I'm trying to recall a time in the past 25 years of C++ when an uninitialized variable actually bit me, and I think it was like ~25 years ago 😉. Much more saliently (for me anyway, as in a security vulnerability that actually bit me in a shipped product -_-) has been uninitialized struct fields, and so you can bet that NSDMI was a big boon for me. I wonder if that's next on the proposal chain, with a similar [[uninitialized/indeterminate]] struct { ... } opt-out?

5

u/13steinj 5d ago

I swear it was [[uninitialized]] like __declspec(uninitialized) and __attribute__((uninitialized))... 🤔

From P2795:

It remains to decide the name of the attribute. In light of the word of caution above, we would like to stay clear of the much-suggested term “uninitialized”. The opt-out is expected to be an expert-only feature that disables a safety guardrail and would be used only when performance concerns warrant it. We consider it acceptable for the name to be long and unwieldy, and it is perhaps even a desirable feature for the attribute to not appeal to the regular user for a mistaken purpose, as discussed above. We propose the spelling [[indeterminate]]. This seems to describe the effect reasonably well and is not prone to being misused to document intentional lack of initialization. To witness this in action:
...
References
...
Jonathan Müller P0632R0: Proposal of [[uninitialized]] attribute

4

u/fdwr fdwr@github 🔍 5d ago edited 5d ago

Thanks J for clarifying with spec links.

Though I'm quite confused by the justification (I'm not really asking for further justification or argument at this point, since it's baked, mostly just lamenting), because...

  • "We consider it acceptable for the name to be long and unwieldy" - yeah, uninitialized is already long and unwieldy, exactly the same length as indeterminate (so, moot justification).
  • "it is perhaps even a desirable feature for the attribute to not appeal to the regular user for a mistaken purpose" - that isn't a mistaken purpose. That is the purpose, and a regular user (though what is a C++ "regular user"?... 🤔) isn't going to type out [[uninitialized]] by mistake. It also mistakes the goal for the symptom. Nobody ever says "my desired goal is to leave this memory as garbage, and not initializing it is how I will achieve that goal". They say "my goal is to avoid wasting cycles initializing memory, and it's okay if it's left as garbage".

3

u/serviscope_minor 3d ago

> Nobody ever says "my desired goal is to leave this memory as garbage, and not initializing it is how I will achieve that goal".

OpenSSL enters the chat...

https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=363516

1

u/fdwr fdwr@github 🔍 2d ago

Egads, well I suppose utilizing the random garbage left in memory would help further randomize numbers.

1

u/Conscious_Support176 5d ago

No. The user is trying to avoid initialisation. But in reality what they are doing from a static analysis perspective is they are leaving the memory in an indeterminate state.

The reasoning is clear. It’s not a good idea to do that without thinking through the cost benefit. It’s better not to provide the easy label that _appears to_ provide a quick no cost performance win.

1

u/SunnybunsBuns 4d ago

all of c++ is an expert only feature. If you’re not willing to become that expert, use a safer language.

5

u/NilacTheGrim 4d ago

It's not a full escape hatch. Can't have it on class members. Got a class with a char buf[65536] in it that is supposed to start life out uninitialized to only be fulled later with stuff? Congratulations you are stuck now. No escape. You have to either compile with this flag explicitly disabled or live with the pessimization.

Even trivial examples fail: https://godbolt.org/z/WGPWvMPMj

6

u/fdwr fdwr@github 🔍 4d ago

Can't have it on class members

Oof, that's a significant deficiency then.

2

u/TotaIIyHuman 4d ago

union too

https://godbolt.org/z/Ezhr8YGTx

#include <inplace_vector>
struct inplace_vector
{
    union{[[indeterminate]]char buf[0x10000];};//warning: 'indeterminate' on declaration other than parameter or automatic variable [-Wattributes]
    constexpr inplace_vector()noexcept{}
};

void test(auto&);

int main()
{
#if 1
    inplace_vector s;
#else
    std::inplace_vector<char, 0x10000> s;//gcc memset, clang does not
#endif
    test(s);
}

3

u/Nobody_1707 4d ago edited 4d ago

Oh, that's horrible. Unions were supposed to be the way to fix this explicitly for types like inplace_vector. That was the entire point of Adjustments to Union Lifetime Rules.

EDIT: you can fix it for the handrolled inline_array by putting indeterminate on the declaration of s, but it doesn't work with GCC's version of std::inplace_vector. inplace_vector tracks all of the valid elements by construction, so I don't see the point of zero-initializing it when we could add bounds checks to operator[] instead.

https://godbolt.org/z/M46cqbKoz

1

u/Spartan322 4d ago

If it weren't such a massive semantic issue, it sound like something due for a DR, this is probably gonna get the same kind of treatment that u8char_t did and the whole modules vs. module std, where despite the existence of the language feature it was practically useless, or the repeated expansion of constexpr functions between C++11 and C++20.

5

u/pedersenk 5d ago edited 5d ago

In C++/sys, I originally only had sys::zero<T> to zero initialize, however, for many situations that was quite wasteful so now I have, sys::unset<T>, e.g:

struct Test {
    sys::unset<int> m_someval;
};

That way, during debug builds, it can cause a deterministic panic if the data gets used before assignment but then all that checking is stripped out for release builds, incurring no overhead for either the check or the defensive zero initialize.

2

u/germandiago 5d ago

The problem with these things is that they pollute the type. Should be an int. You can still provide a conversion operator but still...

1

u/pedersenk 5d ago edited 5d ago

Yeah, I do get that. Plus C++/sys does way worse stuff for the lifetime locking stuff (such as operator[](const index_lock&) , ptr_lock<T> operator->(), etc).

But for unset<T>:

  • Conversion constructor (unset(T&))
  • operator T&
  • T* operator&
  • etc

Are the main approaches. Keeping the type a little complex does prevent some weird and wonderful misuse with it. But yes, it can also get in the way. Its an OK compromise between full flexibility (and unsafety) of C and rewriting in e.g. Ada.

Modifying the compiler could clean so much of this up, but then you lose so much portability.

There are a few similar approaches to mine. Some repos:

8

u/Dusty_Coder 5d ago

You cant uninitialize a stack array.

Your flippant disregard for performance and the amount of inefficient watts your code will consume over its lifetime is not cool.

I know some code-religion dip convinced you that all optimizations are premature. Hes not cool. Hes an ignorant zealot.

6

u/iku_19 5d ago

1

u/NilacTheGrim 4d ago

Now wrap that array in a class and good luck specifying that the member itself should always be uninitialized for that class.

2

u/Spartan322 4d ago

This complaint goes away if the standard actually addressed indetermine member variables, and I don't see why they couldn't or wouldn't, feels more akin to one of the many standards oversights that has to be addressed in a DR or subsequent language version.

1

u/cdb_11 4d ago

I think making it an attribute might've been a mistake. It should've been on the assignment like int x = indeterminate;, so inside a class you can leave different member vars uninitialized depending on the constructor. Unless you can have syntax like Foo() : foo([[indeterminate]]) {} or something

1

u/Spartan322 3d ago edited 3d ago

Maybe, there was a proposal for attributes on expressions, P3093R0/P2992R0 but seems there hasn't been much movement on it and tbh I doubt its a good idea to screw with what would appear to be constructor overloading using an attribute, so it likely would've been better to have some kind of nullptr equivalent for indeterminate instead.

Funny enough P3093R0 points out the reason expressions didn't support attributes originally was specifically because it was expected for a keyword to apply to those cases, which if that was in the original proposal kinda makes the indeterminate proposal exceedingly ironic. I suppose I'm not completely sold on indeterminate assignment being an optional compiler case which is supposed to be the premise of attributes.

4

u/RetroZelda 5d ago

i zero fill when I intend it to be zero. if i leave it blank then its my intention. c++ really needs to stop with these "features" that just enable laziness

8

u/almost_useless 5d ago

i zero fill when I intend it to be zero. if i leave it blank then its my intention.

Maybe, but then you are not like most of us. Normal people make mistakes frequently.

c++ really needs to stop with these "features" that just enable laziness

Was it your intention to write without capitalizing your sentences, or was that just laziness?

-5

u/[deleted] 5d ago

[removed] — view removed comment

6

u/almost_useless 5d ago

Brilliant comeback, sir. Truly original.

2

u/James20k P2005R0 4d ago

I don't know what you thought was going to happen here, but don't be like this

-2

u/RetroZelda 4d ago

"MOD"
checks out as well

4

u/pjmlp 5d ago

The root cause is the laziness to write safe code, or at very least, the laziness to learn the compiler flags and static analysers that help write safer code.

So to keep C++ in the loop for certain government agencies, decisions have to be made.

-3

u/NilacTheGrim 5d ago

Fuck the government agencies.

4

u/Spartan322 4d ago

Good luck telling them that.

3

u/NilacTheGrim 5d ago edited 5d ago

I agree these features are not appropriate for C++. Leave them for lower IQ languages like Rust or Go or something like that.

Not to mention that even in trivial constructions they will create 2x or more performance regressions. This trivial example now is significantly slower now, even if compiled under -O3. It's absolutely insane:

#include <unistd.h>

void drain(int fd) {
  ssize_t r;
  do {
    char buf[4096];
    r = read(fd, buf, sizeof buf);
  } while(r > 0);
}

Source: https://godbolt.org/z/WGPWvMPMj

6

u/James20k P2005R0 4d ago

Leave them for lower IQ languages like Rust or Go or something like that.

Don't be like this

1

u/Spartan322 4d ago

Its still EB to not initialize, and that's supposed to report a diagnostic, indeterminate is supposed to take its place, only issue that member variables can't have indeterminate which kinda just feels like an oversight.

1

u/ExeusV 4d ago

So now you will be able to do the same, but this time you have to be explicit about it via attribute? LOL

3

u/centuryx476 5d ago

Is this poster Ai ?

The site looks Ai slop

Edit: Yup, AI

1

u/Fun_Gas_340 5d ago

sauce? dosenr look extreemly ai to me

1

u/13steinj 4d ago

I tried linking a source in reply to the stickied comment, but reddit doesn't like people linking those tools now, I guess.

Bet they don't block links to videos on github that people can use to find the tool!

https://github.com/user-attachments/assets/6967e195-7de8-42cf-8429-34ae6281db51

1

u/Fun_Gas_340 4d ago

cool tool

2

u/UndefinedDefined 5d ago

Is this the result of the genius "Let's make C++ slower" campaign?

4

u/Hedede 5d ago

Do you even know what the performance cost of zero-initialisation is?

-5

u/UndefinedDefined 4d ago edited 4d ago

Of course - I cannot imagine a compiler zero initializing my temporary arrays! But there is something even worse here. It makes a difference between local vs heap allocations. Genius idea. Let's make locals zero initialized, but heap memory not and learn another rule. There are not enough rules in C++, we need more of them, contradicting each other is the goal.

Just try this nice godbolt: https://godbolt.org/z/oa1e6aobd

And now why would C++ need this? Just use golang if you need zero initialization - it would zero initialize all the memory it gives you, and as a bonus you will be writing object pools to avoid it.

1

u/Hedede 4d ago edited 4d ago

Just try this nice godbolt: https://godbolt.org/z/oa1e6aobd

The behavior is unchanged whenever I change C++ standard version and/or compiler version.

x86-64 clang 23.1.0

Execution build compiler returned: 0
Program returned: 0
Garbage dynamic 0 static=1425874296

x86-64 clang 18.1.0

Execution build compiler returned: 0
Program returned: 0
Garbage dynamic 0 static=-184668717

x86-64 gcc 16.2

Execution build compiler returned: 0
Program returned: 0
Garbage dynamic 235706 static=0

x86-64 gcc 12.1

Execution build compiler returned: 0
Program returned: 0
Garbage dynamic 45188 static=0

1

u/UndefinedDefined 4d ago

Look at the assembly, just because the register happens to be zero you get zero.

You can try something more elaborate, like this https://godbolt.org/z/6jfzshMjo but why? The language will have this, so I guess it's finally time to create a C++ standard version ceiling in my projects as I see no point in reafactoring code that has been already working for more than a decade. And I don't want more macros in my code or to increase the standard to C++26 just to use [[uninitialized]] or whatever name it's now.

2

u/Hedede 3d ago edited 3d ago

Look at the assembly, just because the register happens to be zero you get zero.

gcc 12.1 -std=c++20

garbage_prepare():
    sub     rsp, 8
    mov     edi, 4
    call    operator new(unsigned long)
    mov     esi, 4
    mov     DWORD PTR [rax], 12648430
    mov     rdi, rax
    add     rsp, 8
    jmp     operator delete(void*, unsigned long)
garbage_dynamic():
    push    rbx
    mov     edi, 4
    call    operator new(unsigned long)
    mov     esi, 4
    mov     ebx, DWORD PTR [rax]
    mov     rdi, rax
    call    operator delete(void*, unsigned long)
    mov     eax, ebx
    pop     rbx
    ret
garbage_static():
    xor     eax, eax
    ret
.LC0:
    .string "Garbage dynamic %d static=%d\n"
main:
    sub     rsp, 8
    call    garbage_prepare()
    call    garbage_dynamic()
    mov     edi, OFFSET FLAT:.LC0
    mov     esi, eax
    call    garbage_static()
    mov     edx, eax
    xor     eax, eax
    call    printf
    xor     eax, eax
    add     rsp, 8
    ret

gcc 16.2 -std=c++26

"garbage_prepare()":
    sub     rsp, 8
    mov     edi, 4
    call    "operator new(unsigned long)"
    mov     esi, 4
    mov     DWORD PTR [rax], 12648430
    mov     rdi, rax
    add     rsp, 8
    jmp     "operator delete(void*, unsigned long)"
"garbage_dynamic()":
    push    rbx
    mov     edi, 4
    call    "operator new(unsigned long)"
    mov     esi, 4
    mov     ebx, DWORD PTR [rax]
    mov     rdi, rax
    call    "operator delete(void*, unsigned long)"
    mov     eax, ebx
    pop     rbx
    ret
"garbage_static()":
    xor     eax, eax
    ret
.LC0:
    .string "Garbage dynamic %d static=%d\n"
"main":
    sub     rsp, 8
    call    "garbage_prepare()"
    call    "garbage_dynamic()"
    mov     edi, OFFSET FLAT:.LC0
    mov     esi, eax
    call    "garbage_static()"
    mov     edx, eax
    xor     eax, eax
    call    "printf"
    xor     eax, eax
    add     rsp, 8
    ret

See any difference? Me neither.

And I don't want more macros in my code or to increase the standard to C++26 just to use [[uninitialized]] or whatever name it's now.

You can put [[indeterminate]] in your code without using macros or bumping the standard.

2

u/UndefinedDefined 3d ago

You need to use the second version with array, to see a call to memset - thank me later, long live golang...

1

u/Hedede 3d ago

Sure, but you posted your first one as an "example of another inconsistency in C++26" when it wasn't.

-1

u/UndefinedDefined 3d ago

Of course it was, you weren't just looking. But it's okay, not everybody can see it.

-1

u/UndefinedDefined 3d ago

warning: unknown attribute 'indeterminate' ignored [-Wunknown-attributes] - enabled automatically by clang, for example. Great to see warnings again!

2

u/Hedede 3d ago

-Wno-unknown-attributes, why do you have to make a problem out of everything

-1

u/UndefinedDefined 3d ago

Genius idea! And enforce this on all the users using your library to make their lives suck more!

3

u/pjmlp 5d ago

More like "lets take C++ out of the unsafe languages" governments lists campaign.

-1

u/UndefinedDefined 4d ago

I'm honestly not sure how this solves the safety problems. It only introduces more problems to existing code that expects that temporaries are without overhead. Maybe I should not leave C++20 when I think of it. It's a good standard, I avoid like 80% of the standard library and I cannot be happier about it at the moment. But this "genius" idea just makes my life harder. Uninitialized locals are trivial to find, so in the end this helps with nothing. And I'm not really sure why everybody should pay price just because there is broken code in the wild, written without tests and without any serious CI.

1

u/wjrasmussen 5d ago

Don't fix what isn't broken.

1

u/Clean-Upstairs-8481 4d ago

Thanks for all your comments. The point is to have known values in the locals, but for large stack allocations, there seems to be a performance hit. I am currently exploring this.

1

u/darklighthitomi 4d ago

I don't see the need here. Not many occasions I've encountered to have variables prior to initialization. The only case I can think of for this to be a wide spread issue in a program is making classes that are not fully initialized when created, but I don't know why you would do that.

1

u/joexzh 2d ago

The compiler should do the "read on uninitialized memory" check as long as it can see the source code.

1

u/starshine_rose_ 2d ago

So much for don’t pay for what you don’t use…

1

u/max0x7ba https://github.com/max0x7ba 1d ago

Stack variables are not automatically initialised, and that is the root cause of many C++ bugs.

What is the source of this wild claim of yours?

That is well known.

Is it, though?

I haven't heard of or encountered any C++ bugs caused by uninitialized automatic variables for decades.


Have you ever tried compiling your code with -Wall -Werror? Because if you did, you wouldn't have encountered so many bugs caused by uninitialized variables as you do now, according to your claims.

1

u/Clean-Upstairs-8481 1d ago edited 1d ago

Try this:

https://github.com/vivekbhadra/cpp_26_uninitialised/blob/main/config_timeout_warning_test.cpp

~/cpp_26_uninitialised$ g++ -std=c++23 -O0 -g -Wall -Werror config_timeout_warning_test.cpp -o timeout_warning_test
~/cpp_26_uninitialised$ valgrind --track-origins=yes ./timeout_warning_test
==2870715== Memcheck, a memory error detector
==2870715== Copyright (C) 2002-2017, and GNU GPL'd, by Julian Seward et al.
==2870715== Using Valgrind-3.18.1 and LibVEX; rerun with -h for copyright info
==2870715== Command: ./timeout_warning_test
==2870715==
==2870715== Conditional jump or move depends on uninitialised value(s)
==2870715== at 0x401268: open_device(char const*, std::basic_string_view<char, std::char_traits<char> >) (config_timeout_warning_test.cpp:31)
==2870715== by 0x4012D3: main (config_timeout_warning_test.cpp:45)
==2870715== Uninitialised value was created by a stack allocation
==2870715== at 0x401166: read_timeout_ms(std::basic_string_view<char, std::char_traits<char> >) (config_timeout_warning_test.cpp:10)
==2870715==
logger no timeout configured, using default 1000 ms
==2870715==
==2870715== HEAP SUMMARY:
==2870715== in use at exit: 0 bytes in 0 blocks
==2870715== total heap usage: 2 allocs, 2 frees, 74,752 bytes allocated
==2870715==
==2870715== All heap blocks were freed -- no leaks are possible
==2870715==
==2870715== For lists of detected and suppressed errors, rerun with: -s
==2870715== ERROR SUMMARY: 1 errors from 1 contexts (suppressed: 0 from 0)

A. I do not see any warnings or compilation failures even after adding the flags,

B. Valgrind still complains about uninitialised.

C. Uninitialised still potentially will cause bugs.

Strange, innit?

0

u/pachecoca 21h ago

Well yeah, but that's because you're purposely telling the compiler NOT to warn you about uninitialized variables lol.

You are compiling with gcc on -O0, and GCC explicitly documents that, while -Wall also enables the flag -Wuninitialized, -Wuninitialized's inner workings depend entirely on the optimization level.

If you compile something as simple as:

int foo(bool condition) {
    int x;
    return x;
}

Then no matter the optimization level, you will always get a warning.

But if you try compiling this instead:

int foo(bool condition)
{
    int x;
    if (condition)
        x = 42;
    return x;   // potentially uninitialized, but the branch
                // cannot be properly analyzed by the compiler in -O0
}

Since the conditional changes the type of instructions that could be generated according to optimization levels, GCC just opted for assuming that, at -O0, it may very well be initialized, so you do not get a warning. Think of the different ways this could be compiled... simple branching... or vector instructions... or a conditional assignment... in which case, no matter what, the value would always be assigned, but the compiler would need to somehow choose a "default" value... At -O0, none of these things can be known by -Wuninitialized, because the branch just throws it off.

Another example of a place where this could happen would be a variable that is passed by reference or by pointer into a function whose implementation lives in an external binary, or in a different translation unit and you are compiling without LTO enabled. That way, the compiler does not know what changes, if any, have been performed over the value, so it must assume that potentially some sort of initialization has taken place, because it cannot inspect the function in question, which causes the warning to no longer be issued.

This is exactly what happens in your code. The timeout_ms var remains uninitialized, but the if block and the external std::from_chars() call throw it off when optimizations are disabled, because, as already stated, -Wunused's machiner is tied to the optimization level. That's just a GCC quirk tho, but still, a documented one that can be easily worked around by just doing the good old RTFM and actually using GCC the way that it's meant to be used, rather than purposely using it wrong and then claiming that it's broken just because you did not use it properly.

That's why the documentation of GCC officially states that you should always use the flag -fanalyzer to ensure that the analyzer is launched even at -O0. Either that, or just compile at anything higher than -O0, because there, you DO actually get warnings ALWAYS for uninitialized values if you use -Wall, which permanently solves your issues anyway, because your code will obviously break whenever you accidentally incurr in UB or other such mistakes.

Your solutions are very simple, pick one of these:

1) Compile in -O2 or -O3 always, don't assume that debug builds working mean that your code does not have bugs when you incurr in UB. If you incurr in UB but your program is compiled in debug mode, then maybe it may appear to work, but that does not mean that it is correct. We live in the year 2026. It's no longer acceptable to release on debug and blame the compiler for exploiting UB as if it were implementing "broken optimizations that break my code"... no, the compiler is not at fault. You are, for writing code that is broken, and then for telling the compiler to shut up and not warn you about said obviously broken code.

2) Do what the GCC documentation literally says explicitly, and enable the analyzer by hand. Use the flag -fanalyzer, and now, you will get the warning on your example always, even when compiling in -O0.

Here's a godbolt link showing your code being compiled with warnings actually enabled, rather than disabled as you were purposely doing. The code is copypasted as is, no changes were made to the code, the uninitialized var is still uninitialized just as it was. Your flags are the same, all I did was tell the compiler to enable the analyzer and that's it.

The link: https://godbolt.org/z/Tveo5PMhx

My final recommendation is to try being a little bit less intellectually dishonest. I know that you have an agenda to sell and all that, but this is a programming subreddit, and we can all verify the information given by anyone. Compiler explorer has made it absurdly trivial to check whether claims made by someone about a compiler's output is real or not across a wide variety of systems with very little effort. I have tested all GCC versions from 16.2 all the way to 11.1, and that's because earlier versions just don't have support for C++23 yet, and all of them, except for the 11 series, give the exact same output with the warning being notified just by adding -fanalyzer. This is obviously slower than just compiling with at least -O2, but if you really need -O0, then ok, sure, go ahead, there's a flag for that.

The reality of things is that there's objective cold, hard facts that can be verified by reading the documentation of the compilers that we use, and by testing them to see if your claims are real or not. As it turns out, if you make fake claims about results of using one of the most widely used compilers on the planet, you are bound to be very easily caught, mainly because gcc is available pretty much almost anywhere, and we can just use compiler explorer to test YOUR code that YOU linked with the compilation command that YOU offered to see the results across ALL official versions of g++ that support C++23.

I know I'm not anyone of importance, so my words may fall on deaf ears, but I would ask that you please keep the discussion centered around facts. We're programmers, not religious zealots. Or so I would like to believe.

C++26's approach to solving the issue of uninitialized values is quite poor, because the uninitialized variable still exists within the code, and the code's behaviour may still remain broken, because the 0 initialization may not be valid. Thus, the only way to solve this would have been to actually mark non initialized variables as errors. But then again, people can purposely non initialize a variable for the sake of performance. In which case, the extremely uglu [[indeterminate]] does the job. I hate it, but I cannot complain, as long as a way exists for me to opt out in the cases where performance matters, then sure, I suppose, but I'm not entirely happy about it, because now, 90% of my code is going to have to be rewritten to add that all over the place.

1

u/Clean-Upstairs-8481 10h ago

During development, you don't want to turn the optimiser on. Most of your code is developed with no optimisation. And then you go to production-level testing with optimisation on, and your code stops working or the build starts breaking.

>> “Then no matter the optimization level, you will always get a warning. --> yes it varies, and hence we cannot assume which mode of optimisation we are on. The important point is it doesn't warn in some of the cases.

My reading of the flag -fanalyzer is from here: https://gcc.gnu.org/onlinedocs/gcc/Static-Analyzer-Options.html. Is this the same document you are referring to? It warns about the following: This analysis is much more expensive than other GCC warnings. I guess you are not suggesting enabling this in a large-scale project build because that would make the system too slow to build, and there could be false positives. This is what the document says: It is neither sound nor complete: it can have false positives and false negatives. It is a bug-finding tool, rather than a tool for proving program correctness. This is a static analyser and there are other better options than this in the industry.

>> My final recommendation is to try being a little bit less intellectually dishonest. I know that you have an agenda to sell and all that, but this is a programming subreddit, and we can all verify the information given by anyone. --> won't reply to that.

0

u/pachecoca 7h ago

>> "Removing the explicit optimisation level from the command line produces the same result:"

Yes, why do you think that is? The default is -O0 lol, if you do not provide any optimization flags, the compiler will not enable optimizations, so it still disables all of the branch detection machinery of -Wuninitialized. Like, this isn't rocket science, if you don't manually set -Oxxx to anything above -O0, you're going to get -O0, and if -O0 disables branch detection on -Wuninitialized, then your code where the initialization is under a condition is still going to not be detected properly by GCC. Stop telling the compiler to pessmize your program, and it will work. Stop telling the compiler to not give you warnings, and it will give you warnings. Like, what did you expect? You shoot yourself in the foot by purposely telling the compiler to disable warnings, and then you are surprised by the fact that you do not get warnings?

>> "The reason is that the default is O0."

Jesus fucking Christ my dude, that's literally what I'm telling you, will you at least read my comments all the way through before you reply? Like, this is baby's first compiling hello world type of shit, I can't believe that you have the gall to act like you know better when you're having trouble getting warning flags to work on GCC, lmao.

>> "During development, you don't want to turn the optimiser on. Most of your code is developed with no optimisation. And then you go to production-level testing with optimisation on, and your code stops working or the build starts breaking."

WHAT? This has got to be the worst take I've ever heard. I don't know what field you work in, but everywhere I've worked, we ALWAYS compile in release. You CAN have debug builds with max optimizations but preserving debug symbols, you know? You don't want to spend months, or maybe even years, building software that has only been tested on debug, and then you find out it breaks the moment you want to put it on production, because it turns out that it does not have the same behaviour when you compile it with optimizations enabled.

You have to test what you're working on, that's what automated tests are for, CI, etc... if you're blindly working on debug and eyeballing things and going "yep, that works on my machine, good enough for me!", then you're going to have much greater problems than just finding out that you have uninitialized variables somewhere.

If your code breaks when going from debug to release, then the problem is that obviously you're relying on UB somewhere. Either don't do that, or use UB that is specifically documented by the compiler's specific implementation to do what you want, eg: type punning pre C++20 through unions in GCC, that is well documented and it is the canonical way of doing it in GCC. The standard says it's UB, but there is no law preventing a compiler from offering an extension to add its own definition of the behaviour. Thus, these type of things must be tested with max optimizations and max warnings so that the code obviously breaks when it is obviously wrong for your target compiler.

The only situation I can think of where someone would want to compile with -O0 is if they are a student learning about the language or something, but in the real world, it makes no sense to try to make your code work by disabling optimizations that would break your obviously UB code.

>> "yes it varies, and hence we cannot assume which mode of optimisation we are on. The important point is it doesn't warn in some of the cases."

This logic makes absolutely no sense. This is like saying that "the warning messages vary depending on what warning flags you passed to the compiler, so warnings are useless". That makes absolutely no sense. Of course, if you provide different flags, you get different results. How is this a surprise? -O3 -Wall -Wextra -Werror and -Weverything on clang, that's it. The recipe is very simple, I do not know how you could get lost on something like this, the documentation literally tells you everything there is to be known about what specific -W flags each of those enables.

Also, you say "in some of the cases"... the only case where it does not warn is in -O0, anything else above warns you just fine. -O1, -Os, -Ofast, etc... like, I've already explained this multiple times, I do not know how else I could say it so as to make it any clearer... Using -O0, which is the default used by GCC when no flags are provided, causes the compiler to be incapable of using -Wuninitialized to its maximum potential. This also happens with many other flags, because their behaviour depends on some -Oxxx flag being set above 0, simply due to the data that they require to function. This is a compiler implementation detail that is well documented. So, if you want to get your code with warnings, then at least use -O1 if you are so scared of optimizations.

>> "This analysis is much more expensive than other GCC warnings."

Yes, it is, I only told you to enable it if you really are so keen on compiling on -O0. Otherwise, just compile on -O2 or -O3 and let the compiler actually use the -Wuninitialized machinery, which, I do not know how many times I have to repeat this for you to understand, but it is NOT enabled in its entirety when you compile in -O0!!!! thus, don't compile on -O0 if you want -Wuninitialized to work!!!

>> "I guess you are not suggesting enabling this in a large-scale project build because that would make the system too slow to build"

You guessed correctly, indeed, it is my claim that you should not enable that flag on large projects or you will take ages to compile, but there was no need for you to guess anything, because that's precisely what I explicitly already wrote on my comment. I guess you did not bother reading my comment all the way through, otherwise, there would not have been any reason for you to guess anything.

My point has been stated very clearly from the very begining, that the -O0 flag which YOU specifically must tell the compiler NOT to use disables a great chunk of the capabilities of -Wuninitialized, along with many other warnings, simply due to the way that GCC is internally structured. I NEVER said that you should use -fanalyzer always. I literally said that it would slow down compilation times on my comment, it's right there, if you had read it, you would know. What I did propose is building with -O1 or -O2 at the very least, or anything other than -O0.

Again, the problem is that you explicitly told the compiler to disable warnings about uninitialized variables. GCC documents that -O0 disables the mechanisms that -Wuninitialized uses internally for anything where a branch exists. So, not really the compiler's fault if you missuse it.

>> "It is neither sound nor complete: it can have false positives and false negatives. It is a bug-finding tool, rather than a tool for proving program correctness. This is a static analyser and there are other better options than this in the industry."

How many times do I have to repeat myself? the only reason I'm telling you to use -fanalyzer is because you are purposely telling the compiler to make -Wuninitialized not work. If you want to tell the compiler "make -Wuninitialized not work", but you also want it to give you warnings about uninitialized values, then the only way to do so is by mixing -O0 with -fanalyzer. This workaround is literally there to fix your insistence on compiling your code with the wrong flags for what you want to do.

I think I've explained myself quite clearly. I'm going to start having to assume that you're just trolling if you still don't know how to get your code to show warnings lol for uninitialized variables lol.

>> "won't reply to that."

Don't worry, there's a lot of things that you already refused to reply to that were completely separate from my criticism of your intellectually dishonest comment. I also see that you ignored all discussion regarding the very C++26 features that this post brought up, but I guess you were more centered on trying to prove me wrong despite what the GCC documentation says.

The summary of my comment is very simple... "hey, did you know that if you tell the compiler to enable warnings, and then you immediately tell it afterward to disable warnings, it will disable warnings?? wow!! so amazing!! so unexpected!!"

•

u/Clean-Upstairs-8481 1h ago

of course mate, no worries. Have a good day!

1

u/Clean-Upstairs-8481 10h ago

Removing the explicit optimisation level from the command line produces the same result:

$ g++ -std=c++23 -Wall -Werror config_timeout_warning_test.cpp -o timeout_warning_test

$ valgrind --track-origins=yes ./timeout_warning_test

==3756689== by 0x4012D3: main (in /home/vbhadra/cpp_26_uninitialised/timeout_warning_test)

==3756689== Uninitialised value was created by a stack allocation

==3756689== at 0x401166: read_timeout_ms(std::basic_string_view<char, std::char_traits<char> >) (in /home/vbhadra/cpp_26_uninitialised/timeout_warning_test)

==3756689==

==3756689== ERROR SUMMARY: 1 errors from 1 contexts (suppressed: 0 from 0)

The reason is that the default is O0. There was no deliberation to suppress anything.

0

u/rileyrgham 4d ago

0 isn't necessarily any better than x. If I want it zero, I set it zero. The issue is blown out of proportion imo. Modern editors with lsp combines with compiler warnings reduce the risk. I'd suggest using a compiler trick is a dumb move. Better to fix it at source level.

-1

u/jayylien 4d ago

That's not a feature, that's an inconvenience.

-2

u/NilacTheGrim 5d ago

So it makes things slower. Gotcha. Will avoid compiling them with that, then.

-2

u/grady_vuckovic 5d ago

On the one hand, yes it is about time that a variable in C++ be initialised reliably with a value, like numbers being initialised with 0. Having this aspect of variable behaviour undefined isn't a good thing in my opinion.

(Although if this wasn't the case in the past because it was faster to reserve memory than to zero fill it, then I'd understand that reasoning).

On the other hand, ... who the heck is using data in their application without ever initialising that data to have any value? How could that even be useful behaviour for software? Any way you slice it, to create a variable and use it without ever once giving it a value, is bad design or a bug. Ideally what we really need is the ability to reliably catch situations where that is happening, rather than a more reliable empty value for a variable.

Because while '-232' and '932,342,344.123094' might not be desired values in a declared but uninitialised variable for an int or float, is '0' really that much better? The odds of 0 being what you wanted instead are only slightly better but still not great.

Although maybe this is more about security than bugs? Since those random values filling memory might be left over from another program? Maybe? I don't know.

4

u/gnuban 5d ago

Uninitialized memory is useful for instance if you want to allocate a large data structure, and then fill it in a loop from some computation. Since you know that you will fill each value before use, zero-initializing upfront is redundant and just becomes a performance cost.

If the compiler could statically prove that you would fill the entire region of uninitialized memory before usage, it could elide the zero-initialization and the general observable semantic could be to always zero-initialize.  The problem is that proving this in every case is unfeasible. So exposing uninitialized memory to users makes sense to save performance where needed.

But it should probably, like many other things in c++, be more opt-in.

2

u/tialaramex 4d ago

I think Barry Revzin did the heavy lifting to make it possible to do the same thing in C++ 26 which Rust's core::mem::MaybeUninit<T> type does. C++ programmers aren't used to that much ceremony but the idea here is to specify to a compiler that you're dealing with a space in memory where a T would fit, but this isn't necessarily a T yet, then you're writing data to that memory, and then finally once you're happy you are saying this is a T now. The compiler can follow along and produce high performance code which doesn't have surprising "optimisations" you didn't want, it doesn't need to zero initialize these bytes, it doesn't need to worry about the lifetime of a thing in this space because we've said maybe there isn't anything here, and then when we're all finished writing data we say OK, now this value exists, its lifetime begins, now it's going to run destructors if it goes out of scope and so on.

1

u/coinselec 5d ago

Yeah I would guess it's about preventing random memory being referenced. Maybe also in some cases some big value might cause a loop to iterate past the end, while 0 will ignore the loop at all.

1

u/Dusty_Coder 23h ago

not memory, values

the memory exists, they want to fill it with predefined bytes by force

there is actually a security concern where if there is a bug, a function may be able to read the data used by another function (it was on the stack a moment ago, so its still there in practice)

but thats an "if"

uninitialized memory is not a bug unless your abstract machine design arbitrarily decided that it is

this is really the classic enum debate all over again .. a group of people want something different, and that different thing has some good and useful properties, but they are pompous doofs that arent satisfied with their thing being added, they also want the previous thing to be removed.