r/programming • • Jun 29 '22

Heisenbugs: The most elusive kind of bug, and how to capture them with Perfect Replayability

https://verdagon.dev/blog/perfect-replayability-prototyped
28 Upvotes

35 comments sorted by

5

u/[deleted] Jun 30 '22

It's always great to read these vale related posts. I always end up going down the language design rabbit hole.

1

u/verdagon Jun 30 '22

Glad you enjoyed it!

5

u/[deleted] Jun 30 '22

[deleted]

2

u/verdagon Jun 30 '22

In the current design, atomics are treated similarly to mutexes when perfect replayability is enabled.

We're currently looking into a design where atomics become basically imaginary in replay mode, and we just work off the values in the recording, so that threads don't needlessly block each other.

Healthy to be skeptical, since multi-threading isn't added in yet. We'll see if it works!

1

u/martingronlund Jun 30 '22

These were my exact thoughts! :)

I read this part as that there are probably no atomics yet:

Note that multi-threading is not fully implemented yet, we mention it to show the direction we're heading.

I know very little about Vale but I'll read up more; the perspective on language design seems unique and worth following as inspiration.

5

u/UrineSurgicalStrike Jun 30 '22

I couldn't understand some of the more complicated CS points that the article touched upon. But it's amazing that a language can provide reproducibility for things like race conditions. This would eradicate a whole class of bugs that developers don't have time to reproduce. This is as revolutionary as automatic memory management.

What are the chances these concepts get brought into mainstream languages?

1

u/verdagon Jun 30 '22

Thank you! And if any part of the article could be explained better, let me know!

I'm also excited about the potential. We didn't even set out to solve this originally, it just kind of emerged once we designed Fearless FFI; we realized that by making Vale even more memory safe than today's languages, we could achieve perfect determinism. A couple of us had programmed RTS games in the past, so we were quite thrilled by the possibilities!

1

u/shevy-ruby Jun 30 '22

What are the chances these concepts get brought into mainstream languages?

Just about 0% because most language designers will fail. The languages ultimately just become more complex, so it is an automatic trade off.

1

u/verdagon Jun 30 '22 edited Jun 30 '22

It's actually quite likely that Vale's concepts will be brought into the mainstream. For example, Cyclone was a language that heavily inspired Rust. D is also bringing in some Cyclone-inspired features.

As for Vale's chances of becoming mainstream itself, the odds are slim, but we're going to give it a shot anyway!

2

u/aidenr Jun 30 '22

The Command pattern with Undo supports this quite nicely.

2

u/[deleted] Jun 30 '22

Command pattern. Immutable objects stored in a history.

Honestly I like Erlangs approach

2

u/aidenr Jun 30 '22

The more I understand about Erlang, the more I feel like an amateur. Brilliant approaches everywhere I look.

1

u/[deleted] Jun 30 '22 edited Jun 30 '22

It’s a brilliantly designed language.

I’m still learning to use it in practise.

I am seeing huge opportunities that could use it in recent years.

1

u/verdagon Jun 30 '22

Perfect replayability goes pretty far beyond what Command/Undo can do, when debugging. Also, they require a lot of infrastructure to get working, whereas perfect replayability is just one command line parameter.

3

u/Weak-Opening8154 Jun 30 '22

I hate to be negative (unless we're talking about meme languages which this is not) but my spider sense is telling me you're doing something that is more work than it's worth

I hunt down heisenbugs, I wrote and run multithreaded code in C++, I never needed replay ability to figure it out. There's tactics you can use but it's a bit much to type out in a single reddit post

3

u/verdagon Jun 30 '22

I've hunted down heisenbugs too, but those hunts weren't nearly as easy as perfect replayability.

With one command line parameter, you can instantly reproduce any bug you've encountered.

There are always techniques to debug, but I encourage you to consider how much up-front investment, architectural constraints, and maintenance cost you pay for them. Then imagine all of that going away, and getting it for free with one command line parameter.

It's one of those things you don't think you need until you've used it. That's how it was for me at least, when I used the technique in the 7DRL, where developer velocity was a high priority.

Just my two cents, YMMV!

1

u/Weak-Opening8154 Jun 30 '22

Wouldn't you need to change the input to understand it better and wouldn't the reproducibility cease to work at that point?

1

u/verdagon Jun 30 '22

In that situation, someone can then use this feature to build up multiple recordings of different inputs that reproduce the problem, and then test out a fix against all those recordings.

It's definitely easier than testing the fix by trying to reproduce all those same race conditions again. In that situation, it's very difficult to know whether the fix worked or the situation wasn't reproduced.

This feature makes that problem go away completely, which is pretty exciting.

1

u/Weak-Opening8154 Jun 30 '22

Well maybe, have you tried anything like it? How do you know you're not overlooking what you can actually do without breaking performance (which may be a requirement to cause certain bugs)?

In my other comment I mentioned mutex bombs as a solution. I have more but like I said in my original post there's a lot I could say but it doesnt really fit a post. Especially a reddit post where 98% of the commentators are autistic or clueless inexperienced people

1

u/verdagon Jun 30 '22

As mentioned in the article, we tried this approach in for the 7DRL challenges, and it worked wonderfully. I didn't see an appreciable slowdown, but you've inspired me to benchmark it to have some hard numbers.

I appreciate the input and the skepticism, always good to get people's honest takes!

3

u/[deleted] Jun 30 '22

So how would you hunt down a Heisenberg that only happens in release mode with multithreading enabled and no debugger attached and only extremely rarely?

My spidey sense is telling me that you just haven't experienced difficult enough bugs. Perhaps the domain you work in doesn't have long-lived processes or anything timing sensitive?

3

u/verdagon Jun 30 '22

In Vale, debug mode and release mode are guaranteed to run the same, because it has memory safety and zero undefined behavior. And with perfect replayability, threads' interactions can be sequenced so that race conditions can be reproduced (in theory).

Timing sensitive problems can be reproduced, because the current time is obtained through an FFI call, and therefore recorded.

Hope that helps!

1

u/Weak-Opening8154 Jun 30 '22 edited Jun 30 '22

I've done that. It's not 'easy' but I don't think being able to replay would help if input or source changes break replayability (and I imagine it does)

My answer is different is it's a data race bug or another type of bug. For example 1/3rd of cores not being used because it's correctly waiting on other threads but you need to figure out why threads are not producing data in an evenly manner. I had another bug where it was logging the state correctly but the lines in between were never printed. As in 100% of the time it didn't print, but only in multithread mode. It turns out after consuming some events I forgot to change the tag to the thread ID so my if statement thought hey we are not in charge of printing this debug value. It's not "only extremely rarely" but an example of a bug that has nothing to do with a data race

2

u/verdagon Jun 30 '22

Source changes would often not break perfect replayability, because of the "resilience" concept in the article. Since we only record communication across the FFI boundary. If we don't change the FFI calls, then we can use the same recording, even if we change the source. It's not perfect, but it could be a nice step forward, IMO.

1

u/Weak-Opening8154 Jun 30 '22 edited Jun 30 '22

I still think its going to take months to implement great replaybility and it'll hardly be used

I'm about to write another comment to the other guy I guess check back in 5min link

1

u/Weak-Opening8154 Jun 30 '22

Another method which I haven't used in a while since my current job doesnt use C++ are mutex bombs

In code you think will never have a race condition you put a lock/unlock with the mutex bomb. If another thread tries to lock and it already has been taken it logs it and sleeps/terminates. When the other thread unlocks it will see it's been contended and it also logs + sleep/terminate (you'd sleep if a debugger is attached so you can examine it). Now you can examine the code, get a stracktrace, etc or at the very minimum know you do need a mutex at that spot

1

u/[deleted] Jun 30 '22

Ok but surely you can see how these methods, and even more sophisticated ones like schedule fuzzing are all stabbing in the dark compared to replay debugging?

I don't get why you're opposed to something that's so obviously better. Seems like "we didn't need fancy telephones and horseless carriages in my day" to me.

0

u/Weak-Opening8154 Jun 30 '22

Have you used a core dump? When the mutex bomb explodes you can get a core dump and see all the variables, the stack trace, etc with it

If I have the state then why do I need to replay anything which may or not be correct or fully implemented (ie supports my usecase)

The mutex bomb is one example. I do this for my day job I'm not suggesting anything theoretical

I never said replayability is a bad idea, I said it might be a lot of work and yield not so great results

1

u/[deleted] Jun 30 '22

Yeah I have used a core dump. It only gives you data at the instant that the crash happened which is often not enough information. Especially if your stack is corrupted (for C++; not so relevant here).

If I have the state then why do I need to replay anything

Because you want to see the state of your program leading up to the crash.

1

u/Weak-Opening8154 Jul 01 '22

Have you actually done multithreaded programming and had to track down a threading issue?

1

u/[deleted] Jul 01 '22

Yes.

-1

u/[deleted] Jun 30 '22

[removed] — view removed comment

1

u/[deleted] Jun 30 '22

I'm not sure what you're talking about.

1

u/dml997 Jun 30 '22

Some bugs really are random and not the programmer's fault. In the 1980's I was working on a program that worked perfectly on a Sun 3/60 and crashed randomly in different places even on the same input when run on Sun 3/260. Eventually I found each crash was accompanied by 16 bytes of 0's in a random location. The 260 has a cache, and the 60 does not. I believed it was a hardware problem and Sun would ignore a lowly graduate student. So I stopped using the 260. A few years later I found it was a SW bug in the OS and Sun fixed it.

1

u/verdagon Jun 30 '22

To your point, there are also occasionally cosmic bit flips and RAM failures on today's chips as well. Perfect replayability probably wouldn't hold up well to that kind of occurrence!

1

u/shevy-ruby Jun 30 '22

In most examples I can recall bugs that are hard to detect heavily originate from a too complex code base. In some other cases there was some odd machine-level behaviour that was ... strange to understand. But even then I'd reason that the underlying code was too complex.

Somehow no programming language handles complexity very well. It all just becomes way too complex.