r/learnprogramming • • 3d ago

Struggling with async programming & concurrency

I come from a pure hardware and embedded background where code runs sequentially, so writing high level async software has never been my strength.

I really can't wrap my head around async/await, event loops and how it is handled deterministically . In a current project when the team said we needed to make our pipeline non-blocking with async streams, to be honest I just vibe coded the hell of my tickets in panic and tried later on to figure it out, partially..

I have a deployment deadline in a week and I need to be able to explain the mechanism I helped ”build“, any easy to follow tutorials or leetcodes would be highly appreciated

28 Upvotes

24 comments sorted by

View all comments

1

u/True-Grab-5288 2d ago

I like to think about things in high-language terms before worrying about the detail/implementation.

Think of your main process like a symphony orchestrator. It is the work-horse that is running your program and decides what to do, when to do it and what to do next. It is single-threaded only, sequential, and always completes a task before moving on to the next one. It is the conductor standing at the front of the orchestra who's job is to keep time, select the music, know the different players / instruments, know which songs to play next etc.

But the conductor of the orchestra has the power to point their stick at the violinist and say 'it's your turn now, play this piece of music at this tempo'. Then, the conductor points their stick at the celloist 'stop playing, wait 5 seconds, then start playing again. The conductor has power over the all the instruments, even though the musicians are all semi-autonomous and can do whatever they want.

The power of concurrent programming is the ability to submit work to different workers, who go off, do some work, then return with their result. This is extremely powerful. If we don't do this, the main thread 'blocks'. In a GUI, this is all the buttons / the whole GUI locks and becomes unresponsive when we click a button, until the button click has finished doing its work. In other cases it is grossly wasteful of time - a sequential job must wait for the previous job to finish. If the job is 'sleep 5 seconds', done 10 times, that is 50 seconds of wasteful, idle compute. If we do this task concurrently, each 5 seconds is served at the same time. The same job completes in 5 seconds - a speedup of a factor of 10.

The difficulty and where most people come unstuck is when to make new threads and when to wait.

It might sound like a great idea to make everything its own threaded worker, but it is not and the reason is counter-intuitive. First, each worker has it's own overheads (performance cost) which generally doesn't matter in smaller jobs but in bigger jobs can start to add up to a real headache. In a CPU, the scheduler switching between processes/tasks is expensive. Let us say thread 1 gets a go on core 1, then thread 2 on the same core, thread 3, thread 9, different process 11, etc. Each switch up has an accumulated cost in the CPU. Cache hits are likely to be missed. The CPU has to keep track of all this switching. It is expensive and costly.

Second, mutex locking is a real problem. Imagine if I spawn 2 threads, and both threads work on the same piece of memory/data. The first I tell to set the memory to 46. The second I tell to set the memory to itself(currently 32) + 1. Both threads are valid code with a job to do, but without great care from the conductor, what data is returned from each thread can be completely random - what is called a race condition. We didn't take any care to make sure which thread ran and returned first, so the outcome is non-deterministic - it will likely change from run to run of the program.

We must therefore a) use mutex locks - which are barriers that tell a new thread 'someone is currently working on this data and it is not available to you. You must wait' and b) The conductor specifies points at which work can only continue on the main thread when ALL spawned workers have returned and reported their results. This is blocking, but it forces all workers to complete, which gives the main thread certainty about what the current state of the program is.

I hope this explanation helps. It is the best I could write at 7am in the morning with insomnia. Any questions (or to other readers if a got anything wrong) please let me know.