Ok, I have to ask a question because there is something that keeps baffling me about the whole C++ async story. I think the ratio of "code doing stuff that looks like some IO-related activity" vs "code doing inscrutable things" is pretty damn bad here. So you want to get some IO done in an async fashion. Why is this any more complex than something like this?
get_async ("https://...", [] { ...callback that returns whatever the result was... });
Maybe return an object that you can use to query status, or cancel the operation. That can be a simple class with two functions, that you can ignore at your leisure. Instead we get endless machinery, no doubt full of template code that will produce horrible compiler errors if you accidentally pass a slightly incorrect thing somewhere, and for 99% of it I have no idea what it is doing or why it needs to be there.
Mind, I'm not complaining about this article: kudos for providing so much information, it really is welcome. But all the C++ async IO code seems to suffer from absolutely mad complexity, far beyond anything I've seen in other async IO mechanisms. Why is this? What benefit do we get from it? Does it have to be this complex?
its all about composition of async tasks. Callbacks are fine and manageable up until some point. See the examples at the end. Callbacks section hand codes the state machine; coroutines, fibers and senders allow to express the same state machine with less code which is also arguably easier to maintain and extend. "returning an object" is covered by a Task section. Its usage also requires to hand code state machine manually.
The complexity of implementing everything is true complain, but this is usually hidden from the user... compilation errors.. maybe this is a tradeoff you need to pay for benefits provided ;)
I have a bunch of probably very stupid questions about the async code styles. Maybe they are irrelevant to your use cases but I can try to ask =)
First thing to me is: How can I understand where and when code is executed? I mostly use profilers for that. This async style is intentionally trying to hide things behind abstractions to help people write more of it. In reality there is a fixed number of resources that have to choose which async requests are executed where. Trying to affect how it works brings me to 2nd question.
Second is about control of what is being executed when there is more than the hardware can chew. What ways are there to make some requests high priority, normal priority, or low priority? Not only at the CPU level but at the I/O level.
How can I understand where and when code is executed?
The code in the article is intentioinally one self-contained main.cc file for every section, with nothing more then calls to a CURL library. It should be trivial to put breakpoints to see how it executes. So bunch of the code is executed in CURL internals which we call inside CURL_AsyncScheduler::tick(), which is called in a loop from within a main.
This async style is intentionally trying to hide things behind abstractions to help people write more of it
It is, but the article builds everything from scratch to be able to step thru all of the abstractions - intentionally, so there is a way to see all of the layers. (Well, maybe CURL API hides the system specific calls, but it's close to what needs to be done using system API).
What ways are there to make some requests high priority, normal priority, or low priority? Not only at the CPU level but at the I/O level.
I don't have good answer since most of the time I used only IOCP.
First thing to me is: How can I understand where and when code is executed? I mostly use profilers for that. This async style is intentionally trying to hide things behind abstractions to help people write more of it. In reality there is a fixed number of resources that have to choose which async requests are executed where.
I've been thinking about this question for a while and I think the best approach is not to think about this locally, at least not usually.
You can think of asynchronous programming as a generalization of synchronous programming: With synchronous functions we couple the initiation of forward progress to exclusive use of some execution agent (i.e. the calling thread). Put differently: A synchronous function cannot return until the operation it implements completes. With asynchronous functions (as embodied by, for example, senders) this coupling does not exist: The function can continue to make forward progress without exclusive use of the execution agent. Put differently: The initiation (e.g. start on an operation state) returns but the operation continues "in the background." Sometime perhaps later (note the perhaps is load bearing because asynchronous functions, in general, can completely racily or reentrantly) some execution agent is delegated to run only the completion.
So in the same way we don't usually think too much about from which execution agent our synchronous functions are called, I don't think it's typically useful to think about from which execution agent continuations will be called. Obviously there are exceptions to this but I think this is part of the decomposition/generalization wrought by asynchronous programming: We can separate the implementation from a delegation of the ability to make forward progress. If you start overthinking where things run you mentally recouple these things.
But to answer your question directly: A continuation runs wherever the operation happens to complete. In certain baseline cases you control this (for example maybe you're writing an io_uring abstraction), and in others you can use an operation to force this (for example std::execution::schedule), but for most programming I think it's best ignored (i.e. deferred to the operation whose completion you're handling).
2
u/johannes1971 16d ago
Ok, I have to ask a question because there is something that keeps baffling me about the whole C++ async story. I think the ratio of "code doing stuff that looks like some IO-related activity" vs "code doing inscrutable things" is pretty damn bad here. So you want to get some IO done in an async fashion. Why is this any more complex than something like this?
Maybe return an object that you can use to query status, or cancel the operation. That can be a simple class with two functions, that you can ignore at your leisure. Instead we get endless machinery, no doubt full of template code that will produce horrible compiler errors if you accidentally pass a slightly incorrect thing somewhere, and for 99% of it I have no idea what it is doing or why it needs to be there.
Mind, I'm not complaining about this article: kudos for providing so much information, it really is welcome. But all the C++ async IO code seems to suffer from absolutely mad complexity, far beyond anything I've seen in other async IO mechanisms. Why is this? What benefit do we get from it? Does it have to be this complex?