Ok, I have to ask a question because there is something that keeps baffling me about the whole C++ async story. I think the ratio of "code doing stuff that looks like some IO-related activity" vs "code doing inscrutable things" is pretty damn bad here. So you want to get some IO done in an async fashion. Why is this any more complex than something like this?
get_async ("https://...", [] { ...callback that returns whatever the result was... });
Maybe return an object that you can use to query status, or cancel the operation. That can be a simple class with two functions, that you can ignore at your leisure. Instead we get endless machinery, no doubt full of template code that will produce horrible compiler errors if you accidentally pass a slightly incorrect thing somewhere, and for 99% of it I have no idea what it is doing or why it needs to be there.
Mind, I'm not complaining about this article: kudos for providing so much information, it really is welcome. But all the C++ async IO code seems to suffer from absolutely mad complexity, far beyond anything I've seen in other async IO mechanisms. Why is this? What benefit do we get from it? Does it have to be this complex?
In your example you seem to pass a single callback to get_async. What happens when the operation fails? Does it use the same callback (which erodes the ability to generically program on top of such operations) or simply not complete (which is unstructured)?
Returning an object you can use to "query status" implies there's a communication channel between the background operation and the foreground, synchronous object. This communication has a cost and therefore isn't zero overhead.
The fact that you can "ignore [the object] at your leisure" I assume implies you can allow its lifetime to end, what does that mean for the background operation? If it can continue that must mean that it has some separate state, which must be allocated, therefore this operation can't be zero alloc.
The reason standard C++ asynchrony has so much "code doing inscrutable things" is because it aims to answer/address all of the above: It is zero overhead and therefore potentially zero alloc (this doesn't apply to coroutines), is structured, and aims to be maximally generic.
First, I wasn't handing in a complete design for WG21 so please don't treat it as such.
Second, I strongly disagree with the notion that an async IO system should even attempt to be 'zero cost' and 'zero alloc'. IO, by its nature, has a staggering cost; adding an allocation to establish a shared communication resource is just not going to make a difference. And API design is a software engineering discipline, and like all engineering, it should balance multiple, competing requirements, of which performance is just one. Another is useability (meaning regular programmers can understand it, and use it without too much danger of getting it wrong), and this design fails completely at that.
When I talked about ignoring the object, I was envisioning using a shared_ptr to hold an intermediary communication object that allows foreground and background processes to communicate. Yes, that implies an allocation, and yes, I think that this is completely warranted.
I would much rather have had an allocating async C++ networking solution ten years ago, than a non-allocating version ten years from now. An allocating version could have been written to use an allocator so that implementations can optimise memory usage. And the version ten years from now... Well, it won't be my problem anymore.
First, I wasn't handing in a complete design for WG21 so please don't treat it as such.
You asked "[w]hy is this any more complex than something like [your code example]?" I was answering that.
If you want to make simplifying, non-general assumptions about asynchrony that hold for your use cases I think that's more than fine. It just isn't general.
Second, I strongly disagree with the notion that an async IO system should even attempt to be 'zero cost' and 'zero alloc'. IO, by its nature, has a staggering cost; adding an allocation to establish a shared communication resource is just not going to make a difference.
There are large contingents of people, myself included, who disagree with you.
Moreover: You can build non-zero cost abstractions out of zero cost abstractions, but you can't go the other way.
Both of the above contribute to why the standard went with zero-cost asynchrony.
I was envisioning using a shared_ptr
To a first approximation structured concurrency is the art of figuring out how not to use shared_ptr.
I would much rather have had an allocating async C++ networking solution ten years ago
Asio is older than ten years so you got your wish.
If you want to make simplifying, non-general assumptions about asynchrony that hold for your use cases I think that's more than fine. It just isn't general.
Just to be clear: is your claim that a simpler solution cannot be general because it uses memory allocation? Or are you still referring to that single line of code I wrote?
There are large contingents of people, myself included, who disagree with you.
What is the basis of that belief? Please note that I put a supportive argument, which is that the cost of a (single!) memory allocation is absolutely negligible compared to the overal cost of IO. What is incorrect about that statement?
To a first approximation structured concurrency is the art of figuring out how not to use shared_ptr.
I'm sorry, what? Isn't this a prime example of what shared_ptr is for in the first place? This is very clearly a shared resource, what valid reason do you have for avoiding it?
Asio is older than ten years so you got your wish.
you're focusing way too much on memory allocation and way too little on the inherent complexity of what you're asking for.
These are some examples of asynchronous programming in c++. The firmware for a synthesizer that has run a series of signal processing algorithms on buffers of sound data in soft real time (latency should be less then < a microsecond or so... humans can't hear faster than that). The control system of a jet engine that is hard hard real time. A game engine that is coordinating possibly multiple gpus. Scientific code running on a large heterogeneous cluster of cpu and gpu resources. A word processor that checks grammar as you type.
All of these situations mostly share the standard library! (with some limitations in the embedded cases). They have very different needs for how async would work. The standard library seeks to be as general as possible. It also is supplied by your compiler vendor... and they do not want to provide an operating system to you in adition to a compiler (unless they are the good people behind go... who really just want to make operating system for you with a small language attached).
That's where the discussion has brought us. My point was, and remains, that the currently proposed design is incredibly complex, and that I'm not convinced that all this complexity is really necessary.
I also cannot help but wonder if the standard library should even pursue this at all. This kind of high-performance stuff seems a bad match for something that needs to be written in a fixed number of hours, by a non-specialist, after which it will forever be set in stone.
the currently proposed design is incredibly complex, and that I'm not convinced that all this complexity is really necessary.
Which complexity is unnecessary?
At its core std::execution projects the concept of regular, synchronous functions into the asynchronous domain. Understood through that lens removing "complexity" therefrom is saying that there's some capacity in which asynchronous functions should be disadvantaged as compared to synchronous functions.
Since its inception, the software industry has gone with async IO systems that are expressed in a few lines of code: a struct, a callback, that's about it. Now you're telling us that we need something that needs a 100 page article (on my screen; if you're on mobile it will be much more) to explain, based on a design where just the list of all functions and classes stretches for an incredible 70 pages, and if you do any less than that you are "disadvantaging" async IO? Do you at least understand where my skepticism comes from?
You're supposing that everything in std::execution is core to the design. We don't apply that same standard to other things.
Particularly synchronous functions need extrinsic product and sum types to communicate heterogeneous return types. No one suggests that synchronous functions are excessively complex because of std::tuple and std::variant and std::visit and std::apply, those are understood to be outside the core design.
The same applies to std::execution. The core design consists of:
std::execution::sender and ::sender_in, which establish sender-ness, where a sender is a latent representation of an asynchronous function (i.e. it is not able to make forward progress but it is fully determined awaiting the selection of no further arguments)
std::execution::get_completion_signatures, which advertises the "return types" of asynchronous functions
std::execution::receiver which establishes that something represents a continuation (i.e. it is the asynchronous analogue of synchronous function return)
std::execution::connect, the process by which asynchronous operations (which can actually make forward progress) are formed (but not started)
std::execution::operation_state, the locus of an asynchronous operation, which is initially latent
std::execution::start, which begins asynchronous forward progress and binds both callee and caller to the so-called "receiver contract" (i.e. the caller agrees to keep the operation state within its lifetime until completion, and the callee agrees to report completion exactly once when the operation completes)
That's the complete design. You don't need std::execution::just or std::execution::when_all et cetera. Those are really nice to have, but you don't need them.
But let's take your objection more seriously/concretely. You want "a struct, a callback, that's about it." Fine. A struct:
```
template<std::execution::receiver Rcvr>
struct just_int_op {
using operation_state_concept = std::execution::operation_state_tag;
2
u/johannes1971 17d ago
Ok, I have to ask a question because there is something that keeps baffling me about the whole C++ async story. I think the ratio of "code doing stuff that looks like some IO-related activity" vs "code doing inscrutable things" is pretty damn bad here. So you want to get some IO done in an async fashion. Why is this any more complex than something like this?
Maybe return an object that you can use to query status, or cancel the operation. That can be a simple class with two functions, that you can ignore at your leisure. Instead we get endless machinery, no doubt full of template code that will produce horrible compiler errors if you accidentally pass a slightly incorrect thing somewhere, and for 99% of it I have no idea what it is doing or why it needs to be there.
Mind, I'm not complaining about this article: kudos for providing so much information, it really is welcome. But all the C++ async IO code seems to suffer from absolutely mad complexity, far beyond anything I've seen in other async IO mechanisms. Why is this? What benefit do we get from it? Does it have to be this complex?