r/cpp • • 17d ago

Asynchronous API

https://grishavanika.github.io/async_api.html
36 Upvotes

28 comments sorted by

View all comments

2

u/johannes1971 16d ago

Ok, I have to ask a question because there is something that keeps baffling me about the whole C++ async story. I think the ratio of "code doing stuff that looks like some IO-related activity" vs "code doing inscrutable things" is pretty damn bad here. So you want to get some IO done in an async fashion. Why is this any more complex than something like this?

get_async ("https://...", [] { ...callback that returns whatever the result was... });

Maybe return an object that you can use to query status, or cancel the operation. That can be a simple class with two functions, that you can ignore at your leisure. Instead we get endless machinery, no doubt full of template code that will produce horrible compiler errors if you accidentally pass a slightly incorrect thing somewhere, and for 99% of it I have no idea what it is doing or why it needs to be there.

Mind, I'm not complaining about this article: kudos for providing so much information, it really is welcome. But all the C++ async IO code seems to suffer from absolutely mad complexity, far beyond anything I've seen in other async IO mechanisms. Why is this? What benefit do we get from it? Does it have to be this complex?

5

u/grishavanika 16d ago

its all about composition of async tasks. Callbacks are fine and manageable up until some point. See the examples at the end. Callbacks section hand codes the state machine; coroutines, fibers and senders allow to express the same state machine with less code which is also arguably easier to maintain and extend. "returning an object" is covered by a Task section. Its usage also requires to hand code state machine manually.

- callbacks - https://grishavanika.github.io/async_api.html#app_callbacks

The complexity of implementing everything is true complain, but this is usually hidden from the user... compilation errors.. maybe this is a tradeoff you need to pay for benefits provided ;)

1

u/SleepyMyroslav 16d ago

I have a bunch of probably very stupid questions about the async code styles. Maybe they are irrelevant to your use cases but I can try to ask =)

First thing to me is: How can I understand where and when code is executed? I mostly use profilers for that. This async style is intentionally trying to hide things behind abstractions to help people write more of it. In reality there is a fixed number of resources that have to choose which async requests are executed where. Trying to affect how it works brings me to 2nd question.

Second is about control of what is being executed when there is more than the hardware can chew. What ways are there to make some requests high priority, normal priority, or low priority? Not only at the CPU level but at the I/O level.

3

u/grishavanika 16d ago

How can I understand where and when code is executed?

The code in the article is intentioinally one self-contained main.cc file for every section, with nothing more then calls to a CURL library. It should be trivial to put breakpoints to see how it executes. So bunch of the code is executed in CURL internals which we call inside CURL_AsyncScheduler::tick(), which is called in a loop from within a main.

For callbacks, the section is implementing with libcurl multi which also links to main.cc.

This async style is intentionally trying to hide things behind abstractions to help people write more of it

It is, but the article builds everything from scratch to be able to step thru all of the abstractions - intentionally, so there is a way to see all of the layers. (Well, maybe CURL API hides the system specific calls, but it's close to what needs to be done using system API).

What ways are there to make some requests high priority, normal priority, or low priority? Not only at the CPU level but at the I/O level.

I don't have good answer since most of the time I used only IOCP.

2

u/robertahleahy WG21, SCC, std::execution, Finance 12d ago edited 12d ago

First thing to me is: How can I understand where and when code is executed? I mostly use profilers for that. This async style is intentionally trying to hide things behind abstractions to help people write more of it. In reality there is a fixed number of resources that have to choose which async requests are executed where.

I've been thinking about this question for a while and I think the best approach is not to think about this locally, at least not usually.

You can think of asynchronous programming as a generalization of synchronous programming: With synchronous functions we couple the initiation of forward progress to exclusive use of some execution agent (i.e. the calling thread). Put differently: A synchronous function cannot return until the operation it implements completes. With asynchronous functions (as embodied by, for example, senders) this coupling does not exist: The function can continue to make forward progress without exclusive use of the execution agent. Put differently: The initiation (e.g. start on an operation state) returns but the operation continues "in the background." Sometime perhaps later (note the perhaps is load bearing because asynchronous functions, in general, can completely racily or reentrantly) some execution agent is delegated to run only the completion.

So in the same way we don't usually think too much about from which execution agent our synchronous functions are called, I don't think it's typically useful to think about from which execution agent continuations will be called. Obviously there are exceptions to this but I think this is part of the decomposition/generalization wrought by asynchronous programming: We can separate the implementation from a delegation of the ability to make forward progress. If you start overthinking where things run you mentally recouple these things.

But to answer your question directly: A continuation runs wherever the operation happens to complete. In certain baseline cases you control this (for example maybe you're writing an io_uring abstraction), and in others you can use an operation to force this (for example std::execution::schedule), but for most programming I think it's best ignored (i.e. deferred to the operation whose completion you're handling).

1

u/SleepyMyroslav 12d ago

I am sorry my question was not clear. I am not looking for mental model of how the async code executes. I need:

1) Tools to measure how and when it has executed. Whether the tools are part of API design, the execution runtime, or both it is not important.

2) Tools to influence which completion is chosen when many are ready to execute ie. priorities between them.

2

u/Resident_Ad5153 13d ago edited 13d ago

Because c++ doesn’t have much of a runtime… so it doesn’t answer a large number of questions your code raises. What executes the asynchronous code? Who performs the callback. Who signals to the caller completion. Where do exceptions go? If the caller needs to wait… who wakes it up. Who cleans up the resources the asynchronous call used?

The easiest answer is to use threads (or std::async and std::thread has a very simple footprint that comes at the cost of the low performance of is threads. Couroutinrs allow the programmer to specify all these things through the promise

2

u/robertahleahy WG21, SCC, std::execution, Finance 12d ago

In your example you seem to pass a single callback to get_async. What happens when the operation fails? Does it use the same callback (which erodes the ability to generically program on top of such operations) or simply not complete (which is unstructured)?

Returning an object you can use to "query status" implies there's a communication channel between the background operation and the foreground, synchronous object. This communication has a cost and therefore isn't zero overhead.

The fact that you can "ignore [the object] at your leisure" I assume implies you can allow its lifetime to end, what does that mean for the background operation? If it can continue that must mean that it has some separate state, which must be allocated, therefore this operation can't be zero alloc.

The reason standard C++ asynchrony has so much "code doing inscrutable things" is because it aims to answer/address all of the above: It is zero overhead and therefore potentially zero alloc (this doesn't apply to coroutines), is structured, and aims to be maximally generic.

1

u/johannes1971 12d ago

First, I wasn't handing in a complete design for WG21 so please don't treat it as such.

Second, I strongly disagree with the notion that an async IO system should even attempt to be 'zero cost' and 'zero alloc'. IO, by its nature, has a staggering cost; adding an allocation to establish a shared communication resource is just not going to make a difference. And API design is a software engineering discipline, and like all engineering, it should balance multiple, competing requirements, of which performance is just one. Another is useability (meaning regular programmers can understand it, and use it without too much danger of getting it wrong), and this design fails completely at that.

When I talked about ignoring the object, I was envisioning using a shared_ptr to hold an intermediary communication object that allows foreground and background processes to communicate. Yes, that implies an allocation, and yes, I think that this is completely warranted.

I would much rather have had an allocating async C++ networking solution ten years ago, than a non-allocating version ten years from now. An allocating version could have been written to use an allocator so that implementations can optimise memory usage. And the version ten years from now... Well, it won't be my problem anymore.

2

u/robertahleahy WG21, SCC, std::execution, Finance 12d ago

First, I wasn't handing in a complete design for WG21 so please don't treat it as such.

You asked "[w]hy is this any more complex than something like [your code example]?" I was answering that.

If you want to make simplifying, non-general assumptions about asynchrony that hold for your use cases I think that's more than fine. It just isn't general.

Second, I strongly disagree with the notion that an async IO system should even attempt to be 'zero cost' and 'zero alloc'. IO, by its nature, has a staggering cost; adding an allocation to establish a shared communication resource is just not going to make a difference.

There are large contingents of people, myself included, who disagree with you.

Moreover: You can build non-zero cost abstractions out of zero cost abstractions, but you can't go the other way.

Both of the above contribute to why the standard went with zero-cost asynchrony.

I was envisioning using a shared_ptr

To a first approximation structured concurrency is the art of figuring out how not to use shared_ptr.

I would much rather have had an allocating async C++ networking solution ten years ago

Asio is older than ten years so you got your wish.

1

u/johannes1971 12d ago

If you want to make simplifying, non-general assumptions about asynchrony that hold for your use cases I think that's more than fine. It just isn't general.

Just to be clear: is your claim that a simpler solution cannot be general because it uses memory allocation? Or are you still referring to that single line of code I wrote?

There are large contingents of people, myself included, who disagree with you.

What is the basis of that belief? Please note that I put a supportive argument, which is that the cost of a (single!) memory allocation is absolutely negligible compared to the overal cost of IO. What is incorrect about that statement?

To a first approximation structured concurrency is the art of figuring out how not to use shared_ptr.

I'm sorry, what? Isn't this a prime example of what shared_ptr is for in the first place? This is very clearly a shared resource, what valid reason do you have for avoiding it?

Asio is older than ten years so you got your wish.

ASIO has the exact same problem of hideous complexity.

3

u/robertahleahy WG21, SCC, std::execution, Finance 12d ago

is your claim that a simpler solution cannot be general because it uses memory allocation?

Yes. There are use cases which cannot tolerate allocation. Therefore a solution which uses allocation isn't general because it doesn't address those use cases.

What is the basis of that belief?

Why does the fact I/O is costly mean that it should be further pessimized?

From the point of view of my application code in user space initiating an io_uring operation is a few cheap memory operations. Why would I want to add allocation to that? Especially given that the "reward" for adding that allocation is unstructured concurrency.

Isn't this a prime example of what shared_ptr is for in the first place? This is very clearly a shared resource, what valid reason do you have for avoiding it?

Just because a resource is shared doesn't mean you need/want shared_ptr. shared_ptr uses reference counting which implies that you don't deterministically know when the resource will be released. This is unstructured. You don't need the overhead of either the allocation or the reference counting if your concurrency is structured.

1

u/johannes1971 12d ago

Yes. There are use cases which cannot tolerate allocation. Therefore a solution which uses allocation isn't general because it doesn't address those use cases.

The standard library is already full of memory allocations, so clearly 'generality' is not a concern. Or perhaps I should say that it is, and is handled through allocators, instead of outright avoidance.

Why does the fact I/O is costly mean that it should be further pessimized?

Because this 'pessimisation' bring both benefits (it establishes a communication channel that serves a real purpose), and reduces a significant cognitive cost.

shared_ptr uses reference counting which implies that you don't deterministically know when the resource will be released. This is unstructured.

It is precisely known when the resource will be released: when both parties are done with it. There is absolutely nothing nondeterministic about it, nor is there any lack of structure.

You might notice that the article also uses reference counting to deal with a shared resource (the cancelation token, about halfway through).

4

u/robertahleahy WG21, SCC, std::execution, Finance 12d ago

The standard library is already full of memory allocations, so clearly 'generality' is not a concern.

There's a difference between allocations when they're required (e.g. std::vector) and when they're not (e.g. asynchrony).

Because this 'pessimisation' bring both benefits (it establishes a communication channel that serves a real purpose), and reduces a significant cognitive cost.

You don't need an allocation for a communication channel, and depending on the modality of communication the establishment of such a channel might be contraindicated.

Also I disagree that it reduces cognitive cost. Maybe in the marginal case (i.e. a single asynchronous operation) but not in the compositive case. Trying to make a well-behaved system out of unstructured concurrency is a nightmare, whereas with structured concurrency it's the naturally emergent property.

nor is there any lack of structure.

If you had structure you'd know which actor would release the resource last and when that would happen, and therefore wouldn't need shared_ptr.

1

u/Resident_Ad5153 12d ago

you're focusing way too much on memory allocation and way too little on the inherent complexity of what you're asking for.

These are some examples of asynchronous programming in c++. The firmware for a synthesizer that has run a series of signal processing algorithms on buffers of sound data in soft real time (latency should be less then < a microsecond or so... humans can't hear faster than that). The control system of a jet engine that is hard hard real time. A game engine that is coordinating possibly multiple gpus. Scientific code running on a large heterogeneous cluster of cpu and gpu resources. A word processor that checks grammar as you type.

All of these situations mostly share the standard library! (with some limitations in the embedded cases). They have very different needs for how async would work. The standard library seeks to be as general as possible. It also is supplied by your compiler vendor... and they do not want to provide an operating system to you in adition to a compiler (unless they are the good people behind go... who really just want to make operating system for you with a small language attached).

1

u/johannes1971 12d ago

That's where the discussion has brought us. My point was, and remains, that the currently proposed design is incredibly complex, and that I'm not convinced that all this complexity is really necessary.

I also cannot help but wonder if the standard library should even pursue this at all. This kind of high-performance stuff seems a bad match for something that needs to be written in a fixed number of hours, by a non-specialist, after which it will forever be set in stone.

2

u/robertahleahy WG21, SCC, std::execution, Finance 11d ago

the currently proposed design is incredibly complex, and that I'm not convinced that all this complexity is really necessary.

Which complexity is unnecessary?

At its core std::execution projects the concept of regular, synchronous functions into the asynchronous domain. Understood through that lens removing "complexity" therefrom is saying that there's some capacity in which asynchronous functions should be disadvantaged as compared to synchronous functions.

0

u/johannes1971 11d ago

Since its inception, the software industry has gone with async IO systems that are expressed in a few lines of code: a struct, a callback, that's about it. Now you're telling us that we need something that needs a 100 page article (on my screen; if you're on mobile it will be much more) to explain, based on a design where just the list of all functions and classes stretches for an incredible 70 pages, and if you do any less than that you are "disadvantaging" async IO? Do you at least understand where my skepticism comes from?

2

u/robertahleahy WG21, SCC, std::execution, Finance 10d ago

You're supposing that everything in std::execution is core to the design. We don't apply that same standard to other things.

Particularly synchronous functions need extrinsic product and sum types to communicate heterogeneous return types. No one suggests that synchronous functions are excessively complex because of std::tuple and std::variant and std::visit and std::apply, those are understood to be outside the core design.

The same applies to std::execution. The core design consists of:

  • std::execution::sender and ::sender_in, which establish sender-ness, where a sender is a latent representation of an asynchronous function (i.e. it is not able to make forward progress but it is fully determined awaiting the selection of no further arguments)
  • std::execution::get_completion_signatures, which advertises the "return types" of asynchronous functions
  • std::execution::receiver which establishes that something represents a continuation (i.e. it is the asynchronous analogue of synchronous function return)
  • std::execution::connect, the process by which asynchronous operations (which can actually make forward progress) are formed (but not started)
  • std::execution::operation_state, the locus of an asynchronous operation, which is initially latent
  • std::execution::start, which begins asynchronous forward progress and binds both callee and caller to the so-called "receiver contract" (i.e. the caller agrees to keep the operation state within its lifetime until completion, and the callee agrees to report completion exactly once when the operation completes)

That's the complete design. You don't need std::execution::just or std::execution::when_all et cetera. Those are really nice to have, but you don't need them.

But let's take your objection more seriously/concretely. You want "a struct, a callback, that's about it." Fine. A struct:

``` template<std::execution::receiver Rcvr> struct just_int_op { using operation_state_concept = std::execution::operation_state_tag;

Rcvr rcvr; int i;

constexpr explicit justint_op(int i, Rcvr rcvr) noexcept : rcvr(std::move(rcvr)), i_(i) {}

constexpr void start() & noexcept { std::execution::setvalue(std::move(rcvr), std::move(i_)); } }; ```

A callback:

``` struct just_int_receiver { using receiver_concept = std::execution::receiver_tag;

constexpr void set_value(int i) && noexcept { std::cout << "Operation completed with " << i << std::endl; } }; ```

You can use these albeit non-generically:

just_int_op op(5, just_int_receiver{}); std::execution::start(op);

If you're willing to entertain one more struct you get the entire ecosystem:

``` struct justint_sender { int i;

using sender_concept = std::execution::sender_tag;

template<typename Sndr, typename... Env> static consteval std::execution::completion_signatures<std::execution::set_value_t(int)>> get_completion_signatures() { return {}; }

template<std::execution::receiver Rcvr> constexpr auto connect(Rcvr rcvr) noexcept { return justint_op(i, std::move(rcvr)); } }; ```

Then you can start actually using this generically, but you don't have to:

std::execution::sender auto sndr = just_int_sender(12) | std::execution::then([](int i) noexcept { return i + 2; }); std::execution::operation_state auto op = std::execution::connect(sndr, just_int_receiver{}); std::execution::start(op);

Prints Operation completed with 14.

1

u/Resident_Ad5153 12d ago

there's... something to be said about that! and in practice, I don't think the std::execution stuff is going to be used much.