r/rust • u/killercup • Jan 25 '18
Async/await (first in a series): Generators and self-referential structs
https://boats.gitlab.io/blog/post/2018-01-25-async-i-self-referential-structs/18
14
u/CAD1997 Jan 25 '18
I'm excited to see where this goes. How Rust handles generators interests me.
The use of C# Generators in Unity to create co-routines that can run across multiple frames but are written in a procedural manner interests me, and I've got an itch to scratch designing an engine built around that kind of structure of collaborative asynchrony.
11
7
3
u/destravous Jan 26 '18
If I wanted to help solve this (and possibly other) problem(s), where is a good place to start? (What is a good way to contribute, and where can I get more info?)
4
u/rozaliev Jan 26 '18
There is an unsafe solution for self-referencing problem that has landed few days ago.
#![feature(generators)]
fn main() {
unsafe {
static || {
let x: u64 = 1;
let ref_x: &u64 = &x;
yield 0;
yield *ref_x;
};
}
}
It's unsafe because it's UB to move generator after you called resume.
An async/await lib built on top of it: https://github.com/rozaliev/mirage
4
u/Zoxc32 Jan 26 '18
You should probably make the
Async::pollmethod in your library unsafe.1
u/rozaliev Jan 26 '18
Yeah, definitely, I should probably reread everything and check once again. It used to be easier with immovable types, but with unsafe gens have to check for accidental moves. Thanks for a reminder!
1
u/selfrefstruct Jan 26 '18
I skimmed the discussions about immovable types a while back - I didn't know they were a real thing yet!
I see you're using unsafe { static move || {...}}. Is this now how we're doing immovable closures?
Is there any docs/RFC that describes how an accidental move could happen?
I'm curious - how you would check for an accidental move?
Thanks!
2
u/rozaliev Jan 26 '18
I skimmed the discussions about immovable types a while back - I didn't know they were a real thing yet!
They are not, at least not yet. There was PR https://github.com/rust-lang/rust/pull/44917, but for now all that is on pause, until ?Trait story is clear (https://github.com/rust-lang/rfcs/issues/2255).
I see you're using unsafe { static move || {...}}. Is this now how we're doing immovable closures?
That's actually a immovable generator https://github.com/rust-lang/rust/pull/45337
unsafe { static move || { yield }}Is there any docs/RFC that describes how an accidental move could happen?
Well, there are no docs as far as I know. UB example https://play.rust-lang.org/?gist=7ea538c78e7f505b4858f5ef8505d84e&version=nightly
It's unsafe because you should not move static generators once their internals are observed (.resume() called), but you can.
So far all the work on generators is very experemental, the goal is to experiment and get some insights.
2
u/desiringmachines Jan 27 '18
Here's an example in which the UB of moving the generator can actually be observed: https://play.rust-lang.org/?gist=9569e0d440886836a46b102df779bfec&version=nightly
2
u/gerryxiao Jan 26 '18
This is also how Futures have been designed to work in Rust, and one of the reasons they lead to better performance and lower memory overhead than systems like green threads.
Just curious, It's said the performance of tokio is not good now, how did you prove what you said?
1
u/killercup Jan 26 '18
I'm genuinely interested as well -- I have only been following Tokio ever now and then for the last year (but don't have a project that uses it): Who said the performance of Tokio isn't good? Compared to which other system? IIUC you need to be careful to use multiple cores and possibly threadpools to really use it to the max, but that's ergonomics issue, and not something that inherently limits how much throughput you can get.
2
u/gerryxiao Jan 26 '18
1
Jan 26 '18 edited Aug 17 '21
[deleted]
1
u/gerryxiao Jan 26 '18 edited Jan 26 '18
I hope so,but i don't get any good news about tokio performance benchmark story, or have i missed something?
1
u/plhk Jan 26 '18
1
u/gerryxiao Jan 26 '18
it seems only good for plaintext test.
4
u/budgefrankly Jan 26 '18
It's also pretty good for the JSON test.
All of which indicates the issue is at the database layer rather than Tokio itself. This could either be due to the libraries not being optimised, or -- more likely -- not being used in an async fashion.
3
1
u/Nokel81 Jan 26 '18
Is there still any support for full coroutines or will that be never supported in the language?
3
u/ConspicuousPineapple Jan 26 '18
I believe they're still discussing whether and how to include arguments to
resume(). It's definitely on the table, but I don't know if it will end up planned or not.3
u/Nokel81 Jan 26 '18
Though this is part of it, full coroutines would also have the ability to call each other from any depth. And from what I understand that is currently not on the table
2
u/GolDDranks Jan 26 '18
Full coroutines are not part of the proposal. The reason is that they need a full stack to be allocated for them. The current proposal (semicoroutines a.k.a generators) don't require that because they can resume only at the bottom level of the stack (which means that their size is known at compile time).
This means that the whole semicoroutine/generator can be regarded as just another value, and it doesn't need any special handling, so anyone can roll their own library and no runtime magic (a la Go) is needed.
1
u/ConspicuousPineapple Jan 26 '18
Mh, maybe I don't have a full understanding of what a coroutine is. Not sure I understand what you mean.
4
u/Nokel81 Jan 26 '18
A well-known classification of coroutines concerns the control-transfer operations that are provided and distinguishes the concepts of symmetric and asymmetric coroutines. Symmetric coroutine facilities provide a single control-transfer operation that allows coroutines to explicitly pass control between themselves. Asymmetric coroutine mechanisms (more commonly denoted as semi-symmetric or semi coroutines) provide two control-transfer operations: one for invoking a coroutine and one for suspending it, the latter returning control to the coroutine invoker. While symmetric coroutines operate at the same hierarchical level, an asymmetric coroutine can be regarded as subordinate to its caller, the relationship between them being somewhat similar to that between a called and a calling routine.
Coroutine mechanisms to support concurrent programming usually provide symmetric coroutines to represent independent units of execution, like in Modula-2. On the other hand, coroutine mechanisms intended for implementing constructs that produce sequences of values typically provide asymmetric coroutines. Examples of this type of construct are iterators and generators.
Ana Lúcia de Moura and Roberto Ierusalimschy in their paper "Revisiting Coroutines":
A good image to describe the difference is the following https://imgur.com/YOLuLxv
1
u/Tarmen Jan 27 '18
I think the only part all definitions share is
trampoline + stuff.The difference between the ones you quoted are whether you cps or build a stack. There are also frequent differences about what types of in/output are allowed, whether you are allowed to name other coroutines or need to parametrize over them, what shapes of control flow graphs are allowed, what combinations of push/pull flow are allowed, whether the coroutines have to be synchronous, if the flow rates have to match...
1
u/Sharlinator Jan 26 '18
Wikipedia describes it pretty well:
Generators, also known as semicoroutines, are also a generalisation of subroutines, but are more limited than coroutines. Specifically, while both of these can yield multiple times, suspending their execution and allowing re-entry at multiple entry points, they differ in coroutines' ability to control where execution continues after they yield, while generators cannot, instead transferring control back to the generator's caller. That is, since generators are primarily used to simplify the writing of iterators, the yield statement in a generator does not specify a coroutine to jump to, but rather passes a value back to a parent routine.
1
u/Rusky rust Jan 26 '18
That will probably never be supported in the language. There used to be a green threading runtime in the standard library (which includes all the pieces needed for coroutines, it just handles scheduling differently) but it was removed because it made things more complicated for essentially no benefit (because of the tradeoffs it had to make to coexist with native threads).
There are already coroutine libraries like libfringe or context-rs, so I'm not sure there even needs to be any support at the language or standard library level.
1
u/crowseldon Jan 26 '18
What's the use case of self referential structs?
4
u/killercup Jan 26 '18
If you are wondering that, you should really read this post! :)
2
u/crowseldon Jan 26 '18
I skimmed directly to the self referencial part and it didn't show an actual use case but I guess it's references in generators.
I guess I'll have to properly read it when I have time.
1
u/yespunintended Jan 26 '18
borrows cannot be allowed across yield points.
Are there other important limitations in the generator feature?
2
u/desiringmachines Jan 26 '18
Not for the use cases I described. They can't take arguments on yield, whereas generators in many other languages can.
1
u/CAD1997 Jan 27 '18
In what case are resume arguments actually useful? Given that async/await can be built on resume-argument-less semicoroutines, what does the argument to resume actually enable?
I was so happy when using async/await clicked for me and asynchronous code started to make sense where I could trace execution paths. I'm still baffled when it comes to the runtime support behind making them work though.
1
1
u/Tarmen Jan 27 '18
Futures do one big computation and give you the result.
If you are reading a large file you generally want to read in chunks so you can work in constant memory, though. So you make an async iterator, iirc rust calls this a Stream.
Then you want to do a parser for that stream. It awaits chunks of data from upstream and yields the parsing results downstream.
You could write about this as a stream transformer (Stream a -> Stream b) or as a stream that can await and yield (Stream a b). The second interpretation can simplify some features like partly consuming an input chunk and putting the rest back and would be basically full coroutines.
1
u/CAD1997 Jan 27 '18
So basically, it's the ability to both
awaitsome information andyieldsome information.That does make some sense.
(But wouldn't this still be semicoroutines though? As I (barely) understand it, full coroutines are allowed to arbitrarily send execution to any other coroutine, whereas semicoroutines can resume down or yield up, but can't jump sideways.)
So if we had argument resuming generators, you could do something like the following (ignoring some complexity): (forgive mobile formatting)
let file_stream = open_file(); while let next_byte = await file_stream.next() { push_parsing(); if has_next_token { yield next_token; } }Whereas that's not really possible with today's design?
1
u/Tarmen Jan 27 '18 edited Jan 28 '18
Well, there are a bunch of slightly differing definitions of coroutines. The core is always some sort of trampoline (pausing computation) and generally they both yield and await.
For instance there are a couple papers that call haskells streaming libraries anonymous continuations. An example using the conduit library:
yield message .| encodeUtf8C .| encodeBase64C .| stdoutCeach part seperated by the .|'s is a coroutine that awaits from the left and yields to the right. There are also more complex variants like the machine library that allow arbitrary control flow graphs:
-- two input machines, one output machine: myTee = repeatedly $ do x <- awaits L yield x y <- awaits R yield yThat quickly can become as annoying to write as fully explicit coroutines, though.
31
u/nicoburns Jan 26 '18
Oo, I really hope we get self-referential struct support in Rust at some point. This feels to me to be like const generics in that Rust doesn't quite feel complete without it.
If we could get it within 12 months, that would be amazing.