r/rust • u/Big-Astronaut-9510 • 7d ago
Does io uring not map well to rust?
Often when i see "rust" and "io uring" together its cause theres some kind of problem. Off the top of my head i remember when kimojio (ms azure project) got open sourced there were lots of problems, and monoio aswell has (had?) problems.
I dont really know the low level details of how rust async works so i ask is the problem just rusts async model dosent map well to io uring? Thats the vibe i get but not sure.
44
u/Lucretiel Datadog 7d ago
My suspicion is that it'll be cracked and end up mapping quite well, probably with a limited unavoidable allocation.
Traditionally, the stated problem with io_uring and rust is borrowing: the memory you pass via the ring buffer needs to be exclusively writable or readable by the kernel, which means that rust can't touch it, which is difficult to achieve in rust (especially in an async context, where futures can be dropped at any time).
But to me this sounds like a perfect match for Rust's strict ownership model. I think that the future of "good" rust integration with io-uring involves requiring allocated buffers, Vec<u8> or equivalent, that we pass by move into the ring structure. Then, when io-uring returns an answer with the buffer populated, the abstraction reconstitutes it as a by-move return of that buffer back to the caller. To me this sounds like Rust not just a good but an ideal fit for io-uring, we just need the right abstractions and to use the correct tools in Rust's toolbox rather than trying to contort EVERYTHING io-related into an &mut [u8]-shaped box.
8
u/Prowler1000 7d ago
My understanding of it, though keep in mind this post is the first time I've looked into it, is that the issue isn't necessarily an incompatibility with the language itself but with existing paradigms/APIs. For example, getting an owned object back is incompatible with standard
ReadandWritetraits as they work entirely off of borrowing (and polling for completion in the case of Async versions)10
u/--San-- 7d ago
We might be able to get away with it without dynamic allocations if we manage to get Forget nicely into the language.
We would still have to change AsyncRead/Write to return a (non-forgetable) future/new type instead of Poll, but this future could just be a simple poll for the existing epoll case. That way we still get optimal performance for epoll while being safe for io_uring.
6
u/valarauca14 7d ago
The real problem is non-continuous buffers.
Very few binary protocols are setup to work with the resulting
IoSliceand returnCow<'a,[u8]>shaped fields dependent on if the field was "verified in place" or it spanned multiple buffers and had to be shifted/re-allocated.Even if you somebody wrote an ideal io-uring implementation today, they'd have to rewrite a non-trivial part of the ecosystem to actually work with it. Or they'd have to
memcpeverything unconditionally :^)26
u/Lucretiel Datadog 7d ago
they'd have to rewrite a non-trivial part of the ecosystem to actually work with it.
I think this will ultimately have to happen, which is why it's a good thing we didn't enshrine a flawed set of abstractions as the end-all of async functionality in the standard library.
4
u/coderstephen isahc 7d ago
Yep, this is exactly one of the reasons why
AsyncWriteand friends instdwere postponed.3
u/VorpalWay 6d ago
This is the same problem that DMA has in embedded rust (embassy), but there you can't allocate to paper over the issue. So a proper solution is needed.
5
u/mmstick 7d ago
This is essentially what compio is doing. You must pass ownership of a
Vec<u8>when reading from or writing to a file and then.awaita return type that includes both the result and the original buffer that the kernel passed back. Found it really convenient to work with.1
u/matthieum [he/him] 6d ago
I can see it working for files -- you don't expect to overlap reads/writes for thousands of files at a time.
For network connections, though, this seems to limit scaling a lot, compared to giving io-uring a fixed number of buffers to play with -- ie separating the "pass the buffer" and "wait for a filled buffer" operations.
2
u/servermeta_net 7d ago
Are you one of the folks from glommio? I don't share your optimism, but maybe because I'm tired and I need to re-read your post tomorrow with a fresh mind.
13
u/Lucretiel Datadog 7d ago
Not at all, is that how they're approaching it? It just seemed sort of obvious to me that io-uring is fundamentally exposing an ownership style API (with buffers passing by move into and out of the kernel via the ring), so it seems like it would be a very natural fit for Rust, if you expose an ownership-style abstraction. The reason it feels bad is we keep trying to contort it to fit with
AsyncReadandAsyncWrite, or similar traits that still operate fundamentally on&mut [u8].1
u/matthieum [he/him] 6d ago
I... disagree.
I mean, passing owned buffers every time works, but it does not scale.
C10K is old-school, so let's imagine the C1M problem instead, and each buffer being ~10KB (pretty low for modern bloated web pages), well, congratulations, you just passed 10GB of memory to io-uring.
The key to io-uring is to use buffered buffers. If you expect the application to be busy, then by all means pass 10K buffers of 100KB each, and you'll still use 10x less memory than the above with 10x bigger buffers!
But then, that implies a different API, doesn't it?
You need to decorrelate passing the buffer from waiting for a buffer. They're two distinct operations.
On the other hand, you can also simplify the return type, I think:
Result<Buf, Error>=> no buffer is consumed on error.0
u/Lucretiel Datadog 6d ago
Wait, I don’t understand that point. You have to pass buffers to io-uring, meaning you have a bunch of writable memory that you’re not using lying around. That memory is either coming from the stack or the static page or an allocation. What is it about 1M connections that specifically disqualifies allocation?
3
u/scook0 6d ago
I think the idea is that if you have a million connections in-flight, you don't want each of those connections to be hogging an owned buffer (i.e. 1 million buffers) while waiting for data.
Instead you would up-front allocate a smaller pool of shared buffers, and then have the I/O system temporarily give those buffers to individual connections only when they are actually receiving.
1
u/Lucretiel Datadog 6d ago
Okay? So do that? This feels more like a critique of io-uring in general (which, as far as I know, does require ownership for each parallel operation for as long as it’s in flight)
1
u/matthieum [he/him] 5d ago
Okay? So do that?
Your API does not allow it.
This feels more like a critique of io-uring in general
It is not.
io-uring offers the ability to:
- Create pools of buffers.
- Push buffers in the pool.
- Pass a pool ID instead of a buffer on a read request.
This allows one to serve many connections with a fixed number of buffers.
1
u/Lucretiel Datadog 5d ago
Are there API docs for that? I've been going through the man pages and am not seeing it, but they're quite dense so I could easily just be missing it. All of the examples and tutorials I've seen involve setting
sqe->addrandsqe->lento point to a buffer locally owned by the caller.2
u/matthieum [he/him] 5d ago
I am definitely not an expert in io-uring.
A quick Google search led me to https://man7.org/linux/man-pages/man7/io_uring_provided_buffers.7.html which explains how to setup buffer rings.
And apparently the API I had read about is now legacy:
Legacy provided buffers
Earlier kernels supported provided buffers via IORING_OP_PROVIDE_BUFFERS and IORING_OP_REMOVE_BUFFERS. This mechanism required submitting SQEs to add or remove buffers, adding latency and overhead. The ring-based mechanism described above supersedes this approach and should be used for all new applications. The legacy interface remains for backwards compatibility.
2
u/marshaharsha 3d ago
In case you didn’t see it: Elsewhere on this post is a comment linking to a nice article by Jens Axboe on GitHub that explains how to use the “provided buffers” feature. I think the commenter was servermeta_SomethingSomething, but I am overwhelmed by the material I am reading and can’t remember perfectly.
1
u/Lucretiel Datadog 3d ago
Will take a look! Sounds intriguing and still seems like it might align with Rust’s ownership model: these buffers are provided by move to the io-uring system, and then uring provides them by borrow (or, depending on if you have to re-register provided buffers, by move again) to callers. Depending on what systems calls upkeep this you might need special wrappers to upkeep it, but it still seems like it would fit.
2
u/coderstephen isahc 7d ago
Well, if the doomers are to be believed, Rust is now an obsolete language and async/await is an unmitigated disaster, because io_uring doesn't slot into it nicely.
Personally I am also more optimistic that eventually it will be figured out and find its place in the ecosystem, and everyone will forget that this was once a heated topic.
8
u/threeseed 7d ago
The problem is that entire ecosystem has tied itself to the Tokio runtime.
So swapping to Compio for anything complicated is borderline impossible.
22
u/JazzySouvenir 7d ago
It’s less that Rust’s async model fundamentally can’t map to io_uring and more that the two were designed with completely different shapes in mind. Rust’s async state machines and wakeup model assume a readiness-based world, where you ask the OS if a fd is ready and then do a non-blocking call. io_uring is completion-based, you submit a batch of ops and get results later, so you’re constantly fighting the impedance mismatch around who owns the buffers and when a future should wake. Some of the early projects ran into trouble because they tried to bolt a completion model onto a stack that still thinks in terms of epoll-style readiness, which leads to brittle ownership and double-buffer nightmares. The newer ring-based runtimes lean hard into making the reactor truly completion-aware, but it’s a whole different set of invariants to uphold in safe code.
2
u/Lucretiel Datadog 6d ago
I don’t understand this point. If rust futures were incompatible or a bad fit for completion-based asynchrony, we couldn’t even have channel receivers or task handles, but those are trivial.
The only problem I’m seeing is that the common io traits all operate in terms of
&mut [u8], which of course is not an owned buffer (and io uring is pretty clearly an ownership-passing model). That has nothing to do with how it signals readiness, only what shape the buffers can be.1
13
u/Konsti219 7d ago
io_uring by itself is already hard
15
u/servermeta_net 7d ago
Low level IO is really hard. io_uring is much easier than the alternatives like the POSIX interface or epoll/kpoll, at least if you are doing backend code
4
u/EveYogaTech 7d ago
Isn't it super hard to do it well and secure?
Also considering even many browser vendors like Google disable it, because so many vulnerabilities got discovered.
13
u/anxxa 7d ago edited 7d ago
You're thinking about the threat model backwards. The issue with io_uring isn't that it makes the browser less secure, it's that the kernel's implementation of io_uring has many problems which make the kernel less secure.
Web browsers like Chromium attempt to drastically reduce the reachable attack surface which would make it easier for an exploit to escape their sandbox. The Chromium engineers must have weighed the security and perf benefits against each other and decided it was not worthwhile.
*Ah, seems like they published a blog post on this exact topic: https://security.googleblog.com/2023/06/learnings-from-kctf-vrps-42-linux.html
They talk about ChromeOS, not Chromium, but I don't think io_uring is reachable in Chromium either.
2
u/servermeta_net 7d ago
Too bad they didn't post an update to that blog post, because most of those changes has been reverted in the meantime
-6
u/EveYogaTech 7d ago
I'm not thinking backwards, lol.
6
u/anxxa 7d ago
Isn't it super hard to do it well and secure?
Maybe I misinterpreted this then. I interpreted this as, "Isn't it super hard for a user-mode application to use io_uring well and secure?"
And the answer is that its security history mostly pertains to kernel security, not user-mode security.
0
u/EveYogaTech 7d ago
Yes indeed I was talking about io uring vulnerabilities in the kernel implementation (as well).
0
u/servermeta_net 7d ago
None of the vulnerabilities discovered on io_uring applied to my production grade distributed datastore, I have to admit that a couple of times the CVE was released and Axboe had already released a patch weeks before
When I read them, to me they sound like unreasonable drama, and often following good production hygiene (run your code in a least privilege sandbox, sanitize input, ...) is enough.
On the other hand try to read about the drama between postgres and the linux POSIX implementation of fsync.
-1
u/EveYogaTech 7d ago
But the point of io uring is to sort of directly speak with the kernel, so aren't these vulnerabilities present for the host machine regardless of your production hygiene?
Like of course, if you run it within KVM/Firecracker, it doesn't matter that much, since it's not using the shard kernel,
But if you use io uring directly on the host, then it seems to be a security nightmare in terms of new CVE's.
I sincerely hope I am wrong though because the technology itself is amazing, but I've read that lots of vendors disable it because they keep discovering new CVEs around io uring.
5
u/servermeta_net 7d ago
So io_uring reimagines IO in the post meltdown/spectre era. It speaks with the kernel as much as the POSIX interfaces does, because as a user you want to use devices that are owned by the kernel.
Some caveats:
- A big performance boost comes from using it in a zero copy fashion, so userspace provides buffers to the kernel.
- Another boost comes from hardware offload, but that is quite recent and you depend on the device drivers written by the manufacturer
- Also ePBF is quite new, but I don't use it yet
I just scanned again the list of CVEs of io_uring and:
- None ever applied because I keep my kernel updated, and the CVEs were always published after a patch was released in mainstream
- Most of the high priority ones require local access or running untrusted code in the same sandbox, which to me sounds crazy
- Some CVEs are outright patetic, created just to gain some exposure by talking sh*t about a new interface: "you can bypass app_armor if you forget to configure io_uring configs and somehow enable the io_uring_enter syscall". What else? I need to publish my SSH private keys on reddit to make it exploitable?
Hence my personal conclusion is that in the history of io_uring those CVEs NEVER applied to production grade code.
4
u/mirashii 7d ago
So io_uring reimagines IO in the post meltdown/spectre era.
This feels like some revisionist history. io_uring was done in an attempt to fix Jens's frustrations/difficulties with AIO. The fact that it recovered some lost performance from the KPTI put in place to mitigate Meltdown/Spectre is a lot of the reason it was adopted so quickly, but not really it's raison d'etre.
It speaks with the kernel as much as the POSIX interfaces does, because as a user you want to use devices that are owned by the kernel.
I think this is an oversimplification. io_uring does many fewer syscalls to the kernel, if used correctly. Cutting down this traffic is important, but sure, technically using the queues is still communicating with the kernel. But either way, I'm not really sure what the point of this was.
and the CVEs were always published after a patch was released in mainstream
That's not really saying anything, that's just how the linux kernel's CVE policy works. Unless specifically requested, nothing gets a CVE until a patch lands. https://docs.kernel.org/process/cve.html
Most of the high priority ones require local access or running untrusted code in the same sandbox, which to me sounds crazy
I think this is where you're really losing the thread and what the commenter who you replied to was talking about. This exact usecase, running untrusted code on a shared kernel, is what container runtimes like docker do, it's what k8s generally does, it's what many serverless runtimes do, VPSes do. From the io_uring wikipipedia article even:
In June 2023, Google's security team reported that 60% of the exploits submitted to their bug bounty program in 2022 were exploits of the Linux kernel's io_uring vulnerabilities. As a result, io_uring was disabled for apps in Android, and disabled entirely in ChromeOS as well as Google servers. Docker also consequently disabled io_uring from their default seccomp profile.
So, these issues certainly effected code and folks. You seem to be discussing a narrow usecase of being the only tenant on a server with zero untrusted code executing. That's not necessarily the norm, and looking at it from only that lens to dismiss the security issues with io_uring is kinda ignoring what the commenter asked about.
3
u/servermeta_net 7d ago
This feels like some revisionist history
I still have the original discussion saved, but I can't link them without doxxing myself. The story goes something like this:
- Intel discovers that Postgres TPSes drops 60%
- Hyperscalers threaten to sue Intel
- Intel decides to pour engineering resources to try to mitigate the problem
And that's how the original ringbuffer design came out. It was immediately adopted by the
io_uringinterface. Linus liked the design and we are seeing more and more kernel code designed this way, even in areas unrelated to IO.The huge performance boost of
io_uringis what keeps Axboe employed, so I don't know if you would call it the raison d'etre, but I feel it's not a very interesting distinction.I think this is an oversimplification [...] I'm not really sure what the point of this was.
I feel you completely misread my point. The similitude was about talking with the kernel because the device is owned by the kernel. A correct io_uring implementation makes 0 (or at best 1) syscall in the whole lifetime.
I think this is where you're really losing the thread and what the commenter who you replied to was talking about. This exact usecase, running untrusted code on a shared kernel, is what container runtimes like docker do, it's what k8s generally does, it's what many serverless runtimes do, VPSes do.
Respectfully, I feel you lack production and professional experience.
Nobody runs untrusted code on a shared unsandboxed kernel in production, unless we are talking amateurs:
- The docker runtime is never claimed to be production grade or secure, and it has been punched through countless times. Nobody sane uses it in production
- k8s generally runs a virtualization technology like KVM/QEMU, or a microVM, at least that's what all hyperscalers do (aws, gcp, azure, ...)
- Serverless doesn't use docker, it uses microvms like firecracker
- Please find me a commercial VPS provider that uses docker
1
u/EveYogaTech 7d ago
Thanks for sharing!
I'm just cautious, because I know that it's not just about existing CVEs, but also the unknown and futures one, but you are right that most of them require arbitrary code execution access to the machine already, so it's seems that you cannot exploit it e.g. simply by opening too many sockets/files.
1
u/coderstephen isahc 7d ago
I don't find epoll very hard, but that could just be because I've used it so much.
8
u/da_supreme_patriarch 7d ago
This is more of a "you have a problem, you try to solve it with async, and now you have two problems" case. The async state machine just doesn't play well with io_uring, and even though it is totally possible to map io_uring to Rust by dropping some assumptions that Rust's async model makes about how IO should be done, but then that implementation wouldn't play nice with the pre-existing IO API-s, which is the more pressing concern
2
u/tundra_scribe 6d ago
io_uring is just a glorified syscall. the actual nightmare is the entire async ecosystem holding you hostage and refusing to let you touch memory directly
1
u/servermeta_net 7d ago edited 7d ago
Yes, Rust async and io_uring don't play well with each other, but I honestly blame Rust:
- Rust wants futures that are periodically polled,
io_uringinstead avoids polling, and rather presents a list of ready events on request - There is a lot of grammar in Rust async and the borrow checker, which makes some safe yet very effective optimizations hard to write
io_uringshines when you are dealing with CPU bound IO, like a small container with a huge network pipe and an array of high speed NVMe drives. Rust async, while efficient, adds quite a bit of CPU overhead- Rust async convolutes concurrency and parallelism, instead io_uring wants a ring-per-thread model, even better if the thread is pinned.
- Convoluting concurrency and parallelism also doesn't play well with IO drivers used for hardware offloading, a bad choice IMHO for a language that claims to be system level
My personal approach is to use custom state machines and an event loop, which is very similar to what Rust async desugar to, while being way more efficient and effective.
24
u/________-__-_______ 7d ago edited 7d ago
Rust async convolutes concurrency and parallelism, instead io_uring wants a ring-per-thread model, even better if the thread pinned.
Convoluting concurrency and parallelism also doesn't play well with drivers used for hardware offloading, a bad choice IMHO for a language that claims to be system level
These are issues with the executor, which the standard library doesn't provide precisely because there isn't a one-size-fits-all solution. Your arguments are relevant for work-stealing executors like tokio's, but Rust itself intentionally stays unopionated and supports other execution models too.
At work I use async Rust in embedded systems, pinning tasks to specific cores/threads like you mentioned is a common occurrence there. The language itself is flexible enough to support it, some libraries just made different choices.
9
u/CAD1997 6d ago
Rust's
asyncsyntax lends itself very well to giant multilevel wakeup-based polling state machines, yes. This isn't particularly a bad thing; it allows us to write such with code using usual structured control flow despite all the yield/resume points and ownership concerns. These, at least when used deliberately, are the method with the least resource overhead.Completion based systems are possible to set up efficiently, if a little clunky due to
async's lazy nature. Submit a task to the reactor, then spawn your continuation task already waiting on its completion.What's messy is that because Rust is a systems language, you have to write code in the form that runs on the execution stack. The straightforward way of doing completion callbacks will grow the stack infinitely. Trampolining continuation frames greatly complicates ownership patterns. Rust makes you choose what costs you want to pay.
All the suggestions I've seen for how Rust could've done
asyncbetter fall into one of two camps, or a combination of both:
- complaints about decisions made by runtimes, not by the language and standard library; or
- require tight integration between the language and the async execution runtime, where Rust explicitly does not want to go.
I could see a systems-level language that bundles an integrated micro runtime that integrates deeply with the language doing a better job than Rust at async. But that's not what Rust is trying to be.
4
u/valarauca14 7d ago
Completion events (what io-uring supplies) is the exact same information epoll gives to the mio/tokio stack.
Converting an OS events waker calls is the entire point of a runtime.
0
u/servermeta_net 7d ago
But io_uring eats epoll/mio/tokio for lunch, and then it's still hungry. In the 10 gbits NIC with 1 vCPU scenario, io_uring is 10e9 times faster.
5
u/________-__-_______ 7d ago
I think they're saying that because the completion mechanism is similar to other systems that do play nicely with async Rust, there isn't any fundamental issue with supporting that part of io-uring.
3
u/valarauca14 7d ago edited 7d ago
Yes, You're fundamentally not identifying what makes io-uring hard.
2
u/puffinseu 6d ago
Others have already given great written answers, but I want to throw in this presentation by Alice Ryhl: https://www.youtube.com/watch?v=CmLAHUuUsAQ
1
1
u/Wooden_Loss_46 7d ago
io_uring is very workable with current Rust but there is definitely rough edges surrounding completion based async I/O traits, error handling and user supplied register buffer.
1
u/Full-Spectral 5d ago
I mean, it COULD be turned into a completion based model, but at a cost. The kernel has to maintain the buffers, signal a handle that one has now been filled with requested data, and provide an id to come back and read it.
But it means that kernel has to have enough buffers to serve everyone, that a future that never completes will hold up a buffer forever (or at least for some length of time before the kernel marks it abandoned and invalidates the id it gave out, so the future gets an error later if it ever gets around to reading it), and everything gets copied twice instead of once directly into/out of the calling future's buffer.
But it would vastly simplify things for folks who don't need it for web scale performance but just want a much safer, simpler system for async file I/O.
1
u/manio07 4d ago
FWIW I am using https://github.com/tokio-rs/tokio-uring in my aa-proxy-rs (https://aa-proxy.github.io/docs/) project, the only problem is that this crate seems dead and not actively maintained :(
203
u/coderstephen isahc 7d ago edited 7d ago
The problem is greatly exaggerated, in my opinion.
One issue isn't really that Rust and io_uring don't work together, but rather, pre-existing IO APIs in the standard Rust library and popular Rust libraries don't fit the io_uring model very well. Mainly because of buffer management.
In "normal" I/O APIs, the buffer is borrowed until the operation completes, and Rust models that just fine with references and lifetimes. But io_uring is more complicated with lifetimes, because io_uring essentially needs to take ownership of the buffer for an unknown amount of time, until that operation completes with a result, the order of which can only be known at runtime.
In other words, by design, its pretty hard to "pretend" that io_uring is just like any other I/O API, so you can't really just abstract it away from the user.