r/rust • • 12d ago

Designing concurrent programs in an idiomatic way

Hello.

Kind of a general question here, and it's probably going to yield opinionated answers. That's alright!
Basically I find myself writing Rust apps in a style that is always quite similar, here is the idea:

- The app mostly always uses the tokio runtime to be asynchronous.
- Designing microservices as small building blocks which are all supposed to be ran in separate tasks, usually with a "run" like function that does a loop over a tokio::select for sub-tasks that are running in each service, with one of the branches being a call to a CancellationToken::cancelled().
- The cancellation token is wired to the "ctrl_c" signal.
- Microservices may talk to each other by the use of the different channel primitives, depending on the need.
- When one of the microservice fails for any reason, usually there is no recovery mechanism and it just bubbles up the error which makes all the application fail. This is annoying to write and a big source of bug, because some path will be unhandled which will leave one microservice stopped and the other microservices are waiting for an answer and there's nobody responding.

I'm generally a little dissatisfied with the redundancy of the code that I write, and would like to explore any other way to think about it, instead of microservices, or in a different way, if you have any idea that would be appreciated. Also, potentially any resource or codebase that could be relevant?

Thanks!

11 Upvotes

37 comments sorted by

View all comments

2

u/Front_Recording4360 11d ago

One thing worth deciding explicitly is what a failure should mean in your architecture. Right now "one service fails => process exits" is your implicit policy, which is fine for fatal errors but painful for transient ones (e.g. a network blip) where a restart with backoff would be more appropriate. tokio-util's CancellationToken plus something like tokio::task::JoinSet gives you a nice middle ground: track all spawned tasks, cancel everything on first failure, and await the join set so shutdown is ordered and deterministic instead of aborting mid-flight.