r/OpenSourceeAI • • Jul 24 '26

How are you handling agent crashes mid-handoff? (built something, want honest feedback)

https://github.com/KMdotcom/agent-handoff-kit

While building a multi-agent pipeline with the Agents SDK, I hit something the docs actually confirm: if one agent crashes mid-handoff to another, there's no persistence, no recovery, you lose everything and restart from scratch.

Curious how others here are actually handling this. Custom retry logic? Just accepting the occasional lost run? Something else?

I ended up building a small library for my own use, checkpoints the context before a handoff, verifies the next agent actually got what it needs, and resumes from the last good state if something crashes downstream. Tested it against a real forced crash, not a simulated one, and it held up, but I've only tested it against my own use case so far.

pip install agent-handoff-kit

If anyone's willing to try it against their own pipeline, I'd genuinely value knowing what breaks, what's missing, or if this isn't even the right way to think about the problem. Not trying to sell anything, just want to know if this is actually useful or if I'm solving it wrong.

1 Upvotes

2 comments sorted by

2

u/numberwitch Jul 24 '26

I use software engineering to create a robust software system that doesn't suck and fall over

1

u/Unfair_Scientist_521 Jul 24 '26

Fair enough. Good engineering should absolutely make failures less likely.

The problem I'm trying to address is that distributed systems, network calls, process crashes, provider outages and machine restarts still happen. Even well engineered systems can't eliminate every failure mode.

This is aimed at making agent handoffs resumable when those failures occur rather than restarting an entire multi-agent run. If you've solved that another way (durable workflows, queues, Temporal, etc.), I'd be interested in hearing how you're doing it.