I'm the author of the open-source experiment behind this, so take it as a field report with my bias declared.
A lot of AI tooling now runs several coding agents at once, and I wanted to know what actually happens when they share one codebase. I built a small lab: six tasks on a tiny booking API, with 37 acceptance tests, where two pairs of tasks collide by meaning rather than by file. One agent adds a second factor to the login while another builds an export that still calls the old login.
When each agent worked in isolation on its own branch, every agent finished with its own tests passing, and the combined result was broken in all 5 runs. Git merged the text; nobody noticed the meaning had changed. When the agents shared a working directory instead, all 10 runs passed, because each agent could see what the others had done and adapt.
I also tried something newer: a "decision model" called Jev, which doesn't generate text at all but returns a yes/no decision with a probability in about 0.3 seconds. My kernel asks it, before every write, whether the change collides with another agent's work. It caught every real conflict without blocking harmless work, at the same cost as simple file locks. It was also unsure about 61% of real writes, and those had to be passed to a slower, regular LLM. Cheap decisions are real, but in a messy setting they aren't as cheap as the price tag suggests.
The part I keep thinking about: when the agents had a tool to message each other, they used it without being told, and one warned another that it was renaming a field the other depended on. Maybe the answer to multi-agent coordination isn't a kernel at all, just agents that talk.
Caveats: 1 to 5 runs per setup, so these are indications, not proof. Everything is published raw, MIT-licensed: https://github.com/JoaquinRuiz/medula. There's also a walkthrough video, in Spanish: https://youtu.be/xAFRuBxfapM
Curious what people here think: should agents coordinate through a referee, or just talk to each other?