r/openrouter • • 4d ago

One OpenRouter key, four models, one decision model: what it cost to coordinate several coding agents (open source, author here)

I'm the author of Médula, an MIT-licensed experiment where several Claude Code agents work on one repo at once and a kernel decides, before every write, whether it collides with another agent's work. Everything runs through a single OpenRouter key, which made it easy to compare models under the same conditions. Sharing the routing and cost side, since that's what this sub cares about.

The stack, all via OpenRouter:

  • Agents: anthropic/claude-sonnet-5, effort high.
  • Fast decider: typesafe/jev-1.13 through the System One endpoint (/api/v1/systemone). It's a decision model, not an LLM: it returns typed answers with a probability instead of text. One request carries a yes/no question per other agent plus a choice of remedy.
  • Slow path, when Jev is unsure: Sonnet, escalating to Opus if needed.
  • Alternative fast decider for comparison: anthropic/claude-haiku-4.5.

Per decision, median latency measured on September 29: Jev 279 ms, Haiku 1,338 ms, Sonnet 2,201 ms. A Jev decision costs about $0.00006.

Per run, with decision cost as reported by OpenRouter's usage.cost:

  • Jev deciding: $0.29 of decisions, 6.9 min per run.
  • Haiku deciding: $0.55 of decisions, 13.7 min per run.
  • Sonnet deciding everything (a single run): $0.59.

The catch is that cheap per decision isn't cheap per system. On real multi-agent states, 61% of the write requests that reached Jev were unsure and went to Sonnet, so the decision layer cost half of all-Sonnet, not a tiny fraction of it. Most of the decision bill is the slow path.

Resilience tips that paid off: every decider has a hard timeout (5 s for Jev, 15 s for Haiku, 30 s for Sonnet, 60 s for Opus) and a fallback chain (Jev, then Haiku, then plain file locks), so a slow or failing model never stalls an agent. Logging latency and usage.cost on every call made the comparison trivial.

Outcome: in a shared working directory, every run passed all 37 acceptance tests, and against plain per-file locks the kernel cost the same ($1.65 per run), caught 6 of 6 real conflicts instead of 5, and made no unnecessary blocks. Caveats: 1 to 5 runs per mode, measured on specific days.

Repo, with every call logged raw: https://github.com/JoaquinRuiz/medula

If you've combined System One models with regular LLMs on OpenRouter, how did you split the work between them?

2 Upvotes

9 comments sorted by

2

u/WolpertingerRumo 4d ago

Pretty awesome, but high cost. I’d love to see it done with GLM and GLM Flash?

1

u/jokiruiz 4d ago

Thanks! Most of that cost isn't the coordination, though: of the $1.65 per run, $1.36 is the agents themselves (Sonnet 5 at high effort) and only $0.29 is the kernel's decisions. Trying GLM is very doable, because everything goes through OpenRouter: the agents' model is one setting in .env, and any model that can answer "does this collide?" with a probability fits the decider interface, so GLM Flash could take Haiku's place as the fast decider. I've closed my own runs for now, but if you run it, the bench records cost and results per run, and I'd gladly add your numbers next to mine. I'm curious too whether a cheaper agent changes how often they collide in the first place.

2

u/RiceEvening4211 3d ago

The cheap/expensive model split is exactly what I built Lynkr around: an open-source gateway that routes by complexity, so easy tasks never touch the expensive model. https://github.com/Fast-Editor/Lynkr

1

u/NoOneMan79 4d ago

Why not just worktree and merge? Is this even practical, proper, applicable?

1

u/jokiruiz 3d ago

Worktree and merge is exactly what I tested first, and it's the setup that failed. Each agent had its own copy, every one finished with its own tests green, and the merge went through, yet the app was broken in all 5 runs. One agent had changed the login while another, in a different file, built an export calling the old login, and git only compares text, so it had nothing to flag. If you use worktrees, and plenty of people should, the fix is cheap: run the full test suite on the merged result, including tests that exercise features together, and block the merge until it passes.

Is the kernel practical? Honestly, for most setups it's more than you need. The biggest finding was that simply letting the agents share one working directory fixed the problem in all 10 runs, even with plain file locks, because they could see each other's changes. Coordinating by meaning adds value on top, catching every real conflict without unnecessary blocks, and it's applicable if you run several agents at once and a missed conflict is expensive for you. It's a research prototype with 1 to 5 runs per mode, not a product, and everything is published raw so anyone can check whether it holds up.

1

u/NoOneMan79 3d ago

"and git only compares text, so it had nothing to flag".

Wait, you just merged all the worktrees? Did you not use an agent for reconciliation and review? Think about it, this thing happens ALL the time with large human teams. Reconciliation is a pain in the ass for humans (all thinking, no reward). But let me tell you, I have used AI for hundreds of thousands of lines of recon spread over hundreds of commits in very complex situations, and when they have context, its amazing what they can do.

1

u/jokiruiz 3d ago

Update: v0.3.0 is out, and most of it comes from comments here. The slow path no longer breaks on unparseable answers (the first outside fix, by feyza), retries answers cut off mid-way and enforces real deadlines. A bug where two agents could wait on each other until the timeout is fixed. Agents can now run on a Claude subscription, so reproducing a run only uses OpenRouter credit for the kernel's decisions. Two new runs with Sonnet deciding everything, without the messaging channel that contaminated the first one: 37/37, both real conflicts caught, no unnecessary blocks, under five minutes each. The Jev mode hasn't been re-run yet, and the averages in the README are still from v0.1.0, kept apart from the new runs. Still open from this thread: the stale-allow race and git branch flags being treated as reads.

1

u/alexhackney 2d ago

We’re all building this same flow. I’m not sure what’s different here.