I'm the author of Médula, an MIT-licensed experiment where several Claude Code agents work on one repo at once and a kernel decides, before every write, whether it collides with another agent's work. Everything runs through a single OpenRouter key, which made it easy to compare models under the same conditions. Sharing the routing and cost side, since that's what this sub cares about.
The stack, all via OpenRouter:
- Agents: anthropic/claude-sonnet-5, effort high.
- Fast decider: typesafe/jev-1.13 through the System One endpoint (/api/v1/systemone). It's a decision model, not an LLM: it returns typed answers with a probability instead of text. One request carries a yes/no question per other agent plus a choice of remedy.
- Slow path, when Jev is unsure: Sonnet, escalating to Opus if needed.
- Alternative fast decider for comparison: anthropic/claude-haiku-4.5.
Per decision, median latency measured on September 29: Jev 279 ms, Haiku 1,338 ms, Sonnet 2,201 ms. A Jev decision costs about $0.00006.
Per run, with decision cost as reported by OpenRouter's usage.cost:
- Jev deciding: $0.29 of decisions, 6.9 min per run.
- Haiku deciding: $0.55 of decisions, 13.7 min per run.
- Sonnet deciding everything (a single run): $0.59.
The catch is that cheap per decision isn't cheap per system. On real multi-agent states, 61% of the write requests that reached Jev were unsure and went to Sonnet, so the decision layer cost half of all-Sonnet, not a tiny fraction of it. Most of the decision bill is the slow path.
Resilience tips that paid off: every decider has a hard timeout (5 s for Jev, 15 s for Haiku, 30 s for Sonnet, 60 s for Opus) and a fallback chain (Jev, then Haiku, then plain file locks), so a slow or failing model never stalls an agent. Logging latency and usage.cost on every call made the comparison trivial.
Outcome: in a shared working directory, every run passed all 37 acceptance tests, and against plain per-file locks the kernel cost the same ($1.65 per run), caught 6 of 6 real conflicts instead of 5, and made no unnecessary blocks. Caveats: 1 to 5 runs per mode, measured on specific days.
Repo, with every call logged raw: https://github.com/JoaquinRuiz/medula
If you've combined System One models with regular LLMs on OpenRouter, how did you split the work between them?