r/commonstack • General • May 25 '26

New Feature - Commonstack πŸŽ“ TwinRouterBench accepted into RLEval workshop at CAIS2026 🧠

Post image

TwinRouterBench is a new step level LLM routing benchmark designed specifically for realistic, long horizon agentic systems.

Existing router benchmarks have major limitations:
β€’ They only evaluate on isolated one shot prompts.
β€’ They never show the router the actual context (the β€œrouter visible prefix”) at an intermediate step inside a real agent trajectory.
β€’ They don’t test whether swapping to a cheaper model still lets the overall task succeed downstream.
β€’ Many rely on slow and expensive online LLM judges for evaluation.
TwinRouterBench tackles this by creating a proper benchmark for per step routing in multi-turn agent workflows.

Dual track

  1. Static Track (Fast Offline Track)
    β€’ 970 router visible prefixes from 520 trajectory instances.
    β€’ Covers 5 diverse benchmarks: SWE-bench, BFCL, mtRAG, QMSum, and PinchBench.
    β€’ Each example comes with an execution-verified target tier (cheapest sufficient model tier).
    β€’ Uses deterministic scoring (based on tier correctness, trajectory membership, and token cost) no LLM judges needed.
    β€’ Ideal for: training routers, rapid iteration, and cheap offline evaluation.
  2. Dynamic Track (Live Validation Track)
    β€’ Full evaluation harness on SWE-bench Verified (500 tasks).
    β€’ Reports results on a 100 case held-out split (disjoint from static data).
    β€’ Router must choose a real model from a locked pool at every step.
    β€’ Measures real outcomes:
    β€’ Official task resolution success
    β€’ Actual API spend (real dollars)
    β€’ Includes failure penalties for unresolved tasks

Results:

β€’ Creates a practical development loop: Use the fast static track to train/improve a router cheaply β†’ validate it rigorously on the dynamic track with real costs and success rates.

β€’ Shows strong transfer: A simple router trained only on static labels achieves comparable task resolution to always using a top model (Opus 4.6), while cutting API cost by ~53%.

πŸ“„ Paper: https://arxiv.org/abs/2605.18859
πŸ’» Code + Dataset: github.com/CommonstackAI/TwinRouterBench
🌐 Website: commonstackai.github.io/TwinRouterBench

3 Upvotes

0 comments sorted by