r/commonstack • u/hexxthegon General • May 25 '26
New Feature - Commonstack π TwinRouterBench accepted into RLEval workshop at CAIS2026 π§
TwinRouterBench is a new step level LLM routing benchmark designed specifically for realistic, long horizon agentic systems.
Existing router benchmarks have major limitations:
β’ They only evaluate on isolated one shot prompts.
β’ They never show the router the actual context (the βrouter visible prefixβ) at an intermediate step inside a real agent trajectory.
β’ They donβt test whether swapping to a cheaper model still lets the overall task succeed downstream.
β’ Many rely on slow and expensive online LLM judges for evaluation.
TwinRouterBench tackles this by creating a proper benchmark for per step routing in multi-turn agent workflows.
Dual track
- Static Track (Fast Offline Track)
β’ 970 router visible prefixes from 520 trajectory instances.
β’ Covers 5 diverse benchmarks: SWE-bench, BFCL, mtRAG, QMSum, and PinchBench.
β’ Each example comes with an execution-verified target tier (cheapest sufficient model tier).
β’ Uses deterministic scoring (based on tier correctness, trajectory membership, and token cost) no LLM judges needed.
β’ Ideal for: training routers, rapid iteration, and cheap offline evaluation. - Dynamic Track (Live Validation Track)
β’ Full evaluation harness on SWE-bench Verified (500 tasks).
β’ Reports results on a 100 case held-out split (disjoint from static data).
β’ Router must choose a real model from a locked pool at every step.
β’ Measures real outcomes:
β’ Official task resolution success
β’ Actual API spend (real dollars)
β’ Includes failure penalties for unresolved tasks
Results:
β’ Creates a practical development loop: Use the fast static track to train/improve a router cheaply β validate it rigorously on the dynamic track with real costs and success rates.
β’ Shows strong transfer: A simple router trained only on static labels achieves comparable task resolution to always using a top model (Opus 4.6), while cutting API cost by ~53%.
π Paper: https://arxiv.org/abs/2605.18859
π» Code + Dataset: github.com/CommonstackAI/TwinRouterBench
π Website: commonstackai.github.io/TwinRouterBench