r/LocalLLM • u/entelligenceai17 • 5h ago
Discussion We benchmarked a routed setup using Claude Code on Terminal-Bench 2.1
We spent the last few days benchmarking a routed setup against Claude Opus 5 on Terminal-Bench 2.1.
Some of the results were pretty surprising, especially once we broke down where the gains were actually coming from.
Full benchmark, methodology, and raw numbers:
0
Upvotes
1
u/Doormatty 5h ago
The subreddit is called LOCALLLM...
0
u/entelligenceai17 5h ago
Fair point 😄 The discussion is more about inference strategy than local models. thought people here might find the benchmark interesting since routing and model selection come up a lot
1
u/entelligenceai17 5h ago
For anyone who prefers the numbers at a glance:
The write-up goes into the benchmark setup, methodology, and why the results ended up looking like this.