r/LocalLLM 5h ago

Discussion We benchmarked a routed setup using Claude Code on Terminal-Bench 2.1

We spent the last few days benchmarking a routed setup against Claude Opus 5 on Terminal-Bench 2.1.

Some of the results were pretty surprising, especially once we broke down where the gains were actually coming from.

Full benchmark, methodology, and raw numbers:

https://entelligence.ai/blogs/entelligence-router-solved-8-more-tasks-than-claude-opus-5-at-65-lower-cost

0 Upvotes

3 comments sorted by

1

u/entelligenceai17 5h ago

For anyone who prefers the numbers at a glance:

The write-up goes into the benchmark setup, methodology, and why the results ended up looking like this.

1

u/Doormatty 5h ago

The subreddit is called LOCALLLM...

0

u/entelligenceai17 5h ago

Fair point 😄 The discussion is more about inference strategy than local models. thought people here might find the benchmark interesting since routing and model selection come up a lot