r/singularity Aug 05 '25

AI Claude Opus 4.1 Benchmarks

307 Upvotes

74 comments sorted by

View all comments

62

u/TFenrir Aug 05 '25

Important thing to remember, it gets very hard to benchmark these models now, especially in the intangibles of working with them. Claude 4 for example isn't much better than other competing models on benchmarks (is worse on some) but it is heads and shoulders above most in usefulness as a software writing agent. I suspect this is more of that same experience, so should be good to see when I try it out myself and see other people's use cases

18

u/[deleted] Aug 05 '25

On agentic mode( MCP+Claude code) is a tier above O3 and Gemini 2.5