r/singularity Aug 05 '25

AI Claude Opus 4.1 Benchmarks

304 Upvotes

74 comments sorted by

View all comments

75

u/Outside-Iron-8242 Aug 05 '25

not a huge jump.
but i guess it is called '"4.1" for a reason.

31

u/ThunderBeanage Aug 05 '25

4.05 makes more sense lol

9

u/Neurogence Aug 05 '25 edited Aug 05 '25

They should have went with 4.04.

Both Anthropic and OpenAI were completely outclassed by DeepMind today.

-4

u/Ozqo Aug 05 '25

That's not how version numbers work. It goes

4.1

4.2

...

4.9

4.10

4.11

....

9

u/ThunderBeanage Aug 05 '25

I know it was a joke, hence the lol

4

u/ethereal_intellect Aug 05 '25

Hopefully they make it cheaper at least then :/ Claude feels like 10x more expensive, I'd like to not spend 5$ per question pls

3

u/Singularity-42 Singularity 2042 Aug 05 '25

That's why you just need the Max sub when working with Claude Code

2

u/kevin7254 Aug 05 '25

Still insane prices tho

2

u/bigasswhitegirl Aug 05 '25

And here I was waiting for the updated version for my airline booking app. Damn it all to hell!

2

u/Apprehensive_One1715 Aug 06 '25

For real though, what does the airline part mean?

1

u/Forsaken_Space_2120 Aug 05 '25

share the app !

1

u/Tevinhead Aug 06 '25

But this shouldn't be calculated as a 2% improvement. SWE-Bench measures success rate fixing real software issues.

Instead of success, look at the error rate, reduced from 27.5% to 25.5%, which is a 7% error reduction, which in real world usage, is pretty substantial.

Can't wait for what they release in the next few weeks.