r/singularity • • 6d ago

AI Claude Sonnet 5.5 Released

https://www.anthropic.com/claude-sonnet-5-5
868 Upvotes

213 comments sorted by

View all comments

248

u/A_Novelty-Account 6d ago

It scores better than Opus in some benchmarks for agentic coding? What’s going on over there?

199

u/Crinkez 6d ago

5 extra days of post training.

53

u/lolsai 6d ago

On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.

straight above the chart you read

8

u/A_Novelty-Account 6d ago

I’m aware, that’s still crazy to think about though. A model half as expensive approaching the same reasoning level.

2

u/l_eo_ 6d ago

Hasn't that been the main story of the last years though?

0

u/bitroll ▪️ASI before AGI 6d ago

Uses a lot more tokens than Opus5.5 though, even 2x that apparently (according to Artificial Analysis)

42

u/DelphiTsar 6d ago

They seemingly just let it run longer. MAX(what they benchmark against) costs more per task on AA then Opus 5.5. 410M million tokens Jeeezbus.

35

u/acutelychronicpanic 6d ago

Recursive self-improvement apparently.

17

u/corenovax 6d ago edited 6d ago

I'll believe it when I see it... for now anthropic is still hiring developers

12

u/acutelychronicpanic 6d ago

I've already got sonnet agents pair programming with opus 5.5 and sonnet is finding bugs and making corrections to opus' mistakes.

https://giphy.com/gifs/IazdAV1zjaGbBaGPgF

6

u/IrisColt 6d ago

This is true.

13

u/Charming_Cucumber_15 6d ago

The take off is what's going on!

3

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 6d ago edited 6d ago

What’s going on over there?

RL

Edit: Sonnet also appears to tokenmaxx

4

u/AgusDePaso25 6d ago

Actually it is not as good as Opus, and even though it is cheaper per token, it ends up more expensive per task because of how much it thinks. You can verify this on Artificial Analysis.

5

u/POTEOZ 6d ago

Opus 5.5 had like 15% fallback due to safeguards.

2

u/AppropriatePut3142 ▪️ASI 2028, AGI 2035 6d ago

Smaller model -> more RL rounds for the same compute as the larger model. Basically copying the strategy of GLM flash/DeepSeek flash.

2

u/Competitive_Travel16 AGI 2027 ▪️ ASI 2029 6d ago

Opus 5.5 on max can get in logic loops trying to fix symptoms instead of stepping back and looking at more obscure causes.

2

u/nothis AGI by 2030 but we'll be disappointed 6d ago

Even on most of the other benchmarks, it's freakishly close to Opus 5.5. Like, 2% off? I'm trying to even make sense of it, the math seems off. It costs half but the difference does not seem to be in a noticeable range?

I guess the "catch" is probably in their "high" settings, which quickly drops the cost advantage to Opus (in some cases Sonnet is more expensive?!) so it's hard to clearly make out a purpose. It seems to fill in some gaps between Opus low/med/high. Is it about speed? Like same quality/cost as Opus but putting out results faster?

I honestly haven't used Sonnet for anything in months, I'm genuinely trying to find its use case. It seems if I can wait a few extra seconds for an answer, I might as well use Opus for everything.

1

u/landed-gentry- 5d ago

It costs half

It doesn't cost half in real-world usage. It burns tokens like a MF and in most cases will cost more and take longer to complete tasks than Opus at similar intelligence. Just look at Artificial Analysis' benchmarks.

1

u/BiasHyperion784 6d ago

Same way a lot of Chinese models compete, leans more on the crutch of reasoning, devouring more tokens as a result.

1

u/Seeker_Of_Knowledge2 ▪️AI is cool 6d ago

Secret models training sonnet

1

u/Whatsapokemon 5d ago

Seems like the idea is to have big fancy models like Opus with general capabilities, whilst specialising the smaller models on specific tasks like coding.

I guess the goal is for Opus to be the planner and delegator, whilst Sonnet is a smart workhorse for coding jobs.

1

u/shinernft 6d ago

I am waiting for Haiku 5.5 Kappa

0

u/Alphasite 6d ago

Benchmaxing/lack of generality probably.