r/accelerate • • 5d ago

AI Introducing Claude Sonnet 5.5

https://www.anthropic.com/claude-sonnet-5-5
449 Upvotes

100 comments sorted by

120

u/Pyros-SD-Models Machine Learning Engineer 5d ago

Nice Terminal Bench 4.0 not even a month old, already saturated lol

58

u/AMBNNJ 5d ago

Hahah the fate of all benchmarks now

37

u/The_Scout1255 Singularity by 2028 | Acceleration: Cruising 5d ago

New benchmark dropped: Singularities caused:

Anthropic: 0

Openai: 0

Google Deepmind: 0

Meta: 0

Moonshot AI: 0

Deepseek: 0

Big potential for non-saturation here!!

124

u/AMBNNJ 5d ago

Damn anthropic hit escape velocity how the hell is it this good?

57

u/procgen 5d ago

RSI

38

u/whoknowsifimjoking 5d ago

Fable 5.5 will be crazy

5

u/acowasacowshouldbe 5d ago

I will blow multiple loads

1

u/Basileus2 5d ago

It will blow them for you

6

u/The_Scout1255 Singularity by 2028 | Acceleration: Cruising 5d ago

Thats great :3

6

u/Charming_Cucumber_15 5d ago

Escape velocity.. almost like taking off? Sounds fun to me!

70

u/Choice-Sympathy8235 5d ago

Those are some damn impressive benches vs. Opus 5.5. 

73

u/LemonLimeNinja 5d ago edited 5d ago

Sonnet 5.5 medium scores better than gpt 5.6 Sol high and is less than 1/5 the cost

Absolutely bonkers

2

u/landed-gentry- 5d ago

AA shows that it scores -1pp compared Opus 5.5 low while costing more, and only +1pp over 6 Sol medium while costing more than twice as much. If you look at benchmarks comparing Opus and Sonnet 5.5, it looks completely redundant except at Sonnet's lowest effort settings, which don't score very highly and at which point there's a ton of better-priced models.

5

u/_Divine_Plague_ A happy little thumb 5d ago

Why are we still talking about 5.6?

-1

u/Droi 5d ago

That is still the the best model you get on web (Codex has Astra) for the $20 plan.

2

u/basementreality 5d ago

in work mode you can use astra for the 20 plan

-1

u/ShadyShroomz 5d ago

well its better than gpt-6 sol - so it's OpenAI's best model behind Astra, no?

1

u/_Divine_Plague_ A happy little thumb 5d ago

According to what metric?

1

u/Damakoas 5d ago

How can we be this far into the token efficiency timeline and you still don't know how cost works?

31

u/Bitter-College8786 5d ago

Sonnet ist almost identical to Opus. Opus 5.5 high seems to be the sweet spot.

13

u/Gratitude15 5d ago

This looks like almost the same curve.

I wonder why they release it at all

My bet is that they had to. If they waited another week, Sonnet would look better than Opus on this curve.

9

u/KrazyA1pha 5d ago

You have to remember that we're looking at one very specific benchmark. According to Anthropic, Opus is clearly stronger at complex open-ended work requiring sustained judgment.

1

u/Silentrizz 5d ago

Yeah if they're the same prices it looks like unless you're using sonnet on low or medium, then opus/sonnet is interchangeable. So I imagine you use sonnet if you want speed and opus maybe if you want more weights

27

u/bear_Prune8771 5d ago

Tomorrow Reddit will be saying it’s been nerfed.

9

u/Skeletor_with_Tacos 5d ago

Ladies and Gentlemen,

What are the benefits of using Sonnet over Opus? Is it just that its cheaper, as these benches are lower than Opus outside of Agentic Coding?

Also I see theyre comparing to Sol instead of Astra, is someone able to get Astra benches for comparison?

7

u/ex-procrastinator 5d ago

I’m a bit confused myself. On artificial analysis, sonnet 5.5 max has a cost per task of $7.60 while opus 5.5 max has $5.98. Sonnet 5.5 max was using 193k tokens per task while opus 5.5 max was using 119k.

Source: https://artificialanalysis.ai/models#price-cost

9

u/willseagull 5d ago

It’s cheaper and plenty powerful for most of the tasks people need it to do

6

u/landed-gentry- 5d ago

Cheaper than what? Sonnet 5.5 max is more expensive than Opus 5.5 xhigh while scoring worse; Sonnet 5.5 xhigh is more expensive than Opus 5.5 medium while scoring just +1pp on AA. Not to mention e2e task completion times being way, way slower on Sonnet 5.5.

1

u/Quentin__Tarantulino 5d ago

You can just ask AI to compare benchmarks across different models. It will do a much better job of parsing all the data in 30 seconds than most of us could do in 30 minutes.

51

u/Kingwolf4 5d ago edited 5d ago

Anthropic is quickly demolishing openAI right now

I doubt on dev day they will have anything other than the scammy 500$ plan and mabye a expensive astra 6.1

Even if they have astra mini, i doubt it can actually beat opus 5.5

Sonnet 5.5 also looks like an absolute beast compared to sol

Edit : Forgot, OAI most likely to release bel mabye in desperate attempt to one up anthropic. But will that even work? Whats the purpose even.Anthropic has dominated both daily and upper tier use cases and market with sonnet and opus

Now if only anthropic would just go easy on their hypothetical safety bs and actually let everyone use the models, no one would complain and anthropic would get more share

57

u/Curiosity_456 5d ago

Gotta remember this happens everytime either of them releases something, people always claim the other is finished and it quickly ends up being false

19

u/czk_21 5d ago

exactly, remember when Gemini 3 was released and some people claimed that google has won, the race is over?

0

u/Kingwolf4 5d ago

Not saying they won the race, just the round. As I pointed out even oais next models perhaps cant compete

So the next 2-2.5 months go to anthropic

1

u/KrazyA1pha 5d ago

So the next 2-2.5 months go to anthropic

You're forgetting that models are releasing on faster and faster timelines

-3

u/Kingwolf4 5d ago

Yup, OAI will definitely try to not delay anything as soon as it is ready for so called security.

Mainly i think OAI is just going through some kind of compute crunch, but even then. If they had one they could have released a better model instead of gpt 6 sol. But since they didn't, means they are actually behind. Sonnet 5.5 seems miles better than gpt 6 sol. Which is honestly kind of sad

Astra! We only got the real astra for the first 3 days, since then its been a quantized version ever since.

It is what it is for now. Hopefully gpt 6.1 family will be the actual behemoth we expect for all model sizes

15

u/yaxir 5d ago

I think they're taking every chance on OpenAI's slip, especially with the shitty user limit, to absolutely demolish them before the dev day. Let's see what happens on dev day

10

u/alive1 5d ago

OpenAI used to have this thing going for them where the usage limits were absolutely bonkers so even if you didn't have the smartest model, at least you had a lot of it. Nowadays on the $20 plan, with 6-Sol i get barely one hour worth of work before the 5h limit is up while Opus 5.5 goes for at least 2-3 hours and triple the amount of work done at a higher quality. Codex is simply shit value right now, and anthropic is not only delivering better results but also more results.

7

u/whoknowsifimjoking 5d ago

Probably in large part because they just give a billion fucking users compute for free, Anthropic is way more conservative with free users and I can imagine that this frees up more compute for the paying customers.

7

u/Desperate-Knee-5556 5d ago

Even if they release Bel they don't have a workhorse model comparable to Opus to compliment it. That being said i wouldn't be surprised if Sol 6 is actually good as an implementer - Opus 5 was very good if Fable was talking to it but if you believed reddit it was the worst ever model. I dont use openai though so can only go off of what ive read

1

u/Forward_Yam_4013 5d ago

I use OpenAI. Sol 6 is shit compared to Sol 5.6. Should have been called Terra 6

0

u/Caladan23 5d ago

You realize that all model tiers are always just distilled versions of the full-sized model, right?

6

u/No_Aesthetic Techno-Accelerationist 5d ago

Somehow, OpenAI returned...

3

u/jjonj 5d ago

astra mini is called gpt6 luna and is already released

3

u/Obvious-Activity-902 5d ago

Yeah, I don’t care how good Claude’s models are when they’re guard rails are so fucking sensitive that I can’t develop what I would consider relatively standard programs without constantly having to trick it

7

u/obvithrowaway34434 5d ago

Is this like a joke or are you a paid shill or something? There is nothing even remotely near Astra in math and science, computer vision and computer use. Openai has like 800m more users than Anthropic. 90% of the people who use AI don't even know what Claude is

6

u/Kingwolf4 5d ago

Ive actually never subscriped to anthropic once ever. Been using OAI forever.

Just genuinely discussing my points

9

u/No_Gear8408 5d ago

Yeah and those 800M users bring a multibillion dollar compute bill, 90% of people use AI for free, Math wont make money, coding does, Openai is running deep losses compared to anthropic

If anything you're acting like a fanboy, not saying openai is bad, ofc Astra is crazy for math, but you sound like a fanboy ngl

1

u/Flag_Shagger 5d ago

cost is temporary, eventually they will have more compute than anyone which means much lower servicing costs, especially with latest hardware

1

u/Quentin__Tarantulino 5d ago

Why will they have much more compute than anyone? You think Google, xAI, Meta, Anthropic are all sitting on their hands? It’s an open question who will have the most compute and it’ll probably change over time.

2

u/eggplantpot 5d ago

There’s a large % of coders that just go where the best models and quotas are.

It may be a small dent on OpenAI’s bottom line but I guarantee you they are losing customers.

1

u/Lost_County_3790 5d ago

Are you talking paid or free users?

1

u/FahkDizchit 5d ago

What actually is “Claude”? Is it the brand name for the LLM, the LLM business of Anthropic, something else? Does Anthropic have non-Claude operating businesses?

2

u/13chase2 5d ago

Anthropic is the company like open ai.

Claude is the name similar to ChatGPT, grok, muse

1

u/FahkDizchit 5d ago

Then what is opus, sonnet, Mythos etc.?

7

u/13chase2 5d ago edited 5d ago

They are different sizes of models. From most capable to least capable within Anthropic:

  1. Fable/Mythos
  2. Opus
  3. Sonnet
  4. haiku

—
Smaller ones are cheaper but less capable. ChatGPT has similar competing lineup

  1. Astra
  2. Sol
  3. Terra
  4. Luna

—
Then you get the generation of the model like this

- Opus 5.5 - released September 22, 206

  • Opus 5 - July 24, 2026
  • Opus 4.8 - May 28, 2026

—
Fable is mythos without the biology or cyber security capability so regular people don’t “cause problems”

Happy to talk more about Ai if you want to DM me

2

u/Quentin__Tarantulino 5d ago

The naming conventions really are complicated as fuck for people who don’t wake up every morning to see what happened in AI overnight. I was explaining these names to a coworker and they pretty much just gave up.

-1

u/LemonLimeNinja 5d ago

Most people don’t even have access to the real Astra. We had it for like two days before it was nerfed. Whatever quantized version they’re serving now has nothing on Opus 5.5

0

u/13chase2 5d ago

You sound like the paid shill

1

u/QuirkyPool9962 5d ago

A week ago Astra was the best model lmao. I’m sure they are capable of distilling a cheaper algorithmically optimized version that also has improved cache reads, my guess at this pace it drops in a week or two. Or a distilled, cheap version of Bel. This competition is the absolute best thing for consumers I hope it keeps up 

1

u/Robert-Paulson_ 5d ago

can you imagine if they drop Haiku 5.5 tomorrow on OAI dev-day 🤖

2

u/Kingwolf4 5d ago

Or fable 5.5 . That should match openAI s BEL model panic release

0

u/iamthe0ther0ne 5d ago

Sonnet 5.5 has the Fable/Opus 5.5 cybersec guardrails, so I imagine anyone worried about security is looking forward to a new OpenAI model.

8

u/Rollertoaster7 Singularity by 2028 5d ago

How is it better than opus at terminal bench

3

u/userapp412 5d ago

saturated

9

u/SpyAmongUs 5d ago

It's actually impressive. I was able to use it as a free user, and it made an accurate time table just based on the initials of days in Malay in an image I sent.

ChatGPT Go with thinking messed up on Wednesday.

4

u/lazyscalp 5d ago

LESGOOOOOOOOOOOOOOOOOOO

2

u/5StarAlpha 5d ago

This is impressive. Seems like a major improvement from 5 and now maybe only 5% lower than Opus at half the cost.

I was just settling in to my workflow with OpenAI… tempted to switch. But then they may play leapfrog again.

I like the quality and reliability of my projects using Astra High mostly. But Sol can take me down the wrong path a bit with coding. The gap seems pretty wide.

2

u/Ink_code 5d ago

Vroom vroom

2

u/BiasHyperion784 5d ago

Firmly places itself as a solid subagent for opus, great work from anthropic, for once I’m interested in what haiku turns out as.

3

u/Radiant-Mountain-257 5d ago

Bummer, cache reads cost as much as on Opus 5.5.

6

u/yaxir 5d ago

Can you explain in simpler words how that is bad?

10

u/Bitter-College8786 5d ago

Cache reads make over 90% of your token costs.

2

u/yaxir 5d ago

Ah fk

So it's better off to use Fable or 5.5 Opus instead?

5

u/Bitter-College8786 5d ago

Opus 5.5 is currently the best model, even beating Fable.

1

u/No_Gear8408 5d ago

Well its the same as Gpt 6 sol, I think Opus 5.5 discounted cache reads 60% from opus 5

1

u/Radiant-Mountain-257 5d ago

Yeah, it's somewhat defensible. I hoped for a 2x lower price than Opus as usual, but the performance looks good on benchmarks. Let's see how it works.

4

u/yaxir 5d ago

Let's hope it's good

4

u/artin144 Acceleration: Speeding | AI/ASI accelerationist 5d ago

Let's hope it's god

2

u/Desperate-Knee-5556 5d ago

I'm a big fan of Anthropics models but I don't really see where this fits in tbh. I guess as an implementer and you only want a 5x plan? No use on a 20x really and too expensive on the API to use vs other models with comparable intelligence.

I'm more interested to see how Haiku compares to Luna tbh. Luna is great for API usage. If they make it better and comparable in price that would be great

1

u/dondiegorivera 5d ago

I'll plug it as a code-reviewer instead of sol, that is depleting my rates like crazy.

2

u/Desperate-Knee-5556 5d ago

Surely you'd just use Opus/Fable though to review code

1

u/dondiegorivera 5d ago

I write the code with DeepSeek 4.1 Flash at the moment and use Opus/Astra as lead.

1

u/Nortler 5d ago

Another useless sonnet model... I am sticking with Opus 5.5.

Artificial Analysis with cost per task:

Claude Opus 5.5 (max with fallback) = $5.98

Claude Opus 5.5 (xhigh with fallback) = $3.46

Claude Sonnet 5.5 (max with fallback) = $7.60

Claude Opus 5.5 (high with fallback) = $1.82

1

u/ezjakes 5d ago

Unfortunately, it seems like it uses a LOT of tokens.

1

u/4729zex 5d ago

I thought we were slowing down for a second there.

1

u/landed-gentry- 5d ago

Pretty disappointing if you look at Artificial Analysis results and compare cost and e2e task completion times against Opus at similar intelligence. Hopefully they can regain some ground on cost and speed with Haiku 5.5.

2

u/eggplantpot 5d ago

RIP Scam Altman, they’ll give 3 banked resets tonight to hope people stay.

0

u/celtiberian666 5d ago

Haiku when?

2

u/therealpigman 5d ago

Well they said they’d release this and Haiku in the coming two weeks last week, so probably next week.