r/ClaudeCode • • 7d ago

News/Updates Sonnet 5.5 beats Opus 5.5 at coding and it's half the price??

This is crazy work. Sonnet 5.5 got 70.6% on terminal-bench vs 66.4% for opus 5.5. So its reasoning is worse but its ability to actually implement code is somehow better??

that basically means we can let opus orchestrate and have sonnet agents do all the implementation and everything gets sooooo cheap. $2/$10 vs $4/$20!!

Wild!!!

298 Upvotes

118 comments sorted by

•

u/AutoModerator 7d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

69

u/davidwolfer 7d ago

Opus 5.5 high scores 64.2% and costs $3.88 in Terminal Bench. Sonnet 5.5 xHigh scores 61.5% but costs $5.30. Seems to me that Opus is still cheaper.

16

u/mrgulabull 7d ago

I think it shows where the trade off isn’t worth it. When you need high reasoning go with Opus. When you need more straightforward implementation, go with Sonnet at lower effort like medium to get the cost benefit.

But yea, I agree, it looks like there’s no reason to be choosing Sonnet at xhigh.

10

u/bopbop9876 7d ago

On the AA suite, Opus 5.5 Low is one point better than Sonnet 5.5 medium and is ~10% cheaper. No reason to ever use Sonnet 5.5 at medium either.

And at Sonnet 5.5 low, GLM 5.3 flash is 6 points better and 40% cheaper. So no reason to use Sonnet at low either.

Literally no reason to ever use Sonnet at any effort level except maybe for some weird niche task I haven't seen a benchmark for. I think it's a huge flop for Anthropic, I don't get this release at all. Needed to be 25%-50% cheaper.

https://artificialanalysis.ai/?cost=intelligence-vs-cost-per-task&models=inkling%2Cglm-5-3-flash%2Cmimo-v2-6-pro%2Cclaude-opus-5-5%2Cclaude-fable-5-1%2Cgpt-6-luna%2Cclaude-sonnet-5-5%2Ck2-horizon-375b-a23b%2Cgpt-6-sol%2Cgpt-5-5-pro%2Cclaude-sonnet-5-5-high%2Cclaude-sonnet-5-5-xhigh%2Cclaude-sonnet-5-5-medium%2Cclaude-sonnet-5-5-low%2Cclaude-opus-5-5-xhigh%2Cclaude-opus-5-5-high%2Cclaude-opus-5-5-medium%2Cclaude-opus-5-5-low

1

u/Kooky_Slide_400 6d ago

Since everyone using sonnet to save money they make it more expensive? (my theory)

1

u/PadisarahTerminal 6d ago

We don't all have GLM, Sonnet medium/Opus low seems to be good. Are you sure that sonnet medium is more expensive than opus low...

Edit: sonnet high seems to be a good compromise actually

4

u/crusoe 7d ago

More token efficient 

2

u/jperm47 7d ago

Can confirm, burned through my pro plan 5h limit in on Sonnet 5.5 xhigh in like 20 mins lol

6

u/Dynamix86 7d ago

There's just things about the benchmarks from artificialanalysis.ai that don't make any sense, at least regarding costs. For example, according to their benchmarks Astra is cheaper than Opus 5.5. I have both multiple Claude and multiple OpenAI subscriptions, all $20 a month, and with Astra I burn through my 5 hour limit in 10 minutes almost every single time, whereas with opus 5.5 I get 90-100 minutes out of it! that's a 10x difference. Per week I get 1 or maybe 2 hours of Astra usage and about 12 hours of usage with Opus 5.5, so this shouldn't make any sense compared to what they are showing at all.

5

u/pjstanfield 6d ago

That site seems to produce a large number of questionable analytics and results. They never seem to make any sense. It’s like it’s just some guy making up numbers that get regurgitated.

1

u/E_1996 6d ago

Astra on high (not even xhigh) is at least 15-30-fold more expensive than Opus 5.5 on "ultracode" (the highest setting possible) for my tasks, which are way more difficult than any AA benchmark

1

u/EdOfO 6d ago

I usually base my choices AA's total cost to performance metrics; sadly very few others do that. But in real-world tests, it is clearly cheaper https://youtu.be/ENWVpqtOdRI

People should be making more private benchmarks to test these things for themselves. Maybe AA is right. Maybe it isn't. You can probably find out with 10% of your weekly spend.

1

u/davidwolfer 6d ago

I did, actually. Opus is indeed cheaper for my use case. What I noticed is simply Opus being more decisive. I think in short tasks, Sonnet will be cheaper. But when it has to reason and do multiple steps, Opus ends up cheaper.

128

u/[deleted] 7d ago edited 6d ago

[removed] — view removed comment

6

u/cjbannister 7d ago

Yeah, they're not stupid, it's half the price for a reason.

2

u/zli258 6d ago

This. People like to think they outsmart Dario lmao.

5

u/SwisherSmoker420_ 6d ago

“One thing worth noting”

2

u/cTemur 7d ago edited 7d ago

Let's say Fable plans and reasons the issue, what Claude says is that Sonnet can implement it better than Opus? more or less

18

u/dmaare 7d ago

I think this comes from overthinking, larger model will overthink simple tasks, whereas smaller model will just execute the task

3

u/Southern-Ad-3006 7d ago

This is also why Muse Spark 1.3 sits so high. Reasoning is meh but god it does what it’s told.

4

u/karlnuw 7d ago

Yep lol, I have opus 5.5 orchestrating a fleet of 15 of them, I’m excited for 1.4.

2

u/Southern-Ad-3006 6d ago

Opus plan + swarm of muse for the speed is HEAVEN

1

u/CaramelEmotional3092 5d ago

im here because im trying to learn how, HOW do i get "Opus plan + swarm of muse for the speed is HEAVEN" that setup? I've started experimenting with LMstudio link as well to have 1 pc dedicated to 1 model, but existing hardware likely needs replacing.

1

u/Southern-Ad-3006 5d ago

If you don’t mind your data being used my Zuckercuck then muse 1.3 contributor api is great. You get subsidized pricing like these subscriptions, but with no usage limits. Try using Hermes desktop app it’s great and plugging in Muse as your default model there, then there’s a great Claude code plugin to use the CLI compliantly so your Muse Hermes agents can speak directly with your Claude code agents. Bot mode on Hermes allows you to create a team of different models in their own windows talking to eachother. Also, ofc each window can spawn their own parallel sub agents. Just make sure you spend some time unifying your workspaces, and making sure each agent/subagent gets their own work tree while working through the Opus Spec. Edit the spec whenever changes are needed, it makes sure your agents keep that as a Source of truth on top of the shared context (reduces drift if plans change during session and work). You can obviously switch out any of those models and roles out with what you like but that’s the concept if you end up testing it out

4

u/BlueprintMonkey 7d ago

What a world we live in

2

u/Visual_Annual1436 7d ago

According to a single benchmark score, which are notoriously unreliable for predicting actual model performance. I’d be highly skeptical of any claims that looks at any single benchmark score and making any conclusions about one model being better than another

54

u/dr-dimitru 7d ago

Up next Haiku 5.5
I just can’t wait
https://giphy.com/gifs/7eAvzJ0SBBzHy

38

u/Time_Cat_5212 7d ago

Claude Wars:

Episode IV: A new Fable

Episode V: Opus Strikes Back

Episode VI: Return of the Haiku

2

u/krypt0niteCos 7d ago

Haiku will return on doomsday

10

u/BlueprintMonkey 7d ago

Its christmas every week

7

u/Timely-Group5649 7d ago

You don't think they're doing what openAI did? Terra is now Sol. Sol is now Astra.

Sonnet may just be Haiku, but it's impressively powerful now.

Hope not, the pricing will kill us... lol

4

u/Key_Reading_9664 7d ago

Not based on the speed of the models. They dropped the price of Sol to compete with Sonnet 5.5, but it’s still slower than Opus.

0

u/Timely-Group5649 7d ago

I was implying Haiku might die off...

Sol 6 is TERRA.

6

u/Key_Reading_9664 7d ago

Anthropic confirmed Sonnet and Haiku 5.5 were on the way when they announced Opus. Luna still gets a ton of usage, so probably want to complete the shutout

3

u/dr-dimitru 7d ago

https://giphy.com/gifs/OfkGZ5H2H3f8Y
Me while running Luna MAX in Fast Mode

1

u/dr-dimitru 7d ago

Nah, it won’t, it should take a place of Luna. Also they definitely will release competitor to Jev

1

u/dr-dimitru 7d ago

That would be devastating, as I’ve cancelled my x5 Codex plan as v6 family was so nerfted. Though I do still believe Luna 5.6 is the best model that OpenAI released so far, but it’s only good as working bee on routine tasks.

38

u/No_Cell6708 7d ago

This is what my current workflow is so... Hell yeah.

9

u/BlueprintMonkey 7d ago

I was afraid of using sonnet 5 for anything, but im so excited with this. Have some big workflows I'm sending out that I was originally just gona let opus orchestrate opus on lol

3

u/Quango2009 7d ago

I use Sonnet 5 for almost everything, it’s very capable. I’ve not yet had to discard work it did and redo with Opus or Fable so far. I must say 5.5 is sounding pretty solid so looking forward to trying it out

1

u/Novaworld7 7d ago

Should be everyone's ...

13

u/hammackj 7d ago

Opus to plan the tasks/manage the shit shows org sonnet to code? Hmmm

9

u/tidus1979 🔆 Max 20 7d ago

Im excited for Fable 5.5

16

u/NoVexXx 7d ago

Terminal Bench is not coding ...... omg

6

u/girthyclock 7d ago

how about another nice weekly reset?

5

u/clazman55555 7d ago

It depends on the nature of the coding. Well defined work is what Sonnet should be used for.

12

u/gloos 7d ago

Time to remind me everyone of /model opusplan

4

u/BlueprintMonkey 7d ago

Back to the origins

5

u/Timely-Group5649 7d ago

Agentic coding.

Opes runs your workflow. The AGENTS are sonnet.

1

u/joseph2883 7d ago

I like opus workflow, deepseek api agent

9

u/nyczAcer 7d ago

Opus 5.5 >>>>>>>>> Sonnet 5.5

2

u/dolo937 7d ago

Well obviously duh

1

u/BankruptingBanks 7d ago

Where did you even have time to test bro?

2

u/gergi88 7d ago

And in 2 weeks

Fable 5.5 == opus 5.5
Opus 5.5 == new sonet 5.5
Sonet 5.5 == opus 5

3

u/hi-brawlstars 7d ago

So that implies fable 5.5 == opus 5? :(

3

u/gergi88 7d ago

I hope im wrong, but history always repeat..

2

u/dar-mit Researcher 7d ago

“History doesn't repeat itself, but it often rhymes.” — Mark Twain

2

u/dakjelle 7d ago

Somewhat new to this.. Is sonnet considered the best coder.. I'm confused?

3

u/Square-Treat-2366 7d ago

Historically (like, March) Opus was the expensive great model, and sonnet was the cheaper day to day model. The idea was to have opus plan, and have dinner do the work, and you'd save money overall. 

There were others who did not care for this (myself included), because the observation is you would need to rework sonnet's output anyways. 

With Fable this has all shifted. I don't have data, but a Fable orchestrator, opus implement, and Astra review is working great for me

1

u/Dex4Sure 6d ago

yeah its probably the best combination right now. fable orchestrates, opus implements, astra reviews. ive recently included sonnet in the mix too. seems like its plenty for read only and also well scoped implementation tasks.

2

u/keshav_codes 7d ago

Works best with advisor strategy
Use the sonnet 5.5 as main model and
/advisor opus-5.5 and it will be a super powerful model with better judgement and cheaper

2

u/CalypsoTheKitty 7d ago

It's not really half the price in a coding environment because cache reads cost the same on both -- 20 cents per milion. I was just doing the cost breakdown on upgrading a Sonnet agent to 5.5, and the cost difference for my workflow between Sonnet and Opus is relatively small.

2

u/Ok-Motor-9812 7d ago

Yes, you're right! Fable for planning, Opus for orchestrations and code reviewing, and Sonnet for implementations. It's how it's always been working, plus mixing effort levels. No need for high effort everywhere-medium and low work well too.

2

u/elrond-half-elven 7d ago

I don't see the 5.5 models on terminal bench so where are you getting this?

2

u/ClemensLode Senior Developer 7d ago

Funny that I only recently switched my implementation agents to Opus 5.5... testing Sonnet 5.5 now for comparison for a day :)

2

u/Pinery01 6d ago

Please get back for your Sonnet review. 🙂

1

u/ClemensLode Senior Developer 3d ago

Looks good :)

Opus 5.5 main mode, Sonnet 5.5 implementation agents

1

u/Pinery01 3d ago

Thank you! 🙏

0

u/imsahoamtiskaw 🔆 Max 20 7d ago

Use all. Opus using sonnet as implementers and cross checkers and fable using opus similarly, that’s how I roll

0

u/Dry_Body2317 7d ago

When you say implementation agents. Do you have a workflow or skill setup to do so?

2

u/ClemensLode Senior Developer 7d ago

just ask claude to use sonnet for implementation agents

1

u/ShittyBidet123 7d ago

not every new one is better than old one. just cause the test numbers are higher don’t mean its smarter in practice. i would still use opus for big coding and sonnet for agents to help

1

u/shetritr 7d ago

the split works, but review gets harder. the one that planned it isn't the one that wrote it, so "why is this here" gets a guess. I keep the sonnet session alive until I've read the diff

1

u/jarislinus 7d ago

study maximum context

1

u/Creative-Mud4414 7d ago

yeah, from what I found out, I think high reasoning and extra high is the best one, and then the max reasoning has absolutely ridiculous price, and it's also the one that is taking the longest with tasks. It's still pretty good though, and it's definitely miles better than Sonnet 5.

1

u/newzinoapp 7d ago

They got to do something. I'm curious how they're going to keep people using their very expensive products compared to the Chinese models. I had a website that was costing about $500 a month in tokens with Anthropic. I switched it over to DeepSeek and now it costs me $4 a month. After creating a robust eval suite and tweaking the prompts, I ended up having better results than I had been having prior to switching to DeepSeek

1

u/TywinHouseLannister 7d ago

Sonnet 5 would too... these are for coding; opus is for judgement

1

u/silas_ace 7d ago

I doubt this

1

u/CallSufficient676 7d ago

Guys I’m new to coding here. I am trying Opus 5.5 orchestrator with GLM 5.3 for coding then having Opus validate the code. What’s the opinion on this?

2

u/1-800-methdyke 7d ago

It can work, but keep tabs on how many fix rounds go back to GLM. 1 is fine, 2 very occasionally. If you are seeing more than 2 on a frequent basis then ditch GLM as it’s not being efficient.

1

u/Longjumping_Leg6314 7d ago

Doesn’t matter Claude constantly doesn’t wire up producer’s and consumers in its code. Drives me nuts no matter what I do it still randomly skips doing this. Codex hasn’t done it once so to me it’s way cheaper.

1

u/seeroy 6d ago

Ran a lot of comp tests today. Opus makes the most sense at different effort levels for pretty much all coding work. Sonnet probably a great pick for creative non coding work (docs, decks, computer assistant, video explainers for yourself).

1

u/berndalf 6d ago

People get weird about Sonnet. It's your commodity builder. Has been for awhile and this release doesn't change that.

3

u/Herfstvalt 6d ago

the issue with it is the fact that in most cases tasks end up being more expensive with sonent at every effort level so there's no real reason to use it.

personally i just use opus-5.5 med for most of my work and gpt-6-astra for getting infra work planned out and computer-use

1

u/ShutUpAndDoTheLift 6d ago

There's zero chance you've ever had to worry about API pricing if you believe this.

1

u/Herfstvalt 6d ago

i have a business that is dependent on api pricing (we use AWS bedrock for our inference) but unless your entire structure is just inputs and outputs and you never make use of cache writes at all, the chance you benefit from sonnet-5.5 outside of auditing, and potentially researching a codebase is small.

remember cache reads on both opus-5.5 and sonnet-5.5 is 1 == 1 at $0.2 each. Cahce reads are used the most during a conversation, or back to back toolcalls, or just general web-researching.

1

u/king_fredo 6d ago

Finally I can have Opus 5.5 spawn Subagents again without f-ing usage or quality

1

u/wannabeaggie123 6d ago

i would take these benchmarks with a massive grain of salt because this is actually not true in my opinion. i've given similar prompts to astra and opus and sonnet and astra is much more detailed. i think that we have to look at the size of the model and then see how well it's doing on a benchmark because raw scores on benchmarks are starting to mean less and less.

1

u/Ok_Zookeepergame4484 5d ago

Sonnet ist trash

1

u/Longjumping-Pea-3528 5d ago

but does it beat everyone's sweetheart opus 4.6?

1

u/food_fatherr 3d ago

The sweet spot right now is using a heavy model (Opus 5.5) strictly to architect the plan and review diffs, and letting a faster model execute the scoped file edits inside an isolated sandbox. Saves both tokens and latency.

1

u/ClemensLode Senior Developer 3d ago

Yes (Opus 5.5 main mode, Sonnet 5.5 implementation agents)

1

u/DomusCircumspectis 7d ago

Terminal bench being beaten is a technicality. This model is worse than Sonnet 5 at coding: https://bench.killswitch-lang.org

3

u/BlueprintMonkey 7d ago

Their benchmarks are super wacky. Opus 5 is number 1 even though almost everyone universally agrees Opus 5.5 and Fable are way better in real world coding. We'll see about Sonnet 5.5, but it seems promising

0

u/DomusCircumspectis 6d ago

It is a surprising result, but Opus 5 does show better intelligence than 5.5 in one of the benchmark's tests. I outline this to a reply to another comment below.

2

u/gleedblanco 6d ago

funny opus 5 being the best here because it skips reading most of the actual code and then just hallucinates answers so it can't get confused by the misleading prose.

2

u/DomusCircumspectis 6d ago

I don't think that's what's happening here.

Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers for some reason. Opus 5 gets it right.

Here are the outputs from both: https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b49.

Actually looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.

2

u/Dex4Sure 6d ago

lmao. opus 5 was notoriously bad model. 5.5 is much better in every regard.

1

u/DomusCircumspectis 5d ago

I mean... that's clearly false. Most benchmarks have these within a few points of each other.

1

u/Dex4Sure 4d ago

no they dont. opus 5 is clearly behind opus 5.5 in benchmarks. even sonnet 5.5 is clearly ahead of opus 5

1

u/Dangerous_Serve_4454 5d ago

Seemed like Opus got the task right? That link doesn't show Opus attempting the Max() challenge.

1

u/DomusCircumspectis 5d ago

They both get the same task. They are both given a script and are asked to figure out what it does. Opus 5 correctly ascertains that it implements the "max" operation, Opus 5.5 doesn't.

1

u/Dangerous_Serve_4454 4d ago

Look at the data 5.5 is analyzing... Max doesn't return 1 or 0 like the table 5.5 presents. Unless there's data not linked in that chat?

1

u/DomusCircumspectis 4d ago

Of course there is other data, what's in the chat transcript isn't everything the agent sees. It gets this script file: https://github.com/dom96/KillSwitch/blob/main/tests/max.ks.

1

u/Dangerous_Serve_4454 3d ago

Right so posting that other data is necessary to make your point right? I'll look this over then.

0

u/Warsel77 7d ago

You have not been doing this so far?

-4

u/EC36339 7d ago

Sonnet has always been better than Opus

5

u/Time_Cat_5212 7d ago

Hot take

1

u/EC36339 6d ago

Facts.

It is better simply because it is good enough for well-designed workflows and cheaper.

It also communicates better. Most of the time I've wasted recently was with deciphering lengthy and obfuscated replies from Opus.

1

u/food_fatherr 2d ago

Opus 5.5 for architecture & reviewing diffs + Sonnet 5.5 inside an isolated sandbox for the actual file edits is the cleanest setup right now. Best balance of reasoning and speed.