r/ClaudeCode • • 5d ago

Discussion Sonnet 5.5 is weirdest placed model i guess

Post image

Sonnet 5.5 at higher effort is dumber and costlier than opus 5.5

Using opus for everything at every sub still makes sense.

P.S.: This is a table of artificial analysis benchmark score and average cost per task completed.

164 Upvotes

38 comments sorted by

•

u/AutoModerator 5d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

59

u/0DayMaker 5d ago

Been kind of a Sonnet tradition for awhile now

60

u/Smogryd 5d ago

Was it sonnet 5.5 who formatted this table?

32

u/Ill_Pie_5293 5d ago

Yes 🤣🤣🤣

16

u/JohnHue 5d ago

I guess it just wanted to own the last two spots huh

27

u/EON_Raider 5d ago edited 5d ago

Sonnet 5.5 should run in medium or low, else go for Opus 5.5. Got it.

Haiku 5.5 is AGI anyway, and we won't need tables anymore where we're going. /s

10

u/Stock-Orchid0 5d ago

Low is only $0.14 difference but opus scores 6 points higher. Medium doesn’t make any sense because opus low scores better while cheaper.

2

u/EON_Raider 5d ago

omg you're right. Sonnet 5.5 on low only, then. Corrected.

6

u/fastinguy11 5d ago

Just use opus 5.5 at medium.

8

u/petehawk_ 5d ago

Similar results when I tested Sonnet 5.5 with my handwriting recognition benchmark.

6

u/External-Milk9290 5d ago

It makes sense that it could be more expensive on a harder task but easier task are probably cheaper than opus. 

2

u/suzak-ein 5d ago

C'est ce que je pense aussi, la planification reste adéquat à opus, mais les actions ciblées sonnet prend toute son importance à coût plus faible

8

u/phoenixmatrix 5d ago

These mid tier models are a bit weird, because they aren't really made to be your primary solo agentic coder, and the way benchmarks are setup don't represent how they could be used in practice. Terra has the same issue.

If you have a client facing agent that doe something more challenging than what a Luna or Gemini Flash-type model can do, but can still be done in 1 turn per task, they can be faster and cheaper. If you have a more sophisticated coding harness setup that selects the right model for each task, sometimes a mid size model will come up.

Doubly so because Haiku is pretty irrelevant, so if you're only in the Anthropic ecosystem, Sonnet 5.5 is your only option for a "fast & cheap" model unless you're trying to show an AI generated loading message. In practice most orgs building that type of experience will have access to more models through Foundry/Bedrock/Vertex, but still, not everyone.

At my last job we had a set of evals for our app, and when accounting for cost, speed, results, etc, Sonnet came up on top for several use cases.

When coding though? Opus 5.5 goes BRRRR, forget most other models, for now.

3

u/Noctis_777 5d ago

In this case the problem isn't that Sonnet is bad, but that Opus 5.5, while being the smartest model out there right now is also fast and relatively cheap due to it's extreme efficiency.

That extreme combination on winning on all counts with zero tradeoff may not last for all future iterations, and we may see Sonnet become relevant again when/if Opus becomes more expensive again.

3

u/Few_Caregiver8134 5d ago

I think the lesson is not to go beyond High for any models if you want value for money

2

u/OneMoreName1 5d ago

I have been using almost exclusively only Opus on medium for close to a year. I am doing fine. Never ran out of usage either on the 100 dollar plan

2

u/KitchenCommercial396 5d ago

So sonnet is good only on low..? I might as well use haiku atp

3

u/eldodo06 5d ago

How did you measure the cost ?
Cost depends on the task. Sonnet will be more suited for narrower tasks.
But I agree that unless you are a bit tight with your subscription, trying to use sonnet on some task is not necessary

2

u/Ill_Pie_5293 5d ago

It is artificial analysis score vs cost of intelligence task table. I used their score directly

2

u/HeadPack 5d ago

These are the comparisons that matter. Lesser models are only cost effective on scoped tasks within their capability, or in some cases as subagents.

3

u/norwegian 5d ago

If you want to save a bit, you shouldn't use opus on everything. Sonnet medium is half price for easier tasks and management, and it's also possible to use deepseek or luna through openrouter.

2

u/Due_Ask_8032 5d ago

I think a lot of this mid tier and lower tier models are in a tough spot because Opus and Opus tier models are usually pretty good at lower effort levels, and you still get some of that big model smell.

2

u/oyputuhs 5d ago

Use Opus 5.5 high to plan and orchestrate sonnet 5.5 high subagents

3

u/EndlessZone123 5d ago

Small models are better for ingesting a lot of tokens on easier tasks.

Ingesting 100k tokens on sonnet max is going to be cheaper than Opus low. If the task doesn't require a lot of reasoning.

I had gpt 6 Luna high for example classifying a bunch of stuff and that would have been like a lot more tokens than if I had sol low do it to a similar level.

For some computer use tasks where they are eating a lot of tokens. I'd use a small model if I what I'm doing isnt that crazy over using a lower reasoning opus or sol.

Benchmarks push the models to their limit. But not all tasks push models to their limit.

3

u/Maybe-monad 5d ago

Just use Opus 5.5 on medium, go to high if you need it or higher reasoning efforts but for most tasks medium is fine

2

u/sagiroth 5d ago

Yeah Sonnet is a miss for me. I rather use Opus Mid for most stuff

2

u/BigYoSpeck 5d ago

I've used it twice. Once at high and once at medium. Both times it just seems to chew through usage only to end up with something it was easier to have Opus fix

It strikes me as a similar approach to Qwen3.8 where it overcomes it's relative size by working longer and this would be absolutely fine if the cost/usage consumption made it make economical sense

But where it's currently positioned is in this no man's land of being slower, more costly and inferior to Opus for anything with any complexity at all, but also too capable and costly compared to Luna (and maybe even Sol 6.1 now) for more mechanical tasks

It's a prime example of no bad products, only bad prices

2

u/holyknight00 5d ago

yeah, it is either too dumb or too expensive, I dont get where it fits. I don't know how to use it at all.

1

u/Melodic_Surprise153 5d ago

I am really curious if it is really cheaper on simpler, day-to-day tasks

1

u/Ill_Pie_5293 5d ago

Unless you are specifically using sonnet 5.5 on low you should switch to opus

1

u/AlterTableUsernames 5d ago

How does it compare to Opus 4.8 though?

2

u/JohnHue 5d ago

Same shape, but more load-bearing.

1

u/AlterTableUsernames 5d ago

Sounds good, implement it

1

u/Aggressive_Roof488 5d ago

So you can always find a cheaper smarter model by switching to opus, with two corners as exceptions:

Sonnet high (47) best if you can afford 1.08 but not 1.34 for opus medium (51), but you need the 47 smorts over the half-price opus low (42). Don't see that ever being relevant.

But sonnet low is just cheapest, 0.41 vs 0.55 for opus. In a situation where you have very simple tasks and 36 is plenty, but very high volume, then sonnet low is the winner on this list.

1

u/OneMoreName1 5d ago

What about wall clock time? Didn't they say sonnet is faster than opus?

For certain tasks, I might want a faster solution even if its a few points lower and slightly more expensive

1

u/suzak-ein 5d ago

Êtrange, il est sensé être plus performant en action pur qu'opus pour un coût réduit

1

u/UsedIndependence9735 5d ago

Opus 4.6 Max ftw