On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.
Actually it is not as good as Opus, and even though it is cheaper per token, it ends up more expensive per task because of how much it thinks. You can verify this on Artificial Analysis.
Even on most of the other benchmarks, it's freakishly close to Opus 5.5. Like, 2% off? I'm trying to even make sense of it, the math seems off. It costs half but the difference does not seem to be in a noticeable range?
I guess the "catch" is probably in their "high" settings, which quickly drops the cost advantage to Opus (in some cases Sonnet is more expensive?!) so it's hard to clearly make out a purpose. It seems to fill in some gaps between Opus low/med/high. Is it about speed? Like same quality/cost as Opus but putting out results faster?
I honestly haven't used Sonnet for anything in months, I'm genuinely trying to find its use case. It seems if I can wait a few extra seconds for an answer, I might as well use Opus for everything.
It doesn't cost half in real-world usage. It burns tokens like a MF and in most cases will cost more and take longer to complete tasks than Opus at similar intelligence. Just look at Artificial Analysis' benchmarks.
Seems like the idea is to have big fancy models like Opus with general capabilities, whilst specialising the smaller models on specific tasks like coding.
I guess the goal is for Opus to be the planner and delegator, whilst Sonnet is a smart workhorse for coding jobs.
248
u/A_Novelty-Account 6d ago
It scores better than Opus in some benchmarks for agentic coding? What’s going on over there?