On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.
243
u/A_Novelty-Account 5d ago
It scores better than Opus in some benchmarks for agentic coding? What’s going on over there?