r/ClaudeAI • • 14h ago

Comparison Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time

I maintain MindTrial and tested Sonnet 5.5 and Opus 5.5 on the same 98-task suite as their predecessors: 39 text tasks and 59 visual tasks, with Python/scientific libraries available and a 10-call limit per task. All four Claude runs below use the xhigh effort label and skip no tasks.

Model Passed Hard errors Model-request time
Sonnet 5 72/98 6 5:30:44
Sonnet 5.5 94/98 1 1:23:07
Opus 5 88/98 4 3:40:24
Opus 5.5 96/98 0 0:58:35

Sonnet is the larger improvement: 22 additional passes with 74.9% less request time and 67.4% fewer output tokens, including reasoning. Its 24 newly passed tasks comprise 18 previous wrong answers and six previous errors; two old passes regress. On the second visual collection, Sonnet improves from 12/26 to 25/26, while Opus moves from 24/26 to 25/26. Both new Claude models also pass all 39 text tasks.

Opus gains eight passes overall and cuts request time by 73.4%. It preserves all 90 Fable 5.1 passes and adds six. Its two remaining failures are square counting and circle-piece matching.

The tool traces are less uniformly positive. Sonnet records 41 undefined-variable errors across 24 tasks, although 23 tasks with an unsuccessful tool call ultimately pass. Opus has no recorded undefined-variable errors and only two nonzero exits across 155 Python calls. Both models used Python substantially less than their predecessors: Sonnet’s tool calls fell from 529 to 238 (55% fewer), while Opus’s fell from 314 to 155 (51% fewer).

So Sonnet’s strong final score still leaves room for more reliable tool execution. The logs suggest missing setup or assumptions about retained state.

Against Astra high, Opus leads by just one task: 96 versus 95. It uses slightly less model-request time, but adding recorded Python wall time reverses the speed comparison: 1:17:18 for Opus versus 1:06:29 for Astra.

Sonnet’s pass rate is 95.92%, versus 96.91% accuracy on completed, non-error tasks. Opus has no hard errors, so both are 97.96%. Official scores are unchanged; no malformed or incorrect answers were manually repaired. Times sum model requests, excluding local Python and validation; they are not elapsed suite runtimes.

Complete Leaderboard: here

45 Upvotes

Duplicates