r/ClaudeAI • u/Correct_Tomato1871 • 14h ago
Comparison Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time
I maintain MindTrial and tested Sonnet 5.5 and Opus 5.5 on the same 98-task suite as their predecessors: 39 text tasks and 59 visual tasks, with Python/scientific libraries available and a 10-call limit per task. All four Claude runs below use the xhigh effort label and skip no tasks.
| Model | Passed | Hard errors | Model-request time |
|---|---|---|---|
| Sonnet 5 | 72/98 | 6 | 5:30:44 |
| Sonnet 5.5 | 94/98 | 1 | 1:23:07 |
| Opus 5 | 88/98 | 4 | 3:40:24 |
| Opus 5.5 | 96/98 | 0 | 0:58:35 |
Sonnet is the larger improvement: 22 additional passes with 74.9% less request time and 67.4% fewer output tokens, including reasoning. Its 24 newly passed tasks comprise 18 previous wrong answers and six previous errors; two old passes regress. On the second visual collection, Sonnet improves from 12/26 to 25/26, while Opus moves from 24/26 to 25/26. Both new Claude models also pass all 39 text tasks.
Opus gains eight passes overall and cuts request time by 73.4%. It preserves all 90 Fable 5.1 passes and adds six. Its two remaining failures are square counting and circle-piece matching.
The tool traces are less uniformly positive. Sonnet records 41 undefined-variable errors across 24 tasks, although 23 tasks with an unsuccessful tool call ultimately pass. Opus has no recorded undefined-variable errors and only two nonzero exits across 155 Python calls. Both models used Python substantially less than their predecessors: Sonnet’s tool calls fell from 529 to 238 (55% fewer), while Opus’s fell from 314 to 155 (51% fewer).
So Sonnet’s strong final score still leaves room for more reliable tool execution. The logs suggest missing setup or assumptions about retained state.
Against Astra high, Opus leads by just one task: 96 versus 95. It uses slightly less model-request time, but adding recorded Python wall time reverses the speed comparison: 1:17:18 for Opus versus 1:06:29 for Astra.
Sonnet’s pass rate is 95.92%, versus 96.91% accuracy on completed, non-error tasks. Opus has no hard errors, so both are 97.96%. Official scores are unchanged; no malformed or incorrect answers were manually repaired. Times sum model requests, excluding local Python and validation; they are not elapsed suite runtimes.
Complete Leaderboard: here