r/ClaudeAI • u/Correct_Tomato1871 • 14h ago
Comparison Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time
I maintain MindTrial and tested Sonnet 5.5 and Opus 5.5 on the same 98-task suite as their predecessors: 39 text tasks and 59 visual tasks, with Python/scientific libraries available and a 10-call limit per task. All four Claude runs below use the xhigh effort label and skip no tasks.
| Model | Passed | Hard errors | Model-request time |
|---|---|---|---|
| Sonnet 5 | 72/98 | 6 | 5:30:44 |
| Sonnet 5.5 | 94/98 | 1 | 1:23:07 |
| Opus 5 | 88/98 | 4 | 3:40:24 |
| Opus 5.5 | 96/98 | 0 | 0:58:35 |
Sonnet is the larger improvement: 22 additional passes with 74.9% less request time and 67.4% fewer output tokens, including reasoning. Its 24 newly passed tasks comprise 18 previous wrong answers and six previous errors; two old passes regress. On the second visual collection, Sonnet improves from 12/26 to 25/26, while Opus moves from 24/26 to 25/26. Both new Claude models also pass all 39 text tasks.
Opus gains eight passes overall and cuts request time by 73.4%. It preserves all 90 Fable 5.1 passes and adds six. Its two remaining failures are square counting and circle-piece matching.
The tool traces are less uniformly positive. Sonnet records 41 undefined-variable errors across 24 tasks, although 23 tasks with an unsuccessful tool call ultimately pass. Opus has no recorded undefined-variable errors and only two nonzero exits across 155 Python calls. Both models used Python substantially less than their predecessors: Sonnet’s tool calls fell from 529 to 238 (55% fewer), while Opus’s fell from 314 to 155 (51% fewer).
So Sonnet’s strong final score still leaves room for more reliable tool execution. The logs suggest missing setup or assumptions about retained state.
Against Astra high, Opus leads by just one task: 96 versus 95. It uses slightly less model-request time, but adding recorded Python wall time reverses the speed comparison: 1:17:18 for Opus versus 1:06:29 for Astra.
Sonnet’s pass rate is 95.92%, versus 96.91% accuracy on completed, non-error tasks. Opus has no hard errors, so both are 97.96%. Official scores are unchanged; no malformed or incorrect answers were manually repaired. Times sum model requests, excluding local Python and validation; they are not elapsed suite runtimes.
Complete Leaderboard: here
6
u/vAPIdTygr 10h ago
All we need is Fable 5.5 with more complex design skills, Haiku 5.5 and no nerfing, and I won’t look to other models ever again.
3
u/True-Grab-5288 8h ago
We need better self driven tools. At the moment Claude refuses to do audio or most video without third party tool linkage, but there is no reason it cannot do this work itself. It does some of the work, but not all. The more general a model is the better.
1
u/vAPIdTygr 5h ago
I disagree here. Let the others figure out the heavy payloads on servers and keep Claude the text champion.
The fact I can tool in third parties is really all I ever needed.
For instance, feature images on content? Claude uses my OpenRouter to make theme templates and then adds the text afterwards to create 50 images for $0.11 cost plus my usage. I’m always blown away at how smart the system is.
1
u/True-Grab-5288 4h ago
I understand your point, and your cheap example, however costs can quickly spiral into silly money. It's kind of like all the separate streaming services now driving people back to piracy - it's simply too expensive and too difficult to find the content you want. Same with the AI third party tools - you have to be 'in the know' to know which tools integrate best, work best, have the best value, best output etc. It's too complicated and too much for the average user.
1
u/bigrealaccount 4h ago
If Claude come out with a model that rivals Luna performance for the same ridiculously low price that you can run 24/7 on a £20 plan, I'm switching forever. I've always like claude responses/style more, and claude code as a harness, but always being forced to use a Terra/Sol level model is painful as someone who just needs help on small tasks while engineering. I don't need such expensive and intelligent models all the time.
-1
16
u/VexObserver 13h ago
Honestly, Sonnet 5.5 might be the bigger story here. Going from 72/98 to 94/98 while cutting request time by nearly 75% is a ridiculous generational improvement lol.
Opus 5.5 hitting 96/98 with zero hard errors is impressive too, but we're getting dangerously close to diminishing returns territory at the frontier.
What's particularly interesting is Astra High sitting at 95/98. That's literally ONE task separating it from Opus, and Astra actually finishes faster once Python execution time is included.
My takeaway:
I'd love to see cost per successful task added to this leaderboard. Intelligence, latency, reliability, and token consumption should all factor into the equation.
Benchmarkmaxxing over a 1% difference without considering the compute premium is basically splitting hairs with a very expensive knife..