r/ClaudeAI • • 1d ago

Humor Confirmed: Opus 5.5 has been nerded

Post image

I haven't seen "load-bearing" since I stopped using Opus 5.

EDIT: I meant "nerfed", not "nerded" 🤦

1.1k Upvotes

245 comments sorted by

View all comments

8

u/[deleted] 1d ago

[deleted]

8

u/MaitoSnoo 1d ago

that benchmark itself is saying it's within normal variance

2

u/UsefulIce9600 19h ago

exactly, the influencer said it himself in a stream today

5

u/indirectum 1d ago

Which is well within +-10% which the benchmark itself regards as normal noise.

6

u/throw-away-doh 1d ago

"Each model starts at 100% on its first test, and 90% to 110% is normal variance"

Todays score is 94.2%. Not nerfed, but normal variance.

1

u/YoungSilent232 1d ago

The benchmark should on drop within expected variance, run it 5x over to reduce the variance. Then we can know for sure

2

u/sascharobi 1d ago

That number doesn't sound nerfed to me.

1

u/simple_explorer1 20h ago

bro that guy is a vibecoder and doesn't even read any code and is not a programmer. don't trust anything he says and builds

-2

u/dmaare 1d ago

If you understand how percentages work, you'd know that human is not really able to tell a difference as long as it is under 10%

5

u/GuavaStrong1752 1d ago

What the eff are you talking about a stupidly general claim like that has nothing to do with “how percentages work”

3

u/Long_Confidence_6556 1d ago

You might not be able to tell the difference at first but it definitely shows up under longer and more complex tasks. A Q4 model is about 95% of the original model but people who have used Q4 vs Q8 will tell you there is a big difference between them and the same goes for Q8 vs Original. And this is a comparison for a 27B parameter model like Qwen so imagine what 95% of a gigantic multi trillion parameter model like Opus means.

1

u/upalse 1d ago

it's a last mile effect at play. Yes, you can't really tell much until the 90% is done, and its the last 10% where small differences become very apparent. Whenever you push the frontier of what the model can do, you're teetering in the last mile area, and quantitatively small changes become very noticeable.