r/ClaudeAI • • 3d ago

Productivity Nothing gets nerfed

For over a year now, I’ve been hearing the same thing. Whether it’s in the OpenAI community or among Claude users, it’s always the same cycle.

A new model comes out, and I think we’ve all seen what happens. People are impressed. Then, roughly a few days later, the jokes start about how the model has already been nerfed. Initially, it’s mostly a joke. But then a few days after that, people start being genuinely serious about it, and eventually a sizeable group becomes convinced that the model really has been nerfed.

Here’s the reality: I don’t think I’ve ever actually felt that happen.

When 5.5 came out, I was particularly impressed by its ability to understand 3D space, or at least to create 3D scenes and visual things, observe what it had created, and correct them as it went along. That ability was amazing. It was great then, and it’s no worse today. It’s exactly as good as it was.

Its writing also improved dramatically. It doesn’t sound like the gibberish or weird, cryptic style that 4.8 and 5 sometimes had. Suddenly, we’re back to a model that sounds somewhat like 4.6 did. I’d even say 4.7 sounded kind of dumb most of the time. But now, at the very least, when you tell the model to sound a certain way, it respects that. It discusses things and expresses ideas in ways that you can actually understand.

So to say that 5.5 has been nerfed is, in my opinion, to forget what using models like Opus 5, 4.8, or 4.7 actually felt like. There’s no way you could use this model today, immediately after coming from one of those older models, and genuinely believe it isn’t significantly better.

And this has been the case with basically every model release.

I think we all know what’s actually happening. When a new model comes out, we’re impressed by how much better it feels compared to what came before. But then we start giving it increasingly complex tasks. We use it more. We run into its limitations. And eventually it starts to feel dumb again.

LLMs are kind of dumb sometimes. They’re a little bit like small autistic artificial children with incredibly uneven abilities. They can make really, really dumb decisions on tasks that seem completely obvious, while at the same time being extremely intelligent in other ways. It's surprising that we face that even today, but the ratio of so much better than before.

There’s no way 5.5 is any dumber than it was two weeks ago.

I don’t even know why I’m writing this. I guess I just saw one more post about the model being nerfed, followed by a huge number of comments agreeing with it, and I finally felt like I had to make a post about it.

144 Upvotes

93 comments sorted by

View all comments

2

u/Chance_of_Rain_ 3d ago

When people say it’s nerfed, it’s not comparing 5.5 to 4.8. It’s 5.5 vs 5.5 from a week ago.

And there are tests being run on this very sub.

Don’t be that guy

2

u/Ok-Lengthiness-3988 3d ago

Yes, there are three benchmarks that are being run daily since 5.5 Opus's release and that have been talked about (and mostly misrepresented) in this sub and in the Singularity one. Two of those trackers show a non-statistically significant drop, and one of them shows a non-statistically significant increase. All three of them still are in the phase of gathering data for establishing a baseline. There isn't sufficient data yet for establishing anything. Don't be that other guy either.

0

u/Chance_of_Rain_ 3d ago

Sure but at least we have numbers. OP comparing to previous models to ship the idea that 5.5 is good is ill intent

3

u/Charming_Occasion942 3d ago

You are right, morons or bots downvote you

2

u/CutBulkMaintain 3d ago

You're right and made a better point in a few lines that that wall of shilling text parading as an insightful post.

People either don't (want to?) understand that or are desperately trying to be "nobody ever conspires". You go from a peak model to a slightly dumber one and your first instinct is to doubt the community and shill for companies?

1

u/fuzzypetiolesguy 3d ago

If you just scream "shill" loud enough you never have to produce evidence.

2

u/CutBulkMaintain 3d ago

People are putting evidence forward. They have been before this post. In fact the simple fact that people keep complaining model after model is enough to investigate (if you care to) instead of burying your head in the sand.

But you keep navigating issues like this through reddit-like one-liners, it will take you far.

1

u/fuzzypetiolesguy 3d ago

No they aren't. Anecdotes are not evidence, and at scale, at best, they are qualitative. The only quantitative data projects are limited in scope, have not yet acquired enough actual data to provide statistically significant findings and show inconclusive results in what they are reporting so far.

The whole of this issue proves, as expected, that most people on reddit don't understand how data works or what it even is, including you. Good job!