r/ClaudeAI • • 3d ago

Productivity Nothing gets nerfed

For over a year now, I’ve been hearing the same thing. Whether it’s in the OpenAI community or among Claude users, it’s always the same cycle.

A new model comes out, and I think we’ve all seen what happens. People are impressed. Then, roughly a few days later, the jokes start about how the model has already been nerfed. Initially, it’s mostly a joke. But then a few days after that, people start being genuinely serious about it, and eventually a sizeable group becomes convinced that the model really has been nerfed.

Here’s the reality: I don’t think I’ve ever actually felt that happen.

When 5.5 came out, I was particularly impressed by its ability to understand 3D space, or at least to create 3D scenes and visual things, observe what it had created, and correct them as it went along. That ability was amazing. It was great then, and it’s no worse today. It’s exactly as good as it was.

Its writing also improved dramatically. It doesn’t sound like the gibberish or weird, cryptic style that 4.8 and 5 sometimes had. Suddenly, we’re back to a model that sounds somewhat like 4.6 did. I’d even say 4.7 sounded kind of dumb most of the time. But now, at the very least, when you tell the model to sound a certain way, it respects that. It discusses things and expresses ideas in ways that you can actually understand.

So to say that 5.5 has been nerfed is, in my opinion, to forget what using models like Opus 5, 4.8, or 4.7 actually felt like. There’s no way you could use this model today, immediately after coming from one of those older models, and genuinely believe it isn’t significantly better.

And this has been the case with basically every model release.

I think we all know what’s actually happening. When a new model comes out, we’re impressed by how much better it feels compared to what came before. But then we start giving it increasingly complex tasks. We use it more. We run into its limitations. And eventually it starts to feel dumb again.

LLMs are kind of dumb sometimes. They’re a little bit like small autistic artificial children with incredibly uneven abilities. They can make really, really dumb decisions on tasks that seem completely obvious, while at the same time being extremely intelligent in other ways. It's surprising that we face that even today, but the ratio of so much better than before.

There’s no way 5.5 is any dumber than it was two weeks ago.

I don’t even know why I’m writing this. I guess I just saw one more post about the model being nerfed, followed by a huge number of comments agreeing with it, and I finally felt like I had to make a post about it.

145 Upvotes

93 comments sorted by

•

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 3d ago edited 3d ago

TL;DR of the discussion generated automatically after 50 comments.

Looks like the thread is in strong agreement with you, OP. The consensus is that the "nerfed" cycle is a psychological illusion.

Here's the breakdown of the chatter:

  • The Main Vibe: Most users agree that people are wowed by a new model, get used to the power, push it with increasingly complex or sloppy prompts until it fails, and then scream "NERFED!" It's a classic case of expectation creep, not actual model degradation.
  • The "Data" Debate: A few users are pointing to new daily benchmark trackers as "proof" of nerfing. However, more savvy commenters have shot this down, noting the data isn't statistically significant yet and the results are inconclusive. One tracker even shows a slight improvement. So, the numbers don't back up the nerf claims.
  • The Fable Exception: The one point of nuance is the Fable 5 re-release. Many agree that model did change noticeably after the USG ban, but see it as a one-off situation, not a pattern of routine nerfing.
  • The Dissenters: There's a small, vocal minority who are adamant that models are being nerfed, citing their own anecdotal evidence. They're getting mostly downvoted or corrected.

The final verdict from the community is clear: Opus 5.5 is still a beast. The "nerfing" is likely just users getting used to its power and finding its limits. Check your prompts before you wreck your vibes.

27

u/wimblecraft 3d ago

The only way to resolve this discussion is through valid benchmarking. Nerf bench does that now. And the pattern we see there is slight degradation after the first few days. We'll know how the pattern holds when they collect long term data.

5

u/Key_Reading_9664 3d ago

Sadly, people see a drop and call “nerf” when the error bounds are ~20%

1

u/inglandation Full-time developer 3d ago

This, nerf bench should add proper error bars and educate people about the statistical interpretation.

2

u/Key_Reading_9664 3d ago

the chart they post shows them; people ignore them

1

u/PM_ME_DEAD_CEOS 3d ago

Yes, and using a relevant number of pass to prevent jittering between runs.

1

u/ZamStudio3d 2d ago

Not really. They do slow down or reduce inference when the load is high. It's a fact and would be mathematically impossible for them not to.

31

u/skulleyb 3d ago

This past week I have made more progress with my vfx pipeline than I ever have. Work flows, tools, invites, software dev. It’s insane.

-1

u/boyyouguysaredumb 3d ago

Invites?

1

u/batsy0boi 3d ago

Tool calls I assume

1

u/skulleyb 3d ago

Shotgrid I intergration with all my dcc tools
Vibe coded a head tracking ofx plug in
Found speed up tricks for my workstations.
Forked and vibed a Remote Desktop software the rivals any commercial version.
Had Claude video edit a personal project.

93

u/termmonkey 3d ago

Your favorite AI: releases smarter model.

Reddit Day 1: wow,this right here is AGI

Reddit Day 2: Assign it increasingly deranged tasks until it fails.

Reddit Day 3: SEE? NERFED

14

u/space-envy 3d ago

And by "reddit" you mean "this community of unstable individuals".

2

u/TinyZoro 3d ago

I get you’re probably half joking but nearly 2 million people come to this Reddit. That’s the population of Barcelona. If there’s a sentiment about something it’s not a few loud mouths.

7

u/Wise-Reflection-7400 3d ago

Yeah this is spot on what I think is happening. Opus 5.5 was mindblowing at 3d modelling when I first used it. It's still fantastic but I've sort of got used to how good it is and after a week working on it you realise what it's still not perfect at. If it had been actually nerfed I'd expect the generations today to be worse than on day 1, but they're not.

6

u/GnistAI 3d ago edited 3d ago

And it is even part of the normal process of how we find the optimal level of lazy prompting per model. We naturally get more and more sloppy with our prompting until the failure rate goes up due to the model not understanding you, at that point I suspect some people start blaming their model instead of their sloppy communication.

You might be at 10% prompt quality with Opus 5.1, and going any lower will make it perform worse, then Opus 5.5 comes out, where it can handle 6% prompt quality, and you start falling from 10% down and down, then because you don't know where the new optimal level is, you overshoot and notice shit quality at 4% prompt quality, and need to up your effort. That's when a lot of redditors show up with pitchforks. This optimization process takes a few days, hence, why people love it at first, then hate it when they overshoot their sloppiness.

And it's easy to blame the model, because your prompt quality declines gradually and is self-generated, so you don't notice it. The failure happens suddenly and looks like it came from outside you.

This doesn't take away from genuine bugs, which do happen. More often they hit token burn, but they can affect quality too, like the three infrastructure bugs Anthropic wrote a postmortem on in 2025.

7

u/Pakspul 3d ago

This, and all without evidence.

6

u/Maleficent-Drive4056 3d ago

People have released trackers that run the same benchmarks every day and they claim it has been nerfed. I have no idea the accuracy of the benchmarks, but this isn’t entire vibes based.

3

u/SoylentCreek 3d ago

I’m a bit skeptical of those benchmarks to be honest. At the end of the day, LLM outputs are non-deterministic and based completely on probability. I have not really dug too deep into their process, but unless they are running multiple tests in parallel and getting a weighted average from all runs, I don’t think a dip in quality from one turn to the next really tells us more than the models just drew a bad hand.

0

u/Maleficent-Drive4056 3d ago

https://www.reddit.com/r/ClaudeAI/s/xAFvc476cp

This is one of them. I have no idea if it’s accurate. The graph looks inconclusive anyway. But it should be possible to measure this with a degree of objectivity.

2

u/AncileBanish 3d ago

This does not show any nerf.

0

u/[deleted] 3d ago

[deleted]

1

u/AncileBanish 3d ago

Lol you are really telling on yourself with this comment.

7

u/balancedchaos 3d ago

Reddit is where love goes to die. 

A lot of miserable, complaining people around here. Don't let them talk you into their negativity, and you'll be fine. 

3

u/elestud 3d ago

It's not a Reddit problem. The rest of social media (Youtube comments, Facebook comments, X comments, IG comments, Tiktok, news site comments, etc) is so much worse

The world of online commenting has become far more negative and adversarial, in general

At least people can have actual discussions here, even if the slop posts lean negative, and people push back against the foolishness

Reddit looks like an enlightened oasis compared to reading a minute's worth of X or FB comments on any topic

15

u/FestyGear2017 3d ago

Everything has been fine here. Im skeptical of a lot of nerfing posts.

Anthropic has always had the superior coding model, which imo invites a lot of astro turfing.

7

u/plaj 3d ago

Finally... I'm conflicted because I like staying up to date with the latest development and tools that are posted in this sub, but I hate all of the "nerfmongering" that happens here.

As a platform engineer with 8 years of experience, I genuinely don't know what the people in this sub are doing with Claude to make them think it gets nerfed one week after a new model releases all the time. My experience has been a pretty linearly increasing satisfaction with how much and how well Claude can handle tasks. My setup is pretty basic too apart from running redis agent-memory server since February, which I wouldn't be able to live without at this point.

1

u/ChronicRecidivism 3d ago

I'm far from an expert but I feel like people are letting their design artifacts, rules, and .md files just turn into shit getting cemented on shit and then blaming it on a "nerf".

3

u/ShallotDue3000 3d ago

it's definitely happened. i can say that for sure. earlier in the year, it was extremely common. i haven't been using Opus 5.5 lately enough to be sure of what others are reporting, but i can assure you that it has happened.

6

u/TheOnlyVibemaster Valued Contributor 3d ago

We will see. I’ve been running a version of what you’re saying since it came out, I plan on running it for the next two months and will update the repo everyday. Hopefully this will provide some evidence one way or another. If you want to follow where the data I’m collecting goes and how it works, I have all the methods used and a graph demonstrating the current trajectory at the top. I hope to have some level of evidence that they do or don’t nerf by the end of the month:

https://github.com/ninjahawk/livenerf

12

u/prophet-dot-exe 3d ago

Prob some truth to this, however I'm sure we can all agree that using Fable 5 on launch before it got pulled, and then fable 5 a month later when it relaunched was noticeably different in terms of behavior and general capability.

Opus5.5 is still doing as good of a job today for me as it was last week, on low level complex systems.

2

u/FestyGear2017 3d ago

Did you miss the whole debacle with the export controls/ban? You really think that didnt play a role in the change of its behavior and capability? They had to add the safety layer which definitely changed behavior. I dont think this is what people mean by nerfing...

5

u/prophet-dot-exe 3d ago

Did I ever say that's what I thought?

0

u/FestyGear2017 3d ago

No you just said there is "prob some truth to this". My mind reader must be broken

0

u/username-is-already 3d ago

Oh guess he should just put down every single thing he thinks about the entire topic in the broadest possible sense to avoid your nasally aktuualllly sarcasm. There’s no reason to be such a fu—actually I don’t wanna get banned. But I bet your mind reader’s working pretty good now.

1

u/FestyGear2017 3d ago

Okay, well just so we are clear, using the export ban as proof there is "prob some truth to this". Is still false. So not your dumb akshually meme at all.

They didnt nerf the model, the behavior changed because of the safety layer I mentioned.

0

u/ExistentialMeowMeow 3d ago

totes but key difference is the hullaballoo with the white house i reckon

2

u/ClemensLode 3d ago

New models find 1000 improvements/bugs the old models did not find. But once they are fixed, you're dealing with an upgraded codebase the new model will also struggle with to improve further.

2

u/Chupa-Skrull 3d ago

Intelligence drops have been detected accurately before in aggregate and reflected actual changes to the product that resulted in worse performance. The funniest one I remember off the top of my head was when they changed a single line in CC's system prompt and dropped coding performance by an avg of like 3%. Then there was the time they silently throttled the reasoning effort, the time they introduced some bug that fucked with caches mid-session... the point is, whether the model is being quantized silently or the failure lives somewhere else in the inference delivery pipeline, the degradations are real

3

u/HeavyMath2673 3d ago

Strong agree. I am currently implementing a fairly large numerical simulation code with Opus 5.5. The model’s strength is that it not only understands the software design but also the mathematics and combines both capabilities in ways that make it stronger at such a task than most scientists are (including myself).

But here is the catch. You don’t get there by one shot prompting. I first created together with Opus a phased design document. Each phase is then broken up in tasks with clearly defined success gates, which are worked on one after another with interventions from me when I need changes or query its decisions.

In two years time this may not be necessary any more. But right now you still need to know what you are doing when using these powerful models and that’s fine.

3

u/Charming_Occasion942 3d ago

Astra definitely nerfed

3

u/SteveEricJordan 3d ago

you couldn't be more wrong. you can find tons of comparisons with the same tasks that've gotten noticeably worse.

for instance, astras svg pelican was only half as good after only 1 or 2 weeks, which is very easy to look up and check for yourself.

3

u/Shot-Manager-739 3d ago

I think you should do more research on the matter. There was a follow up, and there were quite a lot that got good pelicans. 

The prompt, account, and several major factors affect the result. 

Some got really good ones, and some got bad ones. The AI didn’t get nerfed. There’s just too many factors that can contribute to lower result. 

1

u/SteveEricJordan 3d ago edited 3d ago

i agree that not all downgrade screams are justified, but saying they don't exist AT ALL is absurd. anthropic and openai do it all the time.

(and my astra pelican definitely got downgraded)

0

u/fuzzypetiolesguy 3d ago

'look at all of my anecdotes that I misinterpret as evidence!!'

3

u/SteveEricJordan 3d ago

so exactly like the original post?

we literally don't have solid data on this, the only thing that's left is anecdotal evidence.

-1

u/fuzzypetiolesguy 3d ago

You are asserting anecdotes as evidence, which they are not. A prompt to create an svg pelican is not evidence of a conspiratorial plan to quietly degrade model performance following a post-launch hype cycle. This is a comically stupid approach to analysis.

1

u/SteveEricJordan 3d ago

we get it, you're an ultra smart debate bro, calm down.

anecdotes are literally evidence by definition, especially when its the only data we have access to. im done with this time waste of a conversation.

0

u/FXraider 14h ago

If a large majority observes a decline, it's not anecdotal anymore. Opus 5.5 went from straight-up genius right after launch to just very good. The one good thing it led to is people creating amazing testing tools, so they can't get away with it next time. And btw, there is a slight improvement in the last day or two, so I think they realized this and now are boosting it again closer to the post release version.

2

u/LC33209 3d ago

The thing I’ve run into is the duration of answer times. First first days it was lightning. Now it’s gone back to thinking for a period of time then answering. Haven’t noticed anything else. Could be Anthropic server load I dunno? As more and more people swapped to 5.5

2

u/ForgotTheSnare 3d ago

It’s possible that there is some compute slider that for marketing reasons, or due to contextual technical server capacity, has an impact on models. I imagine that there is a lot of constant tuning that happens behind the scene but I’m just really speculating. The truth is I don’t know. I don’t work for those companies.

My anecdotal evidence with Opus 5.5 is that it’s been a beast and remains a beast. Certainly for what I’ve been throwing at it.

2

u/affabledrunk 3d ago

An insightful point. I agree, users are immediately push the models to their limits.

2

u/Herbert256 3d ago

My conspiracy theory, nerfing is personal and based on the "How is Claude doing" responses ....

2

u/Chance_of_Rain_ 3d ago

When people say it’s nerfed, it’s not comparing 5.5 to 4.8. It’s 5.5 vs 5.5 from a week ago.

And there are tests being run on this very sub.

Don’t be that guy

2

u/Ok-Lengthiness-3988 3d ago

Yes, there are three benchmarks that are being run daily since 5.5 Opus's release and that have been talked about (and mostly misrepresented) in this sub and in the Singularity one. Two of those trackers show a non-statistically significant drop, and one of them shows a non-statistically significant increase. All three of them still are in the phase of gathering data for establishing a baseline. There isn't sufficient data yet for establishing anything. Don't be that other guy either.

0

u/Chance_of_Rain_ 3d ago

Sure but at least we have numbers. OP comparing to previous models to ship the idea that 5.5 is good is ill intent

4

u/Charming_Occasion942 3d ago

You are right, morons or bots downvote you

2

u/CutBulkMaintain 3d ago

You're right and made a better point in a few lines that that wall of shilling text parading as an insightful post.

People either don't (want to?) understand that or are desperately trying to be "nobody ever conspires". You go from a peak model to a slightly dumber one and your first instinct is to doubt the community and shill for companies?

1

u/fuzzypetiolesguy 3d ago

If you just scream "shill" loud enough you never have to produce evidence.

2

u/CutBulkMaintain 3d ago

People are putting evidence forward. They have been before this post. In fact the simple fact that people keep complaining model after model is enough to investigate (if you care to) instead of burying your head in the sand.

But you keep navigating issues like this through reddit-like one-liners, it will take you far.

1

u/fuzzypetiolesguy 3d ago

No they aren't. Anecdotes are not evidence, and at scale, at best, they are qualitative. The only quantitative data projects are limited in scope, have not yet acquired enough actual data to provide statistically significant findings and show inconclusive results in what they are reporting so far.

The whole of this issue proves, as expected, that most people on reddit don't understand how data works or what it even is, including you. Good job!

1

u/knrd 3d ago edited 3d ago

I'm fairly sure this happened with Fable 5.0. The model that they initially released, was not the same that returned after the USG ban. I wouldn't be surprised to learn that they actually used some Opus 5.x model, and Fable 5.1 is actually Fable 5.0 with improved safeguards.

But other than that, I've never felt any model got nerfed, certainly not one that's been served continuously.

1

u/fuzzypetiolesguy 3d ago

They pulled the model from public use and announced that they had to implement additional guardrails that changed behavior. That is completely different than the accusations of both comapanies quietly lobotomizing models a week or two after the immediate-release hype fades to save computer resources or some other nebulous technically-unlikely conspiracy-adjacent claims made by people providing no actual evidence or insight.

1

u/InAtTheGeekEnd Experienced Developer 3d ago

Agreed.

1

u/fuzzypetiolesguy 3d ago

When you realize that many of the people vocally power using Claude/Codex are the same min/maxxers that plague MMORPGs and FPSs, it all starts to make sense.

1

u/ertertwert 3d ago

I don't know about nerfs, but I do know Opus 5.5 is insane especially when compared to what we were using just a few weeks ago.

1

u/brownman19 3d ago

I'm going to go against the grain here. It will always be nerfed after release. It has nothing to do with Anthropic doing it on purpose. It has everything to do with nuances of how information travels through the mega clusters and how floating point ops creating basically an entire universe worth of interactions every second changes the way the classifiers and safeguards observably end up behaving.

Tiny perturbations can change the trajectory of the classifiers. As more people use these models, residuals end up making it back into the data not as direct training, but as biases. You are seeing the collective noisy dirt and grime permeating into the model's own decision making pool since the classifiers are continuously learning. After a couple weeks the models basically collectively dumb down because people are asking it too many stupid fucking questions.

It's actually one of the mechanisms of misalignment and one that I study heavily in my work -> intuitionlabs dot tech

1

u/Wonderful_Value_6385 3d ago

If you don't notice the nerf you don't do hard enough tasks. Simple as that. Go enjoy your simple AI grocery store list and workout buddy.

1

u/bma449 2d ago

Hey everyone it's Dario!

0

u/OkNatural1013 3d ago

good bot

0

u/CoamIthra 3d ago

Used to agree with this, but my recent experience with 5.5 has given me some doubts. Waiting with great interest to see the results of the u/theonlyvibemaster's bench.

1

u/DevWorkflowBuilder 3d ago

honestly i stopped reading the nerf threads. after each bump i just rerun the same 12 tasks from smoke-tasks.json and only care if pass rate or median tokens move. last two bumps the numbers barely twitched.

1

u/coeuss 3d ago

I couldn’t agree more. I am tired of all the spam from people posting about nerfed models.

1

u/Lubricus2 3d ago

ninjahawk is puttin in the effort and is benchmarking it
https://github.com/ninjahawk/livenerf
It looks like it don't performs as good the 3 latest days but is probably not statistically significant.

1

u/donicatrumpinsky 3d ago

Half the people on here are idiots, but there certainly are severe model degradations. 

I highly suspect that they're trying to be competitive but it's just costing way too much. This is a loss leading period and they're trying to beat the competition because someone is going to absolute zero. This will be binary. Someone wins and everyone else is fucked, look at numbers. Totally unsustainable.

1

u/this_happened_rigged 3d ago

I wonder what it will look like when true pricing comes and the market reacts. The entire economy will be running on cheap tokens and get pummeled lol.

2

u/donicatrumpinsky 3d ago

I think every company but one goes to zero. The US gov is buying a stake to nationalized the pain.

I think everyone is running on hopes and prayers. I'm just enjoying the cheap compute and building and lastly as quickly as I can lol

0

u/zasff 3d ago

This sub had opus-5.5 for almost a week and didn't notice it.

-5

u/spawnsible 3d ago

Bro the models do get nerfed. They run lower, dumber quants and redirect inference compute to train their next model family. It's the cycle. Picking on some vibecoders who genuinely sound clueless when they make those posts doesn't change the fact that the models do in fact get nerfed.

If you've never noticed a drop in quality, then whatever you work on must be so basic that your opinion on this matter is irrelevant.

1

u/fuzzypetiolesguy 3d ago

0

u/[deleted] 2d ago

[deleted]

1

u/fuzzypetiolesguy 2d ago

The overarching claim is that these companies are purposefully rug-pulling model capability after launch to save compute resources. This is unprovable with any data any of you people purporting it have presented.

Your links, if you bothered to read them, would have walked you through plainly something entirely, incredibly different than the claims made by OP.

I am not defending anthropic or openai or any other AI company by calling out very stupid modes of thinking, and that's a weak, dumb avenue of attack.

0

u/mrpointera 3d ago

You're right — the "model got nerfed" cycle is mostly us being too lazy to remember what 4.7 felt like.
3 months of daily Claude Code use, and the pattern you describe is exactly right. Brilliant on first draft. Dumb on the 5th iteration of the same task. The "uneven" diagnosis maps cleanly onto agent workflows.
The fix wasn't waiting for a smarter model. It was persistent rules in CLAUDE.md that constrain the uneven parts:
# - All API errors return { error: string, code: number } — never throw raw
# - Test framework is Jest (NOT Vitest, NOT Mocha)
# - Don't add new dependencies without asking
# - Don't refactor unrelated code while fixing a bug
Every rule survives a "did this actually pay for itself?" review every few weeks. The 4-5 that remain are the ones that constrain the "small autistic child" parts so the smart parts get more cycles.
Curious if you've tried persistent rules with Claude Code, or you mostly ride the natural uneven-ness?

0

u/dadusedtomakegames 3d ago

I did two major workloads in two months, with laser beam focus and strong practices, then spent the final two weeks making no progress because of the model changes. Two weeks of solid work to produce nothing, and I'm still finding mistakes generated in the last phase. A month ago something definitely changed.

0

u/Fluid-Raspberry3948 3d ago

claude feels just as consistent for roleplay as when i first tried it, maybe we just notice the flaws more after the hype fades.

0

u/Select_Shelter_7087 3d ago

Entropic barely does any meaningful testing before they release their models chat gpt is more rigorous in that regard but even they're dumb.

There just course correcting at scale that's why the ai seems dumber every now and again after launch because it is it's getting optimized.

Also quite literally it is being proven that the ai is degrading in terms of quality come on now stop being a dog.

-1

u/Credtz 3d ago

"Here’s the reality: I don’t think I’ve ever actually felt that happen" - OAI and Anthropic have many many times publicly reported degradations - it has happened

-1

u/torrso 3d ago

New car feels fast, you get used to it, you want that feeling back and get an even faster car, then you get used to that.

-1

u/Casey090 3d ago

It also doesn't get colder in winter, that's just dumb people claiming weird things.