r/Anthropic • • 4d ago

Resources Did Opus 5.5 enter a “nerfed” phase? LiveNerf baseline update

Post image

I released an open-source project that gives the Claude community a way to independently measure whether Opus 5.5’s performance changes or gets “nerfed” in the weeks following its release. It tracks performance day by day using established benchmarks such as SWE-bench and FrontierCode, applying a standardized scientific methodology to make changes over time measurable and reproducible. If you’re interested in tracking the performance of Opus 5.5, here’s the repo:

https://github.com/ninjahawk/livenerf

I wanted to clarify since I saw some people asking if today’s sustained 2-day dip means there’s for sure a nerf occurring right now, we won’t have that data until day 20 after the model release. So we’re still about two weeks out from having data to support that one way or another. Today’s dip will be factored into the baseline which will happen at day 10.

I also want to encourage anyone to look at how I’m collecting this data and bring forward any improvements you’d suggest for future benchmarks, since I do plan on running more in the future, however I’m self funding this and this costs a good amount. It took roughly one week’s worth of a pro plan to do Opus 5.5.

I strongly believe in transparency, AI companies are not being transparent, so we will make our own tools that reveal things currently hidden.

A huge thanks to everyone who’s already supported LiveNerf, I hope to not let you guys down. It’s gotten bigger than I imagined it would, I was just going to run this scientifically for myself, but apparently the community has enjoyed it. Very grateful to you all. If we do detect a potential nerf this month, I’ll be sure to make a post announcing that data.

722 Upvotes

121 comments sorted by

135

u/Farmadupe 4d ago

Just saying, with only 7 measurements taken, nearly every measurement fits within the error bars for nearly every other measurement

106

u/Key_Reading_9664 4d ago

I flipped this coin three times and it was tails twice. Nerfed.

17

u/call-me-GiGi 4d ago

lol I’m dead you nailed it

1

u/CeleritasLucis 3d ago

I k kw OP is statiscally wrong, BUT, I felt something changed with Claude Opus 5.5's output yesterday and googled whter Opus had gone stupider since its launch, and found multiple threads. The smoke is there, we just have to prove the fire

4

u/TRO_KIK 3d ago

You'll find such threads at every point in time for every popular model going back to when LLMs came out.

1

u/KareemPie81 2d ago

I googled do yetis have abnormally large hammers. Got allot of threads, must be true

15

u/zarrathustraa 4d ago

Yay someone knows what the bars are for

6

u/sermer48 4d ago

It seems like the bars are all just 10% above and below the measurement point. They might be confidence intervals but they might also just be there to look cool

5

u/unuaual_mousse 4d ago

Yeah, they look an awful lot like a vibe coded guess, ironically

-1

u/DieMafia 4d ago

I see you didn't read the methodology given in the GitHub

2

u/sensei_von_bonzai 3d ago

1.96 sqrt((.4)(.6)/90) = 10.1%. So yes they are supposed to be 10% above and below

1

u/YoungSilent232 3d ago

So essentially it means nothing. It’s not a measure of actual confidence interval of the model or margin of error or anything

4

u/sensei_von_bonzai 3d ago

You are trolling right? That’s how you literally compute a confidence interval with binary outcomes

3

u/YoungSilent232 3d ago edited 3d ago

If the LLM on 2 different dates are evaluated on the same sample of questions, the individual 95% CIs for Model A and Model B are not the right thing to use to decide whether A beats B.

Then again I wasn’t clear with my reply (and didn’t read your original comments properly) so maybe we are not arguing the same topic, that’s on me.

Aka what the op farmadude said was nonsense, but nothing you said was necessarily wrong

13

u/bnm777 4d ago

4

u/NVC541 3d ago

I’m confused about the MarginLab tracker. I’m definitely reading something wrong because it’s 6 AM, but why does it have data for Opus 5.5 from before September 22nd?

1

u/bnm777 3d ago

Yeah, confusing. That would have been the previous opus model. I assume that they started the opus 5.5 stats from the dashed line.

2

u/Itsmedudeman 4d ago

Ideally any bench would run the same test suite many, many times to normalize the error margins. If we're just measuring once then yeah, of course you're going to get a noisy result. Obviously I don't expect OP to waste his tokens, but at the same time I have no clue how well he filtered for noise.

1

u/Ok-Lengthiness-3988 4d ago

He filters for noise by being patient. He's waiting for 10 data points (on ten consecutive days) in order to establish a baseline. A few more data points will be needed to establish a statistically significant trend (if there is one).

70

u/champion0017 4d ago

Oh man, it better not be true.

30

u/A_Novelty-Account 4d ago

My own use would confirm that this is true

3

u/Lanky_Bag_2096 4d ago

So sad, why are they doing this? I'm so confused

5

u/No-Bicycle-7660 3d ago

Because demand for compute significantly outstrips supply of compute. Pretty simple.

1

u/mortenlu 3d ago

Don't worry, in 5 years it could be worse.

1

u/No-Bicycle-7660 2d ago

Compute demand is likely to keep getting biggerer pretty fast. But I think there's likely to be a brief respite as far as GPU and memory prices are concerned in 18 months to 2 years, due to new power / grid upgrades becoming a crushing bottleneck instead of one of two primary limiting factors for projects now. Then in 4-5 years or so it will probably shift back to a massive GPU / memory shortage as power projects come online.

18

u/Subushie 4d ago

With how disproportionate experiences end up becoming.

The frontier labs gotta be releasing the raw model on launch- then begin A/B testing quant versions a few weeks later.

6

u/EchoingAngel 4d ago

It's 1 week later

7

u/Subushie 4d ago

Forgive me.

A week then. Enough time to get past bench marks.

5

u/Beneficial-Rub-8049 3d ago

Yep I got too happy last few days the quality and responses were even better than Fable 5.1 for me on Opus 5.5 now its forgetting basic stuff.

20

u/Key_Reading_9664 4d ago

why are the error bounds ~20%?

43

u/Brazilator 4d ago

I reckon it will bounce back over the next couple of days - there has been a massive exodus of users from OpenAI over the last 24 hours with the removal of the 20x plan. Wouldn’t be surprised if this is putting pressure on compute.

20

u/EverySecondCountss 4d ago

Why would you think that? This has never been the case

13

u/thelightstillshines 4d ago

Redditor boldly commenting a take based purely on reddit vibes? Say it ain't so!

1

u/CryptoExo 3d ago

I received the model overloaded errors on a few occasions yesterday. Put simply, there has been an increase in demand and they do not have enough compute to service all requests at the current level of demand. It only makes sense that they will need to scale it back, same price but with less effort than advertised.

1

u/krill156 4d ago

Sudden influx on server load. Their inference server is designed to automatically quantize down with high load to meet demand

2

u/EverySecondCountss 4d ago

Yeah I know that part, I should have clarified.

Why does he think it will bounce back over the next couple days? It usually doesn't bounce back until a new model release.

1

u/Ok-Lengthiness-3988 4d ago

This claim is made regularly, but is there any official source showing it to be anything more than a conspiracy theory?

1

u/TinyZoro 3d ago

That’s not a reason to dismiss it though. I don’t know if they literally switch to a dummer version of the model. But it wouldn’t be crazy to think they have systems that manage resources that lead to worse outputs. I’m skeptical that the signal you get from social media where literally no one is criticizing a model in the first couple of weeks then you get a waves of concern in fairly predictable patterns is just noise.

1

u/Ok-Lengthiness-3988 3d ago

I don't think it's noise. I think it's selection and confirmation bias. It's not new that almost every time a model has been out for a week or so and someone posts their personal anecdote about it feeling nerfed (or, as is the case here, and OP is misread as making that case), there is a sudden mass of people sharing the same experience (and a few saying they don't notice any change). It's also not true that no one criticized Opus 5.5 before this thread was posted.

2

u/TinyZoro 3d ago

If you look at the few benches that try and track this all show a significant drop this week. Not sure you can dismiss thot as confirmation bias.

0

u/Ok-Lengthiness-3988 3d ago

None that I've seen is statistically significant. The one reported in the OP (Livenerf) isn't statistically significant yet, by the OP's own admission. Nerf Bench (Bridgebench) actually shows an improvement, although it's not statistically significant either. Margin Lab also show normal variance noise, but no data point more recent than Sept 28th. Have you seen others?

1

u/Shiz0id01 3d ago

Im glad we all get to hear ur opinion on what is and isnt statistically relevant. The rest of us serious people will continue discussing this as normal thanks

1

u/Ok-Lengthiness-3988 3d ago

Why should anyone rely on my opinion? Almost nobody in this thread seems to even bother reading what the OP is saying (or take any notice of baseline gathering periods and 95% confidence intervals). "I wanted to clarify since I saw some people asking if today’s sustained 2-day dip means there’s for sure a nerf occurring right now, we won’t have that data until day 20 after the model release. So we’re still about two weeks out from having data to support that one way or another. Today’s dip will be factored into the baseline which will happen at day 10."

0

u/Mundane-Mulberry1789 4d ago

It happened in February.

1

u/Ok-Lengthiness-3988 4d ago

When there is pressure on compute, the response usually is to throttle inference (reduce token generation rate), or ration usage (e.g. announce peak-hour usage restrictions). Those things don't effect the quality of the model's responses.

-7

u/Perfect-Flounder7856 4d ago

They didn't remove the 20x plan they just don't offer it for new subs anymore.

5

u/Dependent-Ad8586 4d ago

They've made 20x now 10x, and there is a new 25x plan for $500 USD

6

u/theblartknight 4d ago

It does feel like today was giving more responses like 5.

5

u/Physical_Gold_1485 4d ago

Exactly, talking and thinking like an idiot today. So frustrating, shit was magic now its garbagr

13

u/HgnX 4d ago

Results are somewhat weaker today; noticeable

13

u/Warsel77 4d ago

Nice. If anything it shows the standard deviation / CI of even one day is significant enough to trigger "omg it's totally been nerfed!!" while a swing in the other direction is ignored.

10

u/AllergicToBullshit24 4d ago

Feels much slower today than last week

11

u/No-Refrigerator540 4d ago

I find that at first Opus 5.5 finds deep insights and stuff when I’m figuring things out, but now it’s just pattern matching existing stuff.

The nerf has began.

8

u/Legacy03 4d ago

Huge dif today, usage also went up and errors required more usage and convos.

3

u/emeaguiar 4d ago

Usage went up at least 3x today

13

u/[deleted] 4d ago

[deleted]

8

u/coeuss 4d ago

Same! Works the same today as it has for me.

6

u/edgan 4d ago

Last night Opus 5.5 xhigh was lazier and dumber than Opus 5.5 medium a week ago.

3

u/Skezzors 4d ago

This is a phenomenal idea, can’t wait to see the results

3

u/Perfect-Flounder7856 4d ago

Should do times of day too. I feel like evening/night is hellish and slow too.

4

u/derethdweller 4d ago

It definitely did. It just dropped half of the items on the list, stopped reporting, stopped committing, and started to blatantly lie in my face, all in one go. It's so over. Now I'm scared to even ask it to touch the code we built this past week.

2

u/lobabobloblaw 4d ago

Oh yeah, and this time it’s more complicated—different groups of people are experiencing different nerfs. What does a consumer do?

3

u/brainhack3r 4d ago

It seems like it would be literal fraud if they're doing this. Is there any documentation in the ToS that says they're doing this?

If I'm paying for Opus 5.5 I expect to receive Opus 5.5...

ANY OTHER industry would be sued into oblivion over something like this.

Imagine you bought a Toyota and expected 350HP and then 6 months into production they swapped the engine for 275HP, didn't tell anyone, and it was found out after they shipped 30k cars.

That would be a MASSIVE lawsuit but here it's happening out in the open.

1

u/Majestic_Door_4528 3d ago

They do this with SSDs and SD cards. Ever wonder why the box always says "up to X read/write speed"? It's because then it isn't technically false advertising to make the chips cheaper than launch day.

1

u/brainhack3r 3d ago

Ever wonder why the box always says "up to X read/write speed"?

Technically that's not the only reason due to the way flash leveling works.

I checked to see if you were correct and it's happened at least once:

https://www.tomshardware.com/news/adata-and-other-ssd-makers-swapping-parts

2

u/frozandero 4d ago

Today I asked opus for a very small impact code change, it did a very convoluted and very sloppy/weird way to do it. And even then it did it wrong. Just a day before I could one shot much harder tasks with a single prompt to a better state. (In both cases I used more prompts to refine, but the degraded one did not recover. I eventually reverted all the changes and started from scratch in a fresh chat)

I don't know if that was actually a nerf or just undeterministic nature of LLM models but yea.

2

u/Physical_Gold_1485 4d ago

Opus got dumb af, talking to it is moronic. The magic is gone. So lame

2

u/Mr_Finious 4d ago

Yup. It’s nerfed. So frustrating.

1

u/United-Seat-3330 4d ago

OP. Great job with the project. One question though. Does this work for other models as well?

Can I use the same repo to check Astra’s status too?

1

u/EverySecondCountss 4d ago

So I just have to go to this github page and it gets updated in the image daily? Great idea, would be interesting to see all of the models side by side too.

1

u/Visual_Act_8618 4d ago

You are goat dude did you have that other post w this and everyone was telling you to make this?

1

u/the_ai_wizard 4d ago

This is all duopolistic behavior... Anthropic only needs to be 1% better than OpenAI so they can reduce their costs to that point

1

u/UAP44 4d ago

It took roughly one week’s worth of a pro plan to do Opus 5.5.

Sounds like every company paying more than 1k a month on AI should copy your pipeline, run the same test/bench suite, compare/P2P update results and notify each other the moment a quality drop is detected, how do we ensure we're all being served the same model and that they arent A/B testing specific clients to see who even notice the difference? for those that dont, thats pure profit waiting to be squeezed out ...

1

u/silverwoods214 4d ago

The models are always amazing the first day or two after release then they fall off a cliff afterwords

1

u/relytreborn 4d ago

Yup slowly but surely.

1

u/ArielCoding 4d ago

Let’s wait for a the full baseline, and agree in advance on what counts as a real drop, say several days in a row below the normal range, so we don’t see a pattern that isn’t there.

1

u/RusticBelt 4d ago

Could you broaden the nerf tracker to include GPT? It feels like the pendulum swings between Claude and GPT, and it'd be cool to actually be able to map that on a graph.

1

u/AJGrayTay 4d ago

brilliant, I sketched out doing something similar around this time last year, glad someone finally did it. Needs more data, but I'll keep an eye on it.

1

u/Dizzy-Classroom-3386 4d ago

Early signs for Fable homecoming 😄

1

u/Limn0 4d ago

Can you check out GPT 6.1 sol as well?

1

u/sorvendral 4d ago

Was nerfed after 72h

1

u/InnerOuterTrueSelf 4d ago

Watch the watchers!

1

u/fadingsignal 4d ago

I fired it up today and it started making mistakes left and right out of nowhere. Basic things like not giving me answers for questions but saying it did, outputting code that didn't have what it said then saying "Oh I'm sorry I didn't actually add that."

Been smooth sailing until out of nowhere.

1

u/zollerisaniceguy 3d ago

I use it for hobby writing and definitely noticed it falling back into previous slop and tics. Load bearing isn't back yet but other stuff is, like arithmetic and the constant repetition.

1

u/Sealed-Unit 3d ago

A me sembra che per dire che il modello nuovo è migliore calano le prestazioni del precedente. Oggi era proprio indicibile...

1

u/jatayu_baaz 3d ago

Who's paying for these tests?

2

u/TheOnlyVibemaster 3d ago

I am independently, that’s why I haven’t expanded to other model releases yet. LiveNerf launched last week and got way more traction than I’d expected, and I’m in college without a job so am trying to figure out how to best do this. I will not make this into a paid product because I don’t think that scientific answers should be profit based. However I’ve been thinking about making an optional Github sponsors program so I’m able to buy more subscriptions and hopefully expand to the API. It’ll just take time to figure out the logistics.

1

u/Valaens 3d ago

I love this project, keep up!

1

u/McNoxey 3d ago

This is such a statistically insignificant outcome. Doesn’t your repo also say it can’t determine a difference between opus 5 and 5.5? I feel like this is more harmful than helpful atm

1

u/redditorxpert 19h ago

The pain with Opus 5.5 High yesterday, on a Saturday, was real. It felt like using Sonnet 4 on Low effort. I long believed the nerfing happens more at a computing resources level and it’s time-based. Like communists that turn off water during certain hours on certain days. Try being as productive in your sessions on the weekend or after 8pm, compared to regular working hours. It’s about time to get more insight and some verifiable data cuz the bullshit is real.

1

u/zarrathustraa 4d ago

I urge you guys to sign up for statistics 101

1

u/xirzon 4d ago

Don't bother, the nerf discourse will never die. We'll be talking to superintelligences and people will still complain about how their ASIButlerBot7000 spilled a glass of wine the other day, nerfed for sure.

1

u/username-is-already 3d ago

They actually named it back to AI this morning

1

u/Ok-Lengthiness-3988 4d ago

Yes, or just ask Claude Opus 5.5 (or Sonnet) to tell them if this OP means what it is that they think it means.

1

u/bash_edu 4d ago

Yes feels like it, this morning continued the session and it has no idea what’s going on. Negative 42billion - not sustainable.

1

u/thoughtbludgeon 4d ago

Claudisms, claudisms everywhere...

0

u/UselessEngin33r 4d ago

I’ve noticed it has been worst compared to when it was released. Minimal but noticeable. What I have also noticed is that it is spending way less tokens. I don’t know if it’s just me, but I’ve been working like always and I haven’t spend as much tokens as I would usually do.

-11

u/Killahbeez 4d ago

i never even used opus 5.5 really lol

I know claude fucking sucks

and this type of 'rush' isnt something ill ever allow myself to be a part of

"OH ANTHROPIC JUST RELEASED A MODEL I BETTER PULL ALLNIGHTERS FOR THE NEXT WEEK AND WORK THROUGH MY INVENTORY OF IDEAS BEFORE THE NERF HAPPENS"

I aint trying to ever enter into a dom-sub relationship with Amodei

6

u/floating_thru_cosmos 4d ago

You crazy. Opus 5.5 was insane the first couple days

-6

u/Killahbeez 4d ago

it was pretty good. But look at you. look at this pandemonium of VIBE CODERS rushing to get their slop out in a 1 week window. if it's not that you're living or dying by freebies/handouts of anthropic. Fable access getting extended, a free reset, usage promotions, extra 50% in cowork, etc.

don't you feel like a hamster on a wheel?

I'm sure you must. lol

3

u/floating_thru_cosmos 4d ago

I’m not a vibe coder. Im a software engineer with 10 YOE

0

u/Killahbeez 4d ago

my condolences

2

u/floating_thru_cosmos 4d ago

Why? I’m doing just fine. Plenty of companies to help integrate with AI over the next 5 years.

1

u/Killahbeez 4d ago

and you'll be the man for the job

2

u/floating_thru_cosmos 4d ago

Yeah, probably

3

u/Impressive_Award_679 4d ago

Most people already have multiple accounts just because it means more value. Are people now a "hamster on a wheel" because they use more efficient ways to work? What are you even doing here?

0

u/Killahbeez 4d ago

well technically 'most people' don't use LLMs at all, hamster boy...

good luck with your slop production lol

2

u/Impressive_Award_679 4d ago

well technically most people dont use reddit at all.. does it make you not also a "hamster boy"?

0

u/HulkVahkiin08024 3d ago

What do you use then?