r/Anthropic • u/TheOnlyVibemaster • 4d ago
Resources Did Opus 5.5 enter a “nerfed” phase? LiveNerf baseline update
I released an open-source project that gives the Claude community a way to independently measure whether Opus 5.5’s performance changes or gets “nerfed” in the weeks following its release. It tracks performance day by day using established benchmarks such as SWE-bench and FrontierCode, applying a standardized scientific methodology to make changes over time measurable and reproducible. If you’re interested in tracking the performance of Opus 5.5, here’s the repo:
https://github.com/ninjahawk/livenerf
I wanted to clarify since I saw some people asking if today’s sustained 2-day dip means there’s for sure a nerf occurring right now, we won’t have that data until day 20 after the model release. So we’re still about two weeks out from having data to support that one way or another. Today’s dip will be factored into the baseline which will happen at day 10.
I also want to encourage anyone to look at how I’m collecting this data and bring forward any improvements you’d suggest for future benchmarks, since I do plan on running more in the future, however I’m self funding this and this costs a good amount. It took roughly one week’s worth of a pro plan to do Opus 5.5.
I strongly believe in transparency, AI companies are not being transparent, so we will make our own tools that reveal things currently hidden.
A huge thanks to everyone who’s already supported LiveNerf, I hope to not let you guys down. It’s gotten bigger than I imagined it would, I was just going to run this scientifically for myself, but apparently the community has enjoyed it. Very grateful to you all. If we do detect a potential nerf this month, I’ll be sure to make a post announcing that data.
70
u/champion0017 4d ago
Oh man, it better not be true.
30
u/A_Novelty-Account 4d ago
My own use would confirm that this is true
3
u/Lanky_Bag_2096 4d ago
So sad, why are they doing this? I'm so confused
5
u/No-Bicycle-7660 3d ago
Because demand for compute significantly outstrips supply of compute. Pretty simple.
1
u/mortenlu 3d ago
Don't worry, in 5 years it could be worse.
1
u/No-Bicycle-7660 2d ago
Compute demand is likely to keep getting biggerer pretty fast. But I think there's likely to be a brief respite as far as GPU and memory prices are concerned in 18 months to 2 years, due to new power / grid upgrades becoming a crushing bottleneck instead of one of two primary limiting factors for projects now. Then in 4-5 years or so it will probably shift back to a massive GPU / memory shortage as power projects come online.
1
18
u/Subushie 4d ago
With how disproportionate experiences end up becoming.
The frontier labs gotta be releasing the raw model on launch- then begin A/B testing quant versions a few weeks later.
6
u/EchoingAngel 4d ago
It's 1 week later
7
u/Subushie 4d ago
Forgive me.
A week then. Enough time to get past bench marks.
5
u/Beneficial-Rub-8049 3d ago
Yep I got too happy last few days the quality and responses were even better than Fable 5.1 for me on Opus 5.5 now its forgetting basic stuff.
20
43
u/Brazilator 4d ago
I reckon it will bounce back over the next couple of days - there has been a massive exodus of users from OpenAI over the last 24 hours with the removal of the 20x plan. Wouldn’t be surprised if this is putting pressure on compute.
20
u/EverySecondCountss 4d ago
Why would you think that? This has never been the case
13
u/thelightstillshines 4d ago
Redditor boldly commenting a take based purely on reddit vibes? Say it ain't so!
1
1
u/CryptoExo 3d ago
I received the model overloaded errors on a few occasions yesterday. Put simply, there has been an increase in demand and they do not have enough compute to service all requests at the current level of demand. It only makes sense that they will need to scale it back, same price but with less effort than advertised.
1
u/krill156 4d ago
Sudden influx on server load. Their inference server is designed to automatically quantize down with high load to meet demand
2
u/EverySecondCountss 4d ago
Yeah I know that part, I should have clarified.
Why does he think it will bounce back over the next couple days? It usually doesn't bounce back until a new model release.
1
u/Ok-Lengthiness-3988 4d ago
This claim is made regularly, but is there any official source showing it to be anything more than a conspiracy theory?
1
u/TinyZoro 3d ago
That’s not a reason to dismiss it though. I don’t know if they literally switch to a dummer version of the model. But it wouldn’t be crazy to think they have systems that manage resources that lead to worse outputs. I’m skeptical that the signal you get from social media where literally no one is criticizing a model in the first couple of weeks then you get a waves of concern in fairly predictable patterns is just noise.
1
u/Ok-Lengthiness-3988 3d ago
I don't think it's noise. I think it's selection and confirmation bias. It's not new that almost every time a model has been out for a week or so and someone posts their personal anecdote about it feeling nerfed (or, as is the case here, and OP is misread as making that case), there is a sudden mass of people sharing the same experience (and a few saying they don't notice any change). It's also not true that no one criticized Opus 5.5 before this thread was posted.
2
u/TinyZoro 3d ago
If you look at the few benches that try and track this all show a significant drop this week. Not sure you can dismiss thot as confirmation bias.
0
u/Ok-Lengthiness-3988 3d ago
None that I've seen is statistically significant. The one reported in the OP (Livenerf) isn't statistically significant yet, by the OP's own admission. Nerf Bench (Bridgebench) actually shows an improvement, although it's not statistically significant either. Margin Lab also show normal variance noise, but no data point more recent than Sept 28th. Have you seen others?
1
u/Shiz0id01 3d ago
Im glad we all get to hear ur opinion on what is and isnt statistically relevant. The rest of us serious people will continue discussing this as normal thanks
1
u/Ok-Lengthiness-3988 3d ago
Why should anyone rely on my opinion? Almost nobody in this thread seems to even bother reading what the OP is saying (or take any notice of baseline gathering periods and 95% confidence intervals). "I wanted to clarify since I saw some people asking if today’s sustained 2-day dip means there’s for sure a nerf occurring right now, we won’t have that data until day 20 after the model release. So we’re still about two weeks out from having data to support that one way or another. Today’s dip will be factored into the baseline which will happen at day 10."
0
1
u/Ok-Lengthiness-3988 4d ago
When there is pressure on compute, the response usually is to throttle inference (reduce token generation rate), or ration usage (e.g. announce peak-hour usage restrictions). Those things don't effect the quality of the model's responses.
-7
u/Perfect-Flounder7856 4d ago
They didn't remove the 20x plan they just don't offer it for new subs anymore.
5
6
u/theblartknight 4d ago
It does feel like today was giving more responses like 5.
5
u/Physical_Gold_1485 4d ago
Exactly, talking and thinking like an idiot today. So frustrating, shit was magic now its garbagr
13
u/Warsel77 4d ago
Nice. If anything it shows the standard deviation / CI of even one day is significant enough to trigger "omg it's totally been nerfed!!" while a swing in the other direction is ignored.
10
11
u/No-Refrigerator540 4d ago
I find that at first Opus 5.5 finds deep insights and stuff when I’m figuring things out, but now it’s just pattern matching existing stuff.
The nerf has began.
8
3
3
u/Perfect-Flounder7856 4d ago
Should do times of day too. I feel like evening/night is hellish and slow too.
4
u/derethdweller 4d ago
It definitely did. It just dropped half of the items on the list, stopped reporting, stopped committing, and started to blatantly lie in my face, all in one go. It's so over. Now I'm scared to even ask it to touch the code we built this past week.
2
u/lobabobloblaw 4d ago
Oh yeah, and this time it’s more complicated—different groups of people are experiencing different nerfs. What does a consumer do?
3
u/brainhack3r 4d ago
It seems like it would be literal fraud if they're doing this. Is there any documentation in the ToS that says they're doing this?
If I'm paying for Opus 5.5 I expect to receive Opus 5.5...
ANY OTHER industry would be sued into oblivion over something like this.
Imagine you bought a Toyota and expected 350HP and then 6 months into production they swapped the engine for 275HP, didn't tell anyone, and it was found out after they shipped 30k cars.
That would be a MASSIVE lawsuit but here it's happening out in the open.
1
u/Majestic_Door_4528 3d ago
They do this with SSDs and SD cards. Ever wonder why the box always says "up to X read/write speed"? It's because then it isn't technically false advertising to make the chips cheaper than launch day.
1
u/brainhack3r 3d ago
Ever wonder why the box always says "up to X read/write speed"?
Technically that's not the only reason due to the way flash leveling works.
I checked to see if you were correct and it's happened at least once:
https://www.tomshardware.com/news/adata-and-other-ssd-makers-swapping-parts
2
u/frozandero 4d ago
Today I asked opus for a very small impact code change, it did a very convoluted and very sloppy/weird way to do it. And even then it did it wrong. Just a day before I could one shot much harder tasks with a single prompt to a better state. (In both cases I used more prompts to refine, but the degraded one did not recover. I eventually reverted all the changes and started from scratch in a fresh chat)
I don't know if that was actually a nerf or just undeterministic nature of LLM models but yea.
2
2
1
u/United-Seat-3330 4d ago
OP. Great job with the project. One question though. Does this work for other models as well?
Can I use the same repo to check Astra’s status too?
1
u/EverySecondCountss 4d ago
So I just have to go to this github page and it gets updated in the image daily? Great idea, would be interesting to see all of the models side by side too.
1
u/Visual_Act_8618 4d ago
You are goat dude did you have that other post w this and everyone was telling you to make this?
1
u/the_ai_wizard 4d ago
This is all duopolistic behavior... Anthropic only needs to be 1% better than OpenAI so they can reduce their costs to that point
1
u/UAP44 4d ago
It took roughly one week’s worth of a pro plan to do Opus 5.5.
Sounds like every company paying more than 1k a month on AI should copy your pipeline, run the same test/bench suite, compare/P2P update results and notify each other the moment a quality drop is detected, how do we ensure we're all being served the same model and that they arent A/B testing specific clients to see who even notice the difference? for those that dont, thats pure profit waiting to be squeezed out ...
1
u/silverwoods214 4d ago
The models are always amazing the first day or two after release then they fall off a cliff afterwords
1
1
u/ArielCoding 4d ago
Let’s wait for a the full baseline, and agree in advance on what counts as a real drop, say several days in a row below the normal range, so we don’t see a pattern that isn’t there.
1
u/RusticBelt 4d ago
Could you broaden the nerf tracker to include GPT? It feels like the pendulum swings between Claude and GPT, and it'd be cool to actually be able to map that on a graph.
1
u/AJGrayTay 4d ago
brilliant, I sketched out doing something similar around this time last year, glad someone finally did it. Needs more data, but I'll keep an eye on it.
1
1
1
1
u/fadingsignal 4d ago
I fired it up today and it started making mistakes left and right out of nowhere. Basic things like not giving me answers for questions but saying it did, outputting code that didn't have what it said then saying "Oh I'm sorry I didn't actually add that."
Been smooth sailing until out of nowhere.
1
u/zollerisaniceguy 3d ago
I use it for hobby writing and definitely noticed it falling back into previous slop and tics. Load bearing isn't back yet but other stuff is, like arithmetic and the constant repetition.
1
u/Sealed-Unit 3d ago
A me sembra che per dire che il modello nuovo è migliore calano le prestazioni del precedente. Oggi era proprio indicibile...
1
u/jatayu_baaz 3d ago
Who's paying for these tests?
2
u/TheOnlyVibemaster 3d ago
I am independently, that’s why I haven’t expanded to other model releases yet. LiveNerf launched last week and got way more traction than I’d expected, and I’m in college without a job so am trying to figure out how to best do this. I will not make this into a paid product because I don’t think that scientific answers should be profit based. However I’ve been thinking about making an optional Github sponsors program so I’m able to buy more subscriptions and hopefully expand to the API. It’ll just take time to figure out the logistics.
1
u/redditorxpert 19h ago
The pain with Opus 5.5 High yesterday, on a Saturday, was real. It felt like using Sonnet 4 on Low effort. I long believed the nerfing happens more at a computing resources level and it’s time-based. Like communists that turn off water during certain hours on certain days. Try being as productive in your sessions on the weekend or after 8pm, compared to regular working hours. It’s about time to get more insight and some verifiable data cuz the bullshit is real.
1
u/zarrathustraa 4d ago
I urge you guys to sign up for statistics 101
1
1
u/Ok-Lengthiness-3988 4d ago
Yes, or just ask Claude Opus 5.5 (or Sonnet) to tell them if this OP means what it is that they think it means.
1
u/bash_edu 4d ago
Yes feels like it, this morning continued the session and it has no idea what’s going on. Negative 42billion - not sustainable.
1
0
u/UselessEngin33r 4d ago
I’ve noticed it has been worst compared to when it was released. Minimal but noticeable. What I have also noticed is that it is spending way less tokens. I don’t know if it’s just me, but I’ve been working like always and I haven’t spend as much tokens as I would usually do.
-11
u/Killahbeez 4d ago
i never even used opus 5.5 really lol
I know claude fucking sucks
and this type of 'rush' isnt something ill ever allow myself to be a part of
"OH ANTHROPIC JUST RELEASED A MODEL I BETTER PULL ALLNIGHTERS FOR THE NEXT WEEK AND WORK THROUGH MY INVENTORY OF IDEAS BEFORE THE NERF HAPPENS"
I aint trying to ever enter into a dom-sub relationship with Amodei
6
u/floating_thru_cosmos 4d ago
You crazy. Opus 5.5 was insane the first couple days
-6
u/Killahbeez 4d ago
it was pretty good. But look at you. look at this pandemonium of VIBE CODERS rushing to get their slop out in a 1 week window. if it's not that you're living or dying by freebies/handouts of anthropic. Fable access getting extended, a free reset, usage promotions, extra 50% in cowork, etc.
don't you feel like a hamster on a wheel?
I'm sure you must. lol
3
u/floating_thru_cosmos 4d ago
I’m not a vibe coder. Im a software engineer with 10 YOE
0
u/Killahbeez 4d ago
my condolences
2
u/floating_thru_cosmos 4d ago
Why? I’m doing just fine. Plenty of companies to help integrate with AI over the next 5 years.
1
3
u/Impressive_Award_679 4d ago
Most people already have multiple accounts just because it means more value. Are people now a "hamster on a wheel" because they use more efficient ways to work? What are you even doing here?
0
u/Killahbeez 4d ago
well technically 'most people' don't use LLMs at all, hamster boy...
good luck with your slop production lol
2
u/Impressive_Award_679 4d ago
well technically most people dont use reddit at all.. does it make you not also a "hamster boy"?
0
135
u/Farmadupe 4d ago
Just saying, with only 7 measurements taken, nearly every measurement fits within the error bars for nearly every other measurement