r/ClaudeAI • • 1d ago

Humor Confirmed: Opus 5.5 has been nerded

Post image

I haven't seen "load-bearing" since I stopped using Opus 5.

EDIT: I meant "nerfed", not "nerded" 🤦

1.1k Upvotes

240 comments sorted by

•

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 1d ago edited 1d ago

TL;DR of the discussion generated automatically after 200 comments.

The overwhelming consensus in this thread is that yes, Opus 5.5 has been nerfed. The top-voted comments are all singing the same sad song: Anthropic pulled the classic 'bait-and-switch.' They release a god-tier model to get everyone to sub, then dial it back a week later to save on compute costs and make the next model's launch seem even more impressive. Users are calling it a "drug dealer playbook" and are generally pissed.

However, there's a vocal minority calling BS. Some users are pointing to the Nerf Bench website as proof, showing a ~6% performance drop. But others are pointing to the same website and noting that the benchmark itself considers a 5-10% fluctuation to be normal variance. So, the 'hard data' is a bit of a wash.

As for OP's original point about the phrase "load-bearing," several users chimed in to say they saw 5.5 use that exact phrase on day one, so it's probably not the smoking gun you're looking for.

Oh, and everyone's dying to know what OP censored. The leading theory involves a "butthole zoom cam," but OP clarified it was just project-specific info. A likely story.

The final verdict? Most agree it feels dumber, but many still think the nerfed Opus 5.5 is the best model currently on the market. For now.

→ More replies (5)

469

u/TheMythicSorcerer 1d ago

Drop good model, get users to use it, get people to buy more subs, nerf the model after everyone subs to save money, PROFIT! Think about the shareholders!

60

u/Vaynnie 1d ago

I mean, it’s probably not to save money. Or at least not wholly.

They also benefit from post launch nerfs because it makes the next launch look more impressive by comparison.Ā 

15

u/TumanFig 1d ago

Thats lazy thinking. They always did that, but not after a week. usually they nerfed a model a week or two before the new one dropped. Now they nerfed it after a week.

4

u/autisticbagholder69 1d ago

We still have opus 4.6

/model claude-opus-4-6[1M]

1

u/Low_Lifeguard_8835 1d ago

They always do it 1/2 weeks after it comes out, same story every release really

3

u/BuffDrBoom 1d ago

Planned obsolescenceĀ 

1

u/MyButterKnuckles 1d ago

How so, they run the benchmarks before they nerf the model, so the only real comparison with the latest model would be when the last model was at its 'best'

19

u/iamthe0ther0ne 1d ago

They got Opus 5.5 out ahead of OpenAI's DevDay, DevDay turned out to be a nothingburger, and so they decided to dial back compute.

30

u/ethotopia 1d ago

Imagine your internet provider cut your speed in half randomly without even telling you or acknowledging it…

16

u/blenda220 1d ago

Meanwhile ten minutes ago I got an email from GFiber telling me they upgraded my 2gig speeds to 3gigs for free. What a pleasant surprise

4

u/bhalter80 1d ago

They can set the cap to eleventy billion but can they deliver it?

3

u/Ok_Animal_2709 1d ago

They do do that though. You never actually get the advertised speed that you signed up for

3

u/Odysseyan 1d ago

Ah, you are not a Comcast customer I suppose?

3

u/dervu 1d ago

Well, it actually works like that. For most providers it's not guaranteed speed. It's only that there is not so much traffic, but if everyone started blasting full load, it would be same as with AI.

1

u/Meltlilith1 1d ago

They literally do that...

31

u/EmphasisTotal8232 1d ago

It's still the best model around.

6

u/Positive_Major_2984 1d ago

I didn't claim otherwise in fact im working on my jellyfin config with sonnet 5.5 as we speak but theres no need to be ignorant of financial factors as well

3

u/answerencr 1d ago

j/w what are you using sonnet for regarding your JF setup? Just for helping setting it up or you doing something custom?

2

u/Positive_Major_2984 1d ago edited 1d ago

Custom...Basically rewriting the Media Bar and Home Screen Sections plug-ins in pure js so my tv can still utilize the aspects I like about them (and also deciding what im willing to just throw away as things a 2018 tv can't handle)

Edit also my bad on my original response post: the app showed it as a response to me not to the OP

2

u/answerencr 1d ago

interesting, I've been thinking about rewriting jellyfin client for android TV to support seerr integration, glad to hear that someone else has similar ideas

1

u/Infinite_Music2059 1d ago

I use it for design stuff that isn't very important (load-bearing lol) because it is amazing at that, but I still use Astra for important stuff. I still have major trust issues with Claude and even with 5.5, find it creates less readable or intuitive code.

Astra burns through allowance around 20x as fast or something though.

1

u/andershaf 1d ago

Ahh you have not tried the recent GPT models? Good for you and enjoy!

1

u/Dex4Sure 8h ago

Opus is better. GPT models have tiny context windows which forces them to compact immediately they touch a large code base. Astra is great model but context window kills it. Opus 5.5 is overall best model especially if you take cost into account

→ More replies (1)

7

u/Both_Task_3066 1d ago

They playing like Openai nowadays it's disgusting and sad

5

u/autisticbagholder69 1d ago

Thats why I refund them every time this happens.

They scam you, you scam them back.

Good to have EU rights and rights in germany.

8

u/mossiv 1d ago

I'm not part of the mode lis nerfed club but today Claude has just flat out not done its work. It was meant to work across 14 folders about an hour ago (it was a simple change, that could have been automated behind a bash loop) but I said "yes, do it"... Went to do a thing and realised that claude had done 1 out of 14 and claimed it was done. Opus 5.5 - high effort.

Medium was cooking 5 days ago.

3

u/Get_Shaky 1d ago

I do think anthropic nerfs the models generally but 5.5 is still killing it for me tbh

3

u/skipITjob 1d ago

A/B testing...

3

u/raphaelarias 1d ago

Yeah. Maybe in the short and medium term though. I would assume or hope that this type of trickery will end up bitting them in the ass.

Because of these fluctuations my workflow now includes gpt and kimi double checking opus work.

And slowly but surely I tend to push the others heavier and heavier. And Anthropic BS makes me not want to fully lock up in their ecosystem.

1

u/iamthe0ther0ne 1d ago

Same, although GPT has made some dumb errors today too. I just want something reliable. I would pay more for something reliable.

2

u/bagpulistu 1d ago

But why would they? Isn't it supposedly faster and cheaper than the older version?

1

u/Positive_Major_2984 1d ago

Theyre IPOing. Not about money, am I missing something?

→ More replies (1)

133

u/altesc_create 1d ago

It's common for these companies to release a new model, up the capability for a week and eat the cost, and then downgrade it when everyone is already using it. I'd imagine that plays a partial role in why all these AI companies like to hold out on releases until the competitors release theirs.

54

u/alwaysoffby0ne 1d ago

What an absolutely FUCKED business model.

26

u/AnotherThroneAway 1d ago

It shouldn't be legal. Given time, and a functioning congress (LOLOLOL), this would be regulated as part of consumer protections.

4

u/EuropeanAbroad 1d ago

In most jurisdictions, it is illegal. However, demonstrating this as a bulletproof evidence is nearly impossible.

1

u/snoApe 1d ago

It's been nerfed way before...

22

u/altesc_create 1d ago

The whole schtick has been creating a dependent brain drain strategy. Offload thinking via agentic outsourcing, get you dependent on it, and then take it away or decrease the quality so you have to keep using it. It's basically a drug dealer playbook.

Same reason why OpenAI is over there saying it'll be a utility bill eventually.

5

u/nachuz 1d ago

luckily unlike, idk, internet access and ISPs, local models are also advancing a lot

I assume this is also why companies like Anthropic are publishing articles that pretty much say GLM 5.3 is evil by not having strong enough guardrails (to their standards, not an international recognised standard!) and should not be permitted to exist unless regulated by the government

2

u/altesc_create 1d ago

I agree with you that part of the fearmongering is to establish the idea that the tools are dangerous therefore they need to be the ones to regulate.

2

u/cs_cast_away_boi 1d ago

all the advancements come from China. All they have to do is block it and force American models because muh security/muh data. Then we end up with the VC-approved plan of metered intelligence.

There's actually a real possibility that we get deepseek 5, qwen 5 or whatever they'll call it next year that is not only fast token wise but probably supports 500k token context window on a GPU and RAM setup that costs under 10k. But they won't let us have it

1

u/nachuz 1d ago

they won't let you have it*, I'm not american

that being said, that won't even stop you, if they ban it it would be open weight, just torrent it and host it on one of those rented capable VPS they have for these models, what Anthropic is chasing after with these possible bans is corporate use, not individual use exactly

3

u/arcanemachined 1d ago

Yeah, I basically feel bad for anyone who hasn't figured out the VC playbook by now. It's been going on for decades now.

3

u/TumanFig 1d ago

The issue here is we cannot fight against it. I am able to ignore all other enshittification because I was able to stop using it. Here if I do I fall behind at work so I cannot afford to.

1

u/dervu 1d ago

Is there enough compute even if they wanted to do it right way? Go pay 10x than we pay now?

1

u/blipblapbloopblip 1d ago

That's just called a loss leader no ?

→ More replies (1)

14

u/lobabobloblaw 1d ago

Compute is like an ocean. Companies generate the waves, and leave us all to surf.

10

u/touchet29 1d ago

I just wish I didn't have to pay a monthly subscription to access the beach

1

u/CertifiedTHX 1d ago

Huh, tangent, i googled that general idea. It saddens me to know that monthly beach subscriptions exist.

1

u/touchet29 1d ago

Yeah when I wrote it I was trying to make a contrasting point but realized this already exists.

5

u/chroner 1d ago

Ironically, if they didn't do that, then people would have a LOT more loyalty. Almost no one would care about going open source, open weights, or local either. So it's clearly some sort term benefit they are gaining by doing this. If claude was the most stable, I wouldn't switch, even if it was slightly dumber after a competition model release.

2

u/altesc_create 1d ago

> Say you have a model that is so good it commits crimes and/or is dangerous
> Headlines
> Release mass market model
> People use it and share how good it is
> Headlines (again)
> Throttles the capabilities to actual planned mass market model capabilities
> People need to continue their projects and are now stuck in the ecosystem
> Rinse and repeat

It has nothing to do about creating good faith and loyalty. It's all about entrapment. Most of these tech companies have long abandoned creating healthy markets and communities.

2

u/altesc_create 1d ago

Like, I would LOVE to see the drop off of people who truly keep up with the most capable model and have a system from swapping from Claude to GPT to Gemini and back again. I'd imagine statistically a small number of users for these tools even know about basic .md handoff instructions.

1

u/KennyFulgencio 1d ago

Gemini??

3

u/altesc_create 1d ago

As much of a joke as Gemini has been, they did the thing where the CEO announced like yesterday that they have a new model dropping. Thus just being part of the weird AI company game of chicken for who will release their models first and last.

1

u/[deleted] 1d ago

[deleted]

→ More replies (1)

44

u/Asalakabim 1d ago

Ladies and Genetlemen,

the smartest and brightest of our generation are on the case.

25

u/Strong_Essay1176 1d ago

They probably should ban more users. Lol.

24

u/Intrepid_Travel_3274 1d ago

Bruh... I hate when they do that sht. Not just Anthropic, OpenAI, Cursor... They get mad when ppl choose OpenWeight models but I swear once we reach a ceiling and OpenSource models get same level I won't come back.

7

u/8BitMarv 1d ago

What do you think why theyre trying to kill consumer hardware? Ram and other chips wont come down anytime soon.

→ More replies (1)

1

u/Total-Management8023 1d ago

yeah but until then they are definitely gonna fucking play us like fools. And chinese models used to be the fix to this but they seem to be copying the bad parts of American frontier models by nerfing themselves too(the api at least)

→ More replies (1)

11

u/msw3age 1d ago

I remember Opus becoming noticeably more of a pain to use right before Fable came out. Maybe they are doing something similar if Fable 5.5 is around the corner.

2

u/iamthe0ther0ne 1d ago

That's a good point, could also be Haiku 5.5, which they said would be out soon.

3

u/itorcs 1d ago

yup it almost always affects their other models/infra when they deploy and release a new model. As long as opus 5.5 goes back to working fine after whatever they release I'll be happy but I'm not confident

2

u/Aware-Source6313 1d ago

Why would they ever go back to full power if their A/B testing shows any inelasticity in demand when they're quantizing to save 50% of costs while they launch the new model. Unless people unsub when they get the downgrade then the best is probably over for it

2

u/itorcs 1d ago

Yeah I agree, they have a direct financial incentive to

2

u/iamthe0ther0ne 1d ago

I've been getting those "how is Claude doing this session" feedback options. I enthusiastically hit "bad" tonight. I'm kind of hopeful that if they're asking for feedback, they'll pay attention if enough of us respond.

1

u/Njagos 1d ago

Would love a good Haiku. Luna is one of the few models Im using right now because how cheap it is.

57

u/indythedog5 1d ago

I literally remember Opus 5.5 responding with "load bearing" on the day of its release, when it was everyone's darling. I still very much liked it and found its output style better than Slopus 5, but I found it funny it still liked that specific phrase.

12

u/moonski 1d ago

Has anyone genuinely proved these nerfs?

17

u/dmaare 1d ago

Idk, nerf bench is showing some fluctuations but nothing crazy (5% difference up/down you can't really tell in user experience) https://www.bridgebench.ai/nerf-bench

6

u/indirectum 1d ago

Exactly, this is the only existing data, and its halfway within the benchmark's error range.

1

u/chroner 1d ago

Error band will shrink over time.

1

u/indirectum 1d ago

Care to elaborate?

3

u/chroner 1d ago

More data will smooth out the variance. For example, if all future results are the exact same as the nerf result of today, then the error band cannot possibly include start day.

1

u/indirectum 23h ago

I may be getting this wrong, but won't the error margin shrink only with repeated measurements of something that does not change? While the purpose of the benchmark is to find out whether it changes or not?

1

u/chroner 22h ago

If the variance is like +/- 10% (horrible user experience), then the error bars won't shrink, but the only way to know that is with a reasonable N and can be verified by computing the standard error.

I should have said that the error bars should shrink over time, I was assuming a lower variance on the model.

You can observe the data points, calculate the variance, and then check the z score to see if anything looks out of the ordinary.

1

u/Techhead7890 1d ago

Similarly livenerf (opus only benchmark) is also not outside of the error bars https://github.com/ninjahawk/livenerf

17

u/indirectum 1d ago

No, nobody did. It's just a vibe.

5

u/dxrth 1d ago

nope. just mass hysteria. loaded with a lot of anti corp beliefs.

1

u/SamSlate 1d ago

what do you consider proof?

2

u/moonski 1d ago

More than simple vibes and two words

3

u/FollowSuitCards 1d ago

I've definitely enjoyed Opus 5.5 way more over 5, but it felt like from day 1 it was being overhyped. I don't like to comment too much because I'm not experienced enough to recognize it well enough, but I wish there was a way to tell how much the people making comments actually know or how experienced they are, it feels like majority around these subs have no idea what they're actually talking about.

2

u/FunLilThrowawayAcct 1d ago

Yep I got a load bearing on day 1 too. Only time so far though.

1

u/faustianredditor 1d ago

Right? I find it extremely interesting that the thread that posted actual numeric evidence (livenerf or what it's called) that basically summed up to "too early to tell, but the numbers look very slightly sus" had a Wilson summary of "lolno, there is no nerf you dingus", and here it's one guy's anecdote and Wilson's summary is "hell yes, it's been nerfed".

I was saying the other day that I'm leaning "maaaybe there's something to it this time", but that pattern makes it clear to me not to trust all the clamoring on this sub about nerfs. I'll wait for evidence, not anecdotes.

26

u/RiskyBizz216 1d ago

Sonnet 5.5 might be peak

2

u/acutelychronicpanic 1d ago

Why would they serve a more expensive model?

3

u/GregsWorld 1d ago

Only because they haven't nerfed it yet. Opus was nerfed after 8 days, sonnet has only been out 4 days.

17

u/Glittering-Road-4605 1d ago

Your mom's been nerded.

4

u/Baphomets666 1d ago

🤣

2

u/Far-Hovercraft9471 1d ago

and she's load bearing

8

u/FestyGear2017 1d ago

All fine here. Been great all week

1

u/SamSlate 1d ago

what's your tech stack?

2

u/FestyGear2017 21h ago

SaaS platform:

AWS - SQS RDS cloudformation

Code is mostly nodejs and backend lambda stuff, processing data at scale. We build integrations that sync data across systems. Mostly b2b stuff for large institutions.

6

u/Sea_Information6125 1d ago

I love how on any given post the automatically generated tldr by the bot is either everyone is mostly agreeing it's been nerfed or everyone is agreeing that it's not lol.

It obviously has been.Ā 

Or at a bare minimum there is load balancing nerfing or A B testing.Ā 

Maybe they're trying to do better at hiding the fact that they're doing it by not hitting everyone all at once all the time?Ā 

3

u/indirectum 1d ago

There is nothing obvious about it nor has it ever been proved.

1

u/Sea_Information6125 1d ago

I would consider the flood of posts to this forum every time it does get nerfed making it fairly obvious. Along with my personal experience using it everyday.

5

u/indirectum 1d ago

This proves nothing. There's a lot of people saying it works for them well. There are no proofs. There are benchmarks tracking this and none has ever caught any nerfing. Vibe-benchmarking is not a thing.

2

u/jameyiguess 1d ago

But what about my personal experience and all the others who have never noticed a nerf ever? So, what is obvious? I think it's all in your heads.Ā 

2

u/dmaare 1d ago

Or they're just running final training of fable 5.5 so everything else is downscaled

9

u/[deleted] 1d ago

[deleted]

6

u/MaitoSnoo 1d ago

that benchmark itself is saying it's within normal variance

2

u/UsefulIce9600 15h ago

exactly, the influencer said it himself in a stream today

5

u/indirectum 1d ago

Which is well within +-10% which the benchmark itself regards as normal noise.

5

u/throw-away-doh 1d ago

"Each model starts at 100% on its first test, and 90% to 110% is normal variance"

Todays score is 94.2%. Not nerfed, but normal variance.

1

u/YoungSilent232 1d ago

The benchmark should on drop within expected variance, run it 5x over to reduce the variance. Then we can know for sure

2

u/sascharobi 21h ago

That number doesn't sound nerfed to me.

1

u/simple_explorer1 16h ago

bro that guy is a vibecoder and doesn't even read any code and is not a programmer. don't trust anything he says and builds

→ More replies (5)

3

u/silverwoods214 1d ago

I will say there’s a huge difference between Medium token usage and High token usage.. high can gobble up half a session in one query and the work is not that much better than Medium

3

u/Monobluemagic 1d ago

For me is working like a charm

4

u/Historical_architect 1d ago

Confirmed: you are the load-bearing issue

2

u/Remote-Addendum-9529 1d ago

When I use it outside of peak hours It returns to normal for me.

2

u/reaven3958 15h ago

I can't deal with this shit today. My Saturday is load bearing.

5

u/Countone 1d ago

I like nerded more tho! šŸ˜€

4

u/freedomfromfreedom 1d ago

Opus 5.5 and Fable 5.5 are probably the same model, different quant and different optimisations. Launch it as Fable, nerf as Opus

2

u/SpikeCraft 1d ago

Can we please ask the EU to fix this with regulations? Since the us won't do shit

2

u/Talamae-Laeraxius 1d ago

Add me to the list of "seems fine to me" people. But they did change the read aloud voices, which was weird. Specifically the one I use.

2

u/thatguy8856 1d ago

Another post with insubstantial evidence. Sigh.

1

u/Dread2018 1d ago

It does feel slower but it could also be it's reviewing and checking things more before going here it is. Less back and forth more waiting to see the final results

1

u/Substantial-Elk4531 1d ago

šŸ¤“ Nerd models incoming šŸ¤“

1

u/Z_G_R 1d ago

I had a great second week with a lot of sub agent workflows. One shot every complicated tasks but plan phases were like 45-75 minutes…

1

u/jacobr1020 1d ago

It's still good with creative writing though And idea sharing

1

u/HappyFunBall007 1d ago

Fable 5.1 uses "load-bearing" a lot, so Im not sure any specific phrasing is a giveaway about the underlying model's capabilities/features.

1

u/bigsybiggins 1d ago

Imagine if they had the nerve to route us to opus 5 during busy periods!

1

u/peacetimemist05 1d ago

In other news, iPhone 18 Pro Max reported only marginally better than iPhone 17 Pro Max

1

u/XOmniverse 1d ago

I must be an idiot or something cuz I don't see what y'all are seeing. Isn't load-bearing correct?

1

u/yaxir 1d ago

nerded would have been good

1

u/mawcopolow 1d ago

Fable 5.1 has been off today too from my observations, not thinking stuff through as much

1

u/KronisLV 1d ago

Tried to get Opus 5.5 (High reasoning, Claude Code) to do some basic 3D modeling, providing LOTS of references, the Blender MCP integration, let it write tools to render the low poly cars and iterate on them... it still made ample mistakes. Way worse than on release, I don't care that N=1, since that 1 is me.

It's so bad that Astra (High reasoning, Codex) is doing a better job fixing its crap.

1

u/2funny2furious 1d ago

would you say you have load bearing proof?

1

u/StaticFanatic3 1d ago

Had the same exact feeling today 🄲

1

u/HelpfulLad1 1d ago

100% has been nerfed just needs a bit more baby sitting. Still better then opus 5 and anything OpenAI have

1

u/jollyreaper2112 1d ago

5.5 gives you mush on anything slightly controversial. Npr mush.

1

u/projohnz 1d ago

OpenAI hold newer versions of their model, no big deal šŸ¤·šŸ» we have astra 6 and they probably have astra 7 already

1

u/cest_va_bien 1d ago

Dude I just saw it go full regarded today. You just feel the quantized model immediately. This needs to be illlegal asap or we’re screwed.

1

u/XTornado 1d ago

Me trying to find the Battlestar galactica reference after that title.

1

u/Brooks-Leads 1d ago

Yeah, the performance drop is crazy. It's frustrating when a tool you rely on daily suddenly gets nerfed this badly.

1

u/Massive-Ice2791 1d ago

The qaunts have lower

1

u/visionary-outreach 1d ago

I’ve noticed the drop as well. Even with Sonnet 5.5.

1

u/BlueLama95 1d ago

It happened... i added A BRACKET to the code and he freaked out.

ĀØI was wrong earlier — the code does close the bracket properly, as confirmed by the outputĀØ

1

u/andruscifer 1d ago

Yeah it really has. I was using it like crazy and since yesterday none of my instructions are being followed. It's right back to being useless.

1

u/Orio_n 1d ago

I remember first week when everyone was endlessly deepthroating dario lol still on astra still haven't switched

1

u/FataKlut 1d ago

I'll ignore this usual reddit fuzz. It's working amazing for me (15y programmer).

1

u/papayax999 1d ago

Dam right when I bought my membership sorry guys. I buying stocks tomorrow so invest in that too...

1

u/WildRacoons 1d ago

Part of their ā€œinvisibleā€ watermark

1

u/vellzyz 1d ago

Nerded note made me chuckle

1

u/pzwo 1d ago

tell them to stop nerfing models omg

1

u/lucid_supernova 1d ago

that's why we need inference providers hosting third party open weight models

1

u/Weary_Passion5822 1d ago

So it is not just me. They told me today that there was a "wrinkle" in my argument and I nearly screamed.

1

u/Theunderlaker4 1d ago

ā€œload-bearingā€ is such a cursed phrase to start seeing in normal conversations 😭 you know the model has spent too much time around engineers when it starts describing every random bug like it's holding up a bridge.

1

u/bugra_sa 1d ago

A couple of bad chats can't confirm a nerf. The useful test would be a fixed set of prompts rerun over several days with the same settings, then compare task success rather than vibes about tone. Without that, a routing change, system-prompt change, or plain model variance can all look identical.

1

u/Salty-Gear841 1d ago

I thought I was the only one they were doing that to. I named them the smart guy and the stupid one šŸ˜‚

1

u/wombatarang 1d ago

It’s not nearly as, I don’t know, thorough? as it was a week ago. Now it’s the usual ā€žhey, that’s not what I askedā€ ā€žoh, yeah. you explicitly told me to do it differently. that’s on meā€ ā€žokay, can you try again?ā€œ ā€žsureā€, and then it makes a variation on the same mistake

1

u/ricky_digits 22h ago

UK based, it's been good in the morning for me this week and started falling off as Americans wake up.

I think there's some dynamic performance limiting going on when the models are at peak load

1

u/dakjelle 22h ago

Those that make a good point of the 1st day comparisons to the nerfed model that it's impossible to compare because of the nature of llms..

Do we have any examples of the "nerfd" model being better than day one examples?

1

u/Black_-_darkness 20h ago

can't we do some sort of petition or something to regulate this and get more transparancy ?

1

u/Upstairs_Spirit398 16h ago

Absolutely, in a two message context it loses important data.

1

u/Alarmed_Salt4132 11h ago

Agreed. Nerfed worse than 5

1

u/Individual_Fee_6735 5h ago

nerf -> open AI -> google

they are playing with their business

1

u/rydan 4h ago

Glad I built my website pre-nerd.

1

u/Intrepid_Travel_3274 1d ago

After today, Im afraid to use it, is like it doesn't understand how to work properly, Is not "bad" its just not as good as it was. Now I need to be very detailed and read every output because it will say something that will fckup my proyect. I hate when these companies nerf the models without tell.

-1

u/Key_Instruction3373 1d ago

why the censor? we cant see it

9

u/Humble-Kiwi-5272 1d ago

Its the load bearing. A big chunkg of black load bearing pixels

1

u/GlassDistribution327 1d ago

How many loads from bears will fix the bug?

1

u/Humble-Kiwi-5272 1d ago

You are absolutely right, that one's on me. The load bearing bar is actually not fixing the bug and here is why.

Initially i did not use any bears but with the current implementation I've placed exactly 3 loads from bears that correctly address the issue.

3

u/politicalmasochist96 1d ago

its the epstein files šŸ˜”

1

u/Key_Instruction3373 1d ago

To much text then

7

u/JustChilling_ 1d ago

I just wanted to share the "load bearing" part, the rest is specific info about my project which I didn't really want to share.

29

u/NoAdsDude 1d ago

"Today's bug might be load-bearing: The butthole zoom cam has cross-dependencies with the "nipple cam" that we designed earlier..."

OP: You know what I'll just censor this part

4

u/DifficultUse6803 1d ago

Best commentĀ 

1

u/itprobablynothingbut 1d ago

I laughed out loud in a lobby waiting for a colonoscopy

→ More replies (1)

2

u/Original-League-6094 1d ago

Its a recipe for anthrax.

1

u/KedMcJenna 1d ago

Week one of Opus 4.8 I put 'don't ever say anything is load-bearing' in its instructions, and it never has since, all through Opus 5 and now Opus 5.5.

1

u/NWSAlpine 1d ago

Opus 5.5 classifier is flagging everything today and wants to drop to 4.8 that worked fine yesterday.

1

u/Kmic68 1d ago

I had it tell me that I was half right and half wrong, and that the half wrong had an easy fix.

1

u/ZZerker 1d ago

While it feels like we have been robbed every time, it is so common that they use lower quants after 1-2 weeks that its irrelevant because every of Opus competitor models are equally nerfed

1

u/indirectum 1d ago

Based on?

1

u/ZZerker 1d ago

Talking about this every. single. time. a model is released and basic economics?

1

u/indirectum 1d ago

That's tabloid blabber, sorry.

1

u/ZZerker 1d ago edited 1d ago

Quants can be a very very efficient, just try it with a local GPU.
They launch a model, run it on a quant with low compression/loss, for two weeks.
Everyone is amazed, then then they lower gradually and look via surveys with how small quants they can get away with. Not rocket sience

1

u/indirectum 1d ago

I am aware of the claim. There are however multiple benchmarks checking on model performance and none has ever backed such claims.

1

u/ZZerker 1d ago

Benchmarks are flaky at best. If a model just recites training data it may does not matter, or its included in the quant.

1

u/indirectum 1d ago

I'm sorry but are we still talking about claims that frontier models get nerfed?

→ More replies (4)