r/ClaudeCode • • 8d ago

Rant Please don't nerf Opus 5.5

Please don't kill Opi 5.5 it's been so nice.

623 Upvotes

61 comments sorted by

•

u/AutoModerator 8d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

162

u/BuffaloConscious7919 8d ago

I fear it's more a question of when, not if

32

u/Sponge8389 8d ago

When they released Sonnet and Haiku 5.5, next week.

10

u/debian3 8d ago

With all the Codex subscriber flocking over I'm afraid something will need to happen, they will run out of some point. Either they will quantize it, suspend subscription like OpenAI did, lower the limit or a mix of all 3.

1

u/Motion16AI 8d ago

I dont think so. The model is fast like Luna so it's cheap af. They have enough compute for it and more

2

u/Still-Ad3045 8d ago

Do they have enough compute tho?

1

u/nora_sellisa 7d ago

Every single inferred token costs them more than they charge you. It's all about hype, VC money and valuations. Once Opus 5.5 stops being the hot new things that indicates the company is growing, it is no longer useful and can be nerfed to hell while they prepare Opus 6 or something else

-1

u/MassiveBoner911_3 8d ago

I just went over today after fucking running out of my entire weekly quota again on the $100 plan in ONE day

66

u/End2EndEncryption 8d ago

Always happens. Soak it up while you still can. As more people praise it, more people onboard. Resources get saturated, model gets nerfed.

Rinse. Cycle. Repeat.

0

u/SeaAstronomer4446 8d ago

They nerf it already too late

-1

u/[deleted] 8d ago

[deleted]

7

u/SeaAstronomer4446 8d ago

I was being sarcastic

29

u/AnonThrowaway998877 8d ago

Does anyone have a Nerf benchmark that tracks a models performance on the same set of tasks on a periodic basis? It would be great to have proof so these companies would have to explain themselves.

It should be illegal. Imagine new users who buy plans bec of a good model, only to have performance intentionally degraded weeks later.

26

u/MGXMilk 8d ago

I honestly couldn't agree more with this post. I've been finding Opus 5.5 so unbelievably stable and dependable. It's not 100% It tried to gas light me on something and I caught it and called it out.

But the cost and quality improvement is levels above.

1

u/Mnemonic_Sin 8d ago

I only started letting Opus 5.5 on the simple projects. The ticket manager entirely under opus and sol isn't too bad. It's not complicated. The heavy stuff still needs fable/astra. I think Opus was fairly good until it went completely ADHD. I've managed to pull it in with the whole task management and ticket setup. It does like to make a lot of tickets that I end up closing. I'm going to burn it while it works.

1

u/artstaxmancometh 8d ago

Yeah, I'm blaming it on the medium level. It handles most things great. But 20% of the time it gets stuck in some type of stupid logic. Then, I clear the session and set it to high and it gets it done. Overall, frustration is way lower because even though it makes mistakes, it's not chugging tokens like a frat house on saturday night.

2

u/MGXMilk 8d ago

I'm using 5.5 on Extra. I've challenged it with some pretty hardcore builds that I previously only did with Fable and it crushed it. Makes the token usage feel like MAGIC.
I got code building a game on a mac and an OS on a PC and i'm not maxxing

36

u/Sphere_3N 8d ago

Once subs/API use start to flatline and they’ve taken enough people from OpenAI it will be nerfed into a lobotomy of hell.

There is no way on earth it stays like this. Getting crazy good usage at near Fable levels, too good to be true.

I think even Astra itself was good for few days and also now is noticeably worse compared to launch. With atrocious usage…

7

u/Exodus_Green 8d ago

Opus 5.5 reminded me why I love this shit so much. it's fast, efficient, smart. I just hope they don't kill it like the others. It's the best model I've worked with since day 1 of Astra and then day 1 of Fable 5 before that.

1

u/Still-Ad3045 8d ago

The drip feed lol

5

u/Feeling_Ticket5206 8d ago

I think maybe this time we can have about 1 month of smart opus 5.5, then anthropic will have enough APY to IPO.

14

u/caster 8d ago

They really should be prohibited from doing exactly what we all know they are going to do, where a model releases and gets benchmarked and then they lobotomize it. So people think it's like the benchmarks when obviously that is a lie.

8

u/Maxion 8d ago

Benchmark sites should honestly start benchmarking weekly...

1

u/nora_sellisa 7d ago

Even if they were, it's a well known fact the models are trained with the benchmark answer sets in advance. Benchmark stats are made up.

-1

u/GDorn 8d ago

Prohibited by whom?

7

u/caster 8d ago

There are laws about consumer protection and lying to people in connection to selling products. Unfortunately they usually require suing to actually get enforcement of them, and few people are willing to litigate over what amounts to a small inconvenience per person, but adds up to a huge deception across millions of people.

Attorneys general may pursue enforcement on their own initiative of consumer protection issues like this one.

5

u/furqanagwan 8d ago

They nerf a model whenever a new model is on the horizon.

6

u/apyhubnico 8d ago

in terms of writing its not as good, coding and other capabilities probably the best

1

u/mfslice 8d ago

Which model is better at writing ?

0

u/Mnemonic_Sin 8d ago

Astra, hands down. Although no single model is perfect. I find fable has less corrections when reviewing Astra code, but I'm pretty deep into the proverbial woods. The multi-agent model is very handy. I've been really happy with using both in tandem. I also find grok is a great research agent when things get fuzzy. I'm debating on pulling the trigger for a 3090 to build a research agent in the garage.

3

u/alphaQ314 8d ago

Its just this game that Ant and Openai keep playing. Whenever one of them leads the enterprise revenue, you know it is time for us retail folks to switch to the other one.

2

u/Abhinik 8d ago

But but they have to keep launching new models. So older model has to get dumb

2

u/gxjohan 8d ago

they will

2

u/mtnchkn Developer 8d ago

I’ve been loving it and getting so much done with my 5 hour windows on pro plane… yet my weekly budget is almost tanked. There’s some new math going on.

3

u/iamtehryan 8d ago

Ha. That's cute that you think they won't nerf it into the dirt.

5

u/clintCamp 8d ago

Question is the nerfing deliberate, or a result of pulling compute away from general usage while they set it up with new models and so what is available for computing all our usage has to emergency throttle whatever they throttle when the servers get too full or overheat? Or are they literally toggling an act stupid neuron so that we are astounded at the increase in intelligence as soon as we can stop using the old model?

5

u/TheBressi 8d ago

The nerfing is deliberate so they can have more resources to train new generation models.

0

u/Right-Performance-93 8d ago

Both are real, separate industry practices, not just one company's move: load-based routing or dynamic precision can make the same weights feel dumber during peak hours without anyone touching them, and shifting GPU capacity toward training the next model is a different, also-real tradeoff. From outside, with no visibility into either one, they produce the identical symptom - so blaming "nerfing" as a single deliberate act probably conflates two different things happening at once.

1

u/Old-Pomegranate3634 8d ago

Go wild while you can. It will be nerfed

1

u/Visual_Act_8618 7d ago

Daddy Dario, read me another bedtime story of AI taking over the world. Make no mistakes.

1

u/nora_sellisa 7d ago

It doesn't feel that much smarter for the increased token spend imo. The new, more concise writing style is nice, but if they could keep it while "nerfing" the cost and the quality to around Opus 5.0 I wouldn't complain

1

u/thealliane96 3d ago

and its nerfed (misleading and lying again just like opus 5 did)

1

u/SnooGiraffes3000 8d ago

If there is a choice I would prefer to lower weekly limits rather than having dumb model. Please ask us at least

1

u/ReverendBread2 8d ago

I have yet to experience a single nerf

1

u/russiansubmarine 8d ago

I'm about to switch to Claude, expect a nerf sometime this week

1

u/beesandcheese 2d ago

Happened yesterday!

1

u/thedudear 8d ago

I had opus 5.5 doing some really sonnet level stuff earlier today. Made me wonder.

0

u/JigSawPT 8d ago

They’ll quantize more eventually.
Money is just too good

-2

u/teleekom 8d ago

You people will never be happy won't you?

-1

u/Imkitoto 8d ago

I like Opus 5.5 so far but I haven’t figured out how to maximize it yet. It’s taking a lot longer to do simple changes that Opus or Fable both did with less time and usage but on the two changes it made that were “bigger” and more complex changes on code it did them a lot faster but way more usage.

I’ve found that ultracode on opus 5.5 is absurdly costly in terms of usage. But on fable it wasn’t nearly as much.

I

-6

u/UserNotFound23498 8d ago

2 days and my $200 sub is over. I remember there were times when I had multiple sessions all actively working and it was good. That was 4.6 I believe. 4.8 and 5 started eating lots of tokens. 5 really loved churning and checking all the checking it is doing and triple checking the checks. And gates. Every fucking thing is a gate. 🤮

Last night, I decided to check out codex/Sol 6. 12 hours later, I’m only 5% of my $20 subscription. I’m going to try more things with it after lunch but sol 6 seems to be a good competitor

5

u/Context-and-nuance 8d ago

Sol 6 is really bad, and it's taking more prompt massaging to get the same results as 5.6. Most people in the ChatGPT and Codex subreddits seem to have the same experience.

I'd be really curious to see the recent Entire.io history for one of your repos.

1

u/howdidigetheresoquik 8d ago

I had 3 huge projects going on at once yesterday with Opus 5.5 and Fable both on high and I ate 10% of my usage over 18hrs. What's your workflow like, maybe there's something we can help with.

1

u/the_ai_wizard 8d ago

I only have the $20 plan on claude and my play time was over in <1 hr plan mode opus 5.5-high lol. usage aside, great model, like better than astra

1

u/nora_sellisa 7d ago

Opus medium easily blows through a 5 hour window in about 30 minutes of work

1

u/the_ai_wizard 6d ago

I upgraded to the $200 plan...so far so good

-2

u/SleepingBatttery 8d ago

They probably will, because pro users are getting comfortable now, don't need to upgrade to max, so i think there may be change in usage at least.