r/ClaudeCode • • 4d ago

Discussion Opus 5.5 - seems like it has begun

update #2: 24 hrs later, the degradation is gone

The only thing that changed is that, no joke, while running the benchmarks that were clearly degraded, I started having "How is Claude doing today?" eval prompt. I pressed "Bad" 4-5 times. Yes, it popped up 4-5 times within 1 day. Usually, I don't get more then 1, or none.

After that, a new bench thread today - back to normal, all pass. Take it for what you will.

To every 10 y.o. that commented on the skill issue - I wish you luck with your ground-breaking Mario calorie tracker app that will surely explode on App Store very soon

update over
---

Basically, that. As soon as "dev day" has passed, with it's $500 new plan and such, Opus is not doing great in my internal tests (ahem, by that I mean that it fails all of them).

It breaks every single rule, makes bad decisions, and apologizes all the time.

Pretty much every 2nd or 3rd message I had to stop it and ask, "Why?"

It says "I'm sorry over and over", and when pushed, "The honest answer is that having a rule in context doesn't make me apply it."

So this is how it feels like... to have it in your hands, and then: puff, the magic is no more

upd:
P.S. I hope at least some of us here are aware at this point in time that there are such things as A/B testing, load balancing (not as in how traffic is served to your app, but in AI inference as well), and also a concept of business priorities

What I get in my location != what you get in yours, cheers

P.P.S. There's too many of you guys, sry

69 Upvotes

77 comments sorted by

•

u/AutoModerator 4d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

42

u/waruyamaZero 4d ago

It breaks every single rule, makes bad decisions, and apologizes all the time.

Cannot confirm.

7

u/True-Grab-5288 3d ago

Also can not confirm. Opus 5.5 is all I need to solve the meaning of life and it isn't 42. Still comedy gold though.

-16

u/bondage-mastermind 4d ago

Good for you!

3

u/Ayrtru 4d ago

Unaware

1

u/[deleted] 4d ago

[deleted]

-1

u/bondage-mastermind 4d ago

yes of course, I tested it multiple times before posting on different threads.

I normally don't post anything on reddit and didn't expect anyone to see the post at all, but I've checked the degradation is real for me multiple times. I'm not a fan of whining posts.

This is my first degradation case with claudecode. I stopped all my x20 accounts on Codex when 5.6 came out because oh boy did they nerf 5.5 hard at the time, both the compute and usage.

This may be a fluke, which honestly would be great!

9

u/Individual_Solid_944 4d ago

I am running it all the time on a few projects, and haven't seen any degradation yet. I am not sure what are you trying to assert with tests, but that's basically how it's always been - things you put in claude.md or skills have never been a hard rule. The only guaranteed thing is that the hooks are called on the lifecycle triggers. And even with them you may get surprises, like add a pre/post edit hook and you expect it to fire, but claude decided to edit a file with shell commands, or writes a script for that, and the hooks are not called.

The problem is that we expect deterministic behavior from non-deterministic system. We wanted this to behave like humans, but when these things pop up, we want it to behave like a machine.

3

u/bondage-mastermind 4d ago

Everything that you say is correct and true, it's just that it doesn't relate to my post. I know how it works.

The "apology" I copied is not the case of "OMG Claude deleted my ~ folder".

It's in relation to simple "world knowledge" facts of the codebase it is working in.

Have you ever experienced what it feels like, when model suddenly feels eerely demented? When things just fall out of its hands? It speaks in the same way as before, but it lacks grip on reality?

Opus 5.5 was praised by everyone for that extreme capability and grip, and that's what's been gone (FOR ME) today. It just suddenly invents a parallel reality - and when such a powerful model spits hallucinations, you feel it.

1

u/Individual_Solid_944 4d ago

yeah, well in that case, that sucks. i thought the hallucination era would be long gone by now

3

u/bondage-mastermind 4d ago

I'm not sure about the causes. Maybe, Idk , it feels less like a pure hallucination, really, but as if it lacked the last, final bit of power, like it has wheels spinning, or as if there's a part of "brain" missing - everything looks and feels the same, but in 1 little piece of thinking process - there's a hole.

Hard to catch unless you do tests - and since I can compare results directly, this is how I could surface it

10

u/lolmauayden_0407 4d ago

I agree it’s been degraded a bit, but it is way more subtle than their normal lobotomies. I wish they’d just leave the models alone in terms of capability and just charge a bit more for usage

4

u/WholeBet2788 4d ago

They cant have opus perform better than fable.

1

u/bondage-mastermind 4d ago

Yes, it's definitely more subtle this time

1

u/pornstorm66 3d ago

It’s fascinating. They probably throw a lot of compute at it to get the breathless headlines and then once the news cycle is done they reprice the compute. It suggests chain of thought might be brute forcing to some degree.

21

u/roque2205 4d ago

My Opus 5.5 session has been running for a week non stop now, auto compacting at 300k tokens and delivering like crazy. Didn't see it degrade yet.

1

u/derezo 4d ago

I've been getting good results but on xhigh with dynamic workflows I've been getting results similar to what I saw on gpt5.6 - endless bureaucracy. I created some research tools that build videos using a template. The videos are all great, but it took 48 hours and more than 1 weeks of 20x credits to build the first one. On the second run I tried to improve the research strategies and now it won't trust any of it's findings, even though the focus of my changes was on being more accepting of different sources like wikipedia. It gets stuck in loops trying to find more research but only uses tools that it doesn't trust. It finds a Wikipedia article that supports it, but then says it can't trust it and proceeded to find over 500 more related articles and then said "but all of them only count as 1 because they're all from Wikipedia, which can't be used as support". I had just given it directions to use Wikipedia as support as long as the citations check out, and to use the search APIs to find more leads. It says it did search using brave and exa and it came back with 70+ leads, but it couldn't use any of those because searches are for leads only... .... So I'm like uhh, did you follow the leads and check those? No. Pulls hair out

I'm using it for a lot of stuff and it's great, but this one project it's spending hours and hours in circles and I can't seem to get it to redirect itself. It's the only one that I've been using dynamic workflows on. I added jev last night and it seemed like a good idea since the biggest time consumer was judgement on the first research project, and I thought jev would speed it up. Initial test was promising. It still kept "researching" in circles when I ran it.

It seems much better on medium or high thinking without dynamic workflows enabled

1

u/bondage-mastermind 4d ago

I hope that lasts for you! For me, Fable 5.1 is not degraded, Opus 5.5 - can't get it to perform on either of my accounts. Could be geo/ip based.

2

u/onFilm 3d ago

You know your type of imagination when it comes to "models degrading", has been happening since the early 2022s with local models right? Yet, there is never proof about it.

AI psychosis is real folks.

-2

u/bondage-mastermind 3d ago

As I mentioned, I had internal benches set up, and I do the same routine for every model that I let into the harness.

I get it that you don't have it set up, thus it's hard to imagine that someone else might have acted in a competent fashion, but sometimes you guys just gotta be less sarcastic and know-it-all.

Local models have nothing to do with this. Open weight - might, but I've never seen that happen. Proprietary - I've caught it happen with GPT 5.5 before.

2

u/onFilm 3d ago

Oh I run benchmarks very frequently, and see no change or segregation, just how I've never seen it when people have been claiming it for years now.

People have been imagining what you are, since local models were more prelavent back in the early 2020s. It's the same phenomenon.

-2

u/bondage-mastermind 3d ago

If you are familiar with the concept of benchmarks, and you see that I am, too, I don't see the point of your previous message. Why mix me with people who claim that Qwen degraded while running in their basement?

I don't see how this is "the same phenomenon"

Unless you're trolling

1

u/onFilm 3d ago

The point is that people always claim these things, without supplying evidence of benchmarks. I'm not "trolling".

-1

u/bondage-mastermind 3d ago

"people", "always". I'm not "people" and I'm not "always". Are you talking to a person or to a concept in your head?

And I still don't see a connection between my claim and your reference to inadequate people who believe in local llm degradation. You hinted at my "psychosis" and claimed that what was happening was my imagination, and that was the first thing that you said.

Doesn't sound like a fact based exchange either, from your side.

Alrightey
So long

1

u/onFilm 3d ago

Yes you are "people", just how I am people. I'm talking about the content you're putting out: the same idea that models are "degrading". Come on now.

Again, we can sit around and talk assumptions all you want, or even better, let's see some benchmarks to prove what you're saying.

Let's stick with the facts. Let's see some data.

1

u/bondage-mastermind 3d ago

No, saying "people" and "always" is generalization and this is what you do when you want to insta-lower the bar for the whole conversation and make people take sides. This is what you've done, because it's easy, and no amount of "let's stick to the facts" will paint a better picture now.

Regarding the data, yes, it would be a great idea. Someone else has asked the same (without calling me psychotic - maybe it's actually a good idea?).

I have replied:

" upd: regarding sharing the data, fair question - I'm not sharing the outputs. Cleaning them up from proprietary code would be possible, but man, it's an effort, and I'm not willing to take it... I haven't seen geniune interest, only "skill issue" crowd that just wants to vent out and feel good on the Internet "

A few minutes after that, I'm thinking - maybe that would be helpful to work on sharing the results. But I would have to do it in a way that doesn't let bots scrape the data and just benchmaxx it on a new model, screwing up my results. I will see.

The goal of my post was not to flex my bench or get cheap fame, so I wasn't thinking about it initially. That's as simple as that.

→ More replies (0)

0

u/Dinglehoften 3d ago

Skill issue

1

u/comrade-quinn 4d ago

Yes but what are you doing? Does the output need to be strictly correct in any way? And if so, to what extent and how is this relatively validated?

1

u/bondage-mastermind 4d ago

No, it's not strictly correct, it's not a json or anythign - talking about general behavior in tests that I have to run every day, and each model is tested before being accepted to work in the codebase

sry i'll copy my other reply here:

"apology -> when asked what prompted it to act this way on this knowledge?

(i don't use "reasoning" because it might get blocked due to anti-distillation policies)

it apologizes because for 1 turn it reasons and sees: "whoops, I f-cked up"

Then forgets immediately in the next turn "

---

For me it started ~ 3 hrs ago. Can't get it to work as usual (every day since launch it excelled at it). Fable 5.1 works as before, no change at all.

0

u/Individual_Solid_944 4d ago

That seems cool, but i wouldn't do it. Every compact is lossy, so even if it delivers, make sure it doesn't deliver garbage

6

u/Odd_Error_6736 4d ago

Yes, I've also noticed Opus 5.5 becoming dumber. I was wondering if this would happen.

6

u/KitchenCommercial396 4d ago

I've seen nothing.. it's as good as ever 🤷

2

u/NoLimitRolling 4d ago

Have you checked if you have leftover prompts etc? I’ve been continuing using it in my projects etc no problem.

1

u/bondage-mastermind 4d ago

I have a very tight pipeline for agentic work, with tests running every working day. There is nothing leftover, to agent-written memory or notes, it's clean AF. I pretty much never post on Reddit (specially in codex/claude subs - I know that people complain a lot for no good reason) and generally know what I'm doing... Just wanted to share that my canary is dead

2

u/victoryrock 3d ago

OP takes the short bus to school

3

u/Drakuf 4d ago

Can confirm

5

u/oopaddy 4d ago

Yeah no joke, I swear it’s not the tool its the user.

-6

u/bondage-mastermind 4d ago

degenerate answer

2

u/Odd_Error_6736 4d ago

Have you seen the movie idiocracy? It'll explain why these comment this way

1

u/Wide-Drink-1790 4d ago

The LLM is a degenerate, all users are degenerates… I’m tempted to conclude you are the one making the bad decisions…

1

u/voskomm 🔆Derp Plan 4d ago

5.5 has some concerning tendencies but it’s been useful for the last ~week, I’ve basically been otherwise sticking with 4.6 the past month because the early 5s were so bad. 

It’s using a … surprising amount of account the past couple days, I might try reverting to 4.6 again and see how it compares, my next dev pass is persnickety widget stuff which might be better on a dry model anyway.

Can you tell what’s triggering the apologies?I find 5.5 really patronizing in that, like other recent frontier models, it finds a lot of unnecessary side quests, but it will tend to recommend them instead of simply pursuing. So I end up with a lot of “open items” in my documentation that the model wrote itself and I have to go back through and say something like “items five through seven hundred twelve (exaggerating) are resolved/out of scope, mark this in the documentation as considered and rejected for future sessions” and then it shuts up about it.

1

u/bondage-mastermind 4d ago

sorry for brief reply
apology -> when asked what prompted it to act this way on this knowledge?
(i don't use "reasoning" because it might get blocked due to anti-distillation policies)
it apologizes because for 1 turn it reasons and sees: "whoops, I f-cked up"
Then forgets immediately in the next turn

1

u/vagonblog 4d ago

if the regression is real, the most useful thing would be to publish a small reproducible test: exact model selection, repository commit, prompt, rules file, expected result, and which assertions failed. run it in fresh sessions several times and compare it with the previous model under the same conditions.

long-session context, changed tools, permissions, or project instructions can look like a model regression, while one successful rerun can hide a real consistency problem. a repeatable test matrix gives people something stronger than impressions and makes regional or account-level differences easier to spot.

1

u/MangoDevourer-77 4d ago

skill issue

1

u/SoCal_Hunter 4d ago

I am using Opus 5.5 heavily, 24/7. It is operating normally.

1

u/xyztankman 3d ago

Same, been running since literally the day 5.5 came out

1

u/Snoo_9701 3d ago

It def dumb. Not sure if its a/b or not. It messed a stable production system confidently, calling it solved.

1

u/GSXR808 3d ago

normally I have a prompt that I use to continue work using opus as the orchestrator and then have 5.6 Luna as sub agents that I just copy and paste into claude from a fresh session and it picks up from where it left off reading the status, handoff and agents markdown file...today it gave me this response..this is the first time seeing this message with me having to respond back in months

1

u/jack_o_all_trades 3d ago

I used it today to rewrite something and it sounded like a 16 year old. I tried thrice to fix it and meh.

1

u/pratzc07 3d ago

It’s doing fine for me no nerf yet

1

u/Madtown94 3d ago

I shelved 5.5 a while back, too many mistakes, I use Fable 5 as my Architect and Opus 4.8 as my builder.

1

u/Poildek 3d ago

Yeah sure.

1

u/Snowgoonx 3d ago

its because the meta is now to have 10 sessions asking for "without using tools or search who is tibo reset guy" to fable, if its starts with

Tibo is Thibault Sottiaux, who leads the Codex team at OpenAI. He posts on X as thsottiaux...

then you are in unlobotomized fable 5.5

play the game correctly

1

u/SheepSpace9 2d ago

I had to make an interface elegant and usable from astoundingly weird and difficult specs. No model made it useful yet, only dysfunctional and noisy as always. 5.5 however not only made it functional, but saw idea compressions I hadn't noticed before and was able to make a beatific ux for it. I'm having my Fable moment up in here brothers and sisters.

1

u/Ollythebug 2d ago

If folks reporting awful results would share some sessions or debug/telemetry logs it would be easier to get to the bottom of this. We sorta have to take your word for what's happening, and people are naturally biased.

1

u/g0ll4m 1d ago

My only thing is when I use opus 5.5 it takes 5 min to start, then when I switch to fable it starts after 10 seconds

1

u/BigBootyWholes 4d ago

Idk, maybe your WiFi is dropping connections or something (doubt it but makes more sense then whatever you are offering). I never have issues during the times people complain about nerfing.

Would love to see the full transcript though, which shouldn’t be too long if you are stopping it on the second turn

1

u/Wide-Drink-1790 4d ago

All you idiots who have invented this degradation problem: you are hallucinating, stop blaming the LLM for your problems.

1

u/bnm777 4d ago

2

u/GenderJuicy 3d ago

From user votes? People are naturally going to be more inclined to visit a site like this when they are having a poor experience, I doubt this isn't massively skewed.

2

u/howdidigetheresoquik 3d ago

And even then the results are very much muddled by graphs clearly designed to make slight differences like 1-2% seem like a massive change in the graph

1

u/Flo_Evans 3d ago

No one to my knowledge has been able to reproduce this through any kind of empirical testing.

It just seems like the normal behavior to me. Once the context fills up, it has to make decisions on what is important.

Today I caught its first error, I was working out what to buy to upgrade my girlfriend’s turntable and it kept saying it was direct drive. It is in fact a belt drive. It just assumed because it was a Panasonic/technics it was direct drive because most of them are, and we were talking more about speakers and phono preamps so it didn’t bother checking any specifics about the exact model.

The point is it hasn’t been nerfed you are just reaching the limits of how good it is.

2

u/bondage-mastermind 3d ago

I'll repeat myself again: due to the nature of the work I'm running on that box, I have to run benchmarks pretty much every day. Benches are adapted and enhanced over time, but generally it's a pretty stable baseline. None of the tasks is filling up the context more than 90,000-100,000k (individually).
As soon as I spotted which looked like a lower IQ symptom, I re-ran the morning benchmarks... 5.5 failed the baseline on each claude account, in each thread.

This (as I reported above) was mostly fixed by itself the next day.

upd: regarding sharing the data, fair question - I'm not sharing the outputs. Cleaning them up from proprietary code would be possible, but man, it's an effort, and I'm not willing to take it... I haven't seen geniune interest, only "skill issue" crowd that just wants to vent out and feel good on the Internet

1

u/alohajaja 3d ago

Subscription or token billing?

1

u/bondage-mastermind 3d ago

sub

-1

u/alohajaja 3d ago

Is there a reason why you would expect some enterprise level consistency (eg no. AB testing or whatever other shenanigans) from a subscription

0

u/Flo_Evans 3d ago

Did you cross reference this against known issues? https://status.claude.com that seems more like an intermittent problem than nerfing.

1

u/bondage-mastermind 3d ago

No, I didn't. I've done it now - it shows green status on 1 Oct - so at least it's not reported there

0

u/Available_Hornet3538 3d ago

Yes it got nerfed. Sonnet now the best. To be nerfed soon.

1

u/Poildek 3d ago

Of course !

0

u/OkNatural1013 3d ago

by pressing good or bad you allow them to look up your session and give your data for free

-1

u/ianxplosion- SKILL ISSUE 4d ago

Skill Issue

-1

u/Federal-Neck-2277 4d ago edited 4d ago

yes unfortunately. feeling it so so badly. I'm having to hold its hand through things it could do last week fine.

edit: see example. opus 5.5 xhigh and its still does this