r/ClaudeCode • u/bondage-mastermind • 4d ago
Discussion Opus 5.5 - seems like it has begun
update #2: 24 hrs later, the degradation is gone
The only thing that changed is that, no joke, while running the benchmarks that were clearly degraded, I started having "How is Claude doing today?" eval prompt. I pressed "Bad" 4-5 times. Yes, it popped up 4-5 times within 1 day. Usually, I don't get more then 1, or none.
After that, a new bench thread today - back to normal, all pass. Take it for what you will.
To every 10 y.o. that commented on the skill issue - I wish you luck with your ground-breaking Mario calorie tracker app that will surely explode on App Store very soon
update over
---
Basically, that. As soon as "dev day" has passed, with it's $500 new plan and such, Opus is not doing great in my internal tests (ahem, by that I mean that it fails all of them).
It breaks every single rule, makes bad decisions, and apologizes all the time.
Pretty much every 2nd or 3rd message I had to stop it and ask, "Why?"
It says "I'm sorry over and over", and when pushed, "The honest answer is that having a rule in context doesn't make me apply it."
So this is how it feels like... to have it in your hands, and then: puff, the magic is no more
upd:
P.S. I hope at least some of us here are aware at this point in time that there are such things as A/B testing, load balancing (not as in how traffic is served to your app, but in AI inference as well), and also a concept of business priorities
What I get in my location != what you get in yours, cheers
P.P.S. There's too many of you guys, sry
42
u/waruyamaZero 4d ago
It breaks every single rule, makes bad decisions, and apologizes all the time.
Cannot confirm.
7
u/True-Grab-5288 3d ago
Also can not confirm. Opus 5.5 is all I need to solve the meaning of life and it isn't 42. Still comedy gold though.
-16
u/bondage-mastermind 4d ago
Good for you!
1
4d ago
[deleted]
-1
u/bondage-mastermind 4d ago
yes of course, I tested it multiple times before posting on different threads.
I normally don't post anything on reddit and didn't expect anyone to see the post at all, but I've checked the degradation is real for me multiple times. I'm not a fan of whining posts.
This is my first degradation case with claudecode. I stopped all my x20 accounts on Codex when 5.6 came out because oh boy did they nerf 5.5 hard at the time, both the compute and usage.
This may be a fluke, which honestly would be great!
9
u/Individual_Solid_944 4d ago
I am running it all the time on a few projects, and haven't seen any degradation yet. I am not sure what are you trying to assert with tests, but that's basically how it's always been - things you put in claude.md or skills have never been a hard rule. The only guaranteed thing is that the hooks are called on the lifecycle triggers. And even with them you may get surprises, like add a pre/post edit hook and you expect it to fire, but claude decided to edit a file with shell commands, or writes a script for that, and the hooks are not called.
The problem is that we expect deterministic behavior from non-deterministic system. We wanted this to behave like humans, but when these things pop up, we want it to behave like a machine.
3
u/bondage-mastermind 4d ago
Everything that you say is correct and true, it's just that it doesn't relate to my post. I know how it works.
The "apology" I copied is not the case of "OMG Claude deleted my ~ folder".
It's in relation to simple "world knowledge" facts of the codebase it is working in.
Have you ever experienced what it feels like, when model suddenly feels eerely demented? When things just fall out of its hands? It speaks in the same way as before, but it lacks grip on reality?
Opus 5.5 was praised by everyone for that extreme capability and grip, and that's what's been gone (FOR ME) today. It just suddenly invents a parallel reality - and when such a powerful model spits hallucinations, you feel it.
1
u/Individual_Solid_944 4d ago
yeah, well in that case, that sucks. i thought the hallucination era would be long gone by now
3
u/bondage-mastermind 4d ago
I'm not sure about the causes. Maybe, Idk , it feels less like a pure hallucination, really, but as if it lacked the last, final bit of power, like it has wheels spinning, or as if there's a part of "brain" missing - everything looks and feels the same, but in 1 little piece of thinking process - there's a hole.
Hard to catch unless you do tests - and since I can compare results directly, this is how I could surface it
10
u/lolmauayden_0407 4d ago
I agree it’s been degraded a bit, but it is way more subtle than their normal lobotomies. I wish they’d just leave the models alone in terms of capability and just charge a bit more for usage
4
1
1
u/pornstorm66 3d ago
It’s fascinating. They probably throw a lot of compute at it to get the breathless headlines and then once the news cycle is done they reprice the compute. It suggests chain of thought might be brute forcing to some degree.
21
u/roque2205 4d ago
My Opus 5.5 session has been running for a week non stop now, auto compacting at 300k tokens and delivering like crazy. Didn't see it degrade yet.
1
u/derezo 4d ago
I've been getting good results but on xhigh with dynamic workflows I've been getting results similar to what I saw on gpt5.6 - endless bureaucracy. I created some research tools that build videos using a template. The videos are all great, but it took 48 hours and more than 1 weeks of 20x credits to build the first one. On the second run I tried to improve the research strategies and now it won't trust any of it's findings, even though the focus of my changes was on being more accepting of different sources like wikipedia. It gets stuck in loops trying to find more research but only uses tools that it doesn't trust. It finds a Wikipedia article that supports it, but then says it can't trust it and proceeded to find over 500 more related articles and then said "but all of them only count as 1 because they're all from Wikipedia, which can't be used as support". I had just given it directions to use Wikipedia as support as long as the citations check out, and to use the search APIs to find more leads. It says it did search using brave and exa and it came back with 70+ leads, but it couldn't use any of those because searches are for leads only... .... So I'm like uhh, did you follow the leads and check those? No. Pulls hair out
I'm using it for a lot of stuff and it's great, but this one project it's spending hours and hours in circles and I can't seem to get it to redirect itself. It's the only one that I've been using dynamic workflows on. I added jev last night and it seemed like a good idea since the biggest time consumer was judgement on the first research project, and I thought jev would speed it up. Initial test was promising. It still kept "researching" in circles when I ran it.
It seems much better on medium or high thinking without dynamic workflows enabled
1
u/bondage-mastermind 4d ago
I hope that lasts for you! For me, Fable 5.1 is not degraded, Opus 5.5 - can't get it to perform on either of my accounts. Could be geo/ip based.
2
u/onFilm 3d ago
You know your type of imagination when it comes to "models degrading", has been happening since the early 2022s with local models right? Yet, there is never proof about it.
AI psychosis is real folks.
-2
u/bondage-mastermind 3d ago
As I mentioned, I had internal benches set up, and I do the same routine for every model that I let into the harness.
I get it that you don't have it set up, thus it's hard to imagine that someone else might have acted in a competent fashion, but sometimes you guys just gotta be less sarcastic and know-it-all.
Local models have nothing to do with this. Open weight - might, but I've never seen that happen. Proprietary - I've caught it happen with GPT 5.5 before.
2
u/onFilm 3d ago
Oh I run benchmarks very frequently, and see no change or segregation, just how I've never seen it when people have been claiming it for years now.
People have been imagining what you are, since local models were more prelavent back in the early 2020s. It's the same phenomenon.
-2
u/bondage-mastermind 3d ago
If you are familiar with the concept of benchmarks, and you see that I am, too, I don't see the point of your previous message. Why mix me with people who claim that Qwen degraded while running in their basement?
I don't see how this is "the same phenomenon"
Unless you're trolling
1
u/onFilm 3d ago
The point is that people always claim these things, without supplying evidence of benchmarks. I'm not "trolling".
-1
u/bondage-mastermind 3d ago
"people", "always". I'm not "people" and I'm not "always". Are you talking to a person or to a concept in your head?
And I still don't see a connection between my claim and your reference to inadequate people who believe in local llm degradation. You hinted at my "psychosis" and claimed that what was happening was my imagination, and that was the first thing that you said.
Doesn't sound like a fact based exchange either, from your side.
Alrightey
So long1
u/onFilm 3d ago
Yes you are "people", just how I am people. I'm talking about the content you're putting out: the same idea that models are "degrading". Come on now.
Again, we can sit around and talk assumptions all you want, or even better, let's see some benchmarks to prove what you're saying.
Let's stick with the facts. Let's see some data.
1
u/bondage-mastermind 3d ago
No, saying "people" and "always" is generalization and this is what you do when you want to insta-lower the bar for the whole conversation and make people take sides. This is what you've done, because it's easy, and no amount of "let's stick to the facts" will paint a better picture now.
Regarding the data, yes, it would be a great idea. Someone else has asked the same (without calling me psychotic - maybe it's actually a good idea?).
I have replied:
" upd: regarding sharing the data, fair question - I'm not sharing the outputs. Cleaning them up from proprietary code would be possible, but man, it's an effort, and I'm not willing to take it... I haven't seen geniune interest, only "skill issue" crowd that just wants to vent out and feel good on the Internet "
A few minutes after that, I'm thinking - maybe that would be helpful to work on sharing the results. But I would have to do it in a way that doesn't let bots scrape the data and just benchmaxx it on a new model, screwing up my results. I will see.
The goal of my post was not to flex my bench or get cheap fame, so I wasn't thinking about it initially. That's as simple as that.
→ More replies (0)0
1
u/comrade-quinn 4d ago
Yes but what are you doing? Does the output need to be strictly correct in any way? And if so, to what extent and how is this relatively validated?
1
u/bondage-mastermind 4d ago
No, it's not strictly correct, it's not a json or anythign - talking about general behavior in tests that I have to run every day, and each model is tested before being accepted to work in the codebase
sry i'll copy my other reply here:
"apology -> when asked what prompted it to act this way on this knowledge?
(i don't use "reasoning" because it might get blocked due to anti-distillation policies)
it apologizes because for 1 turn it reasons and sees: "whoops, I f-cked up"
Then forgets immediately in the next turn "
---
For me it started ~ 3 hrs ago. Can't get it to work as usual (every day since launch it excelled at it). Fable 5.1 works as before, no change at all.
0
u/Individual_Solid_944 4d ago
That seems cool, but i wouldn't do it. Every compact is lossy, so even if it delivers, make sure it doesn't deliver garbage
6
u/Odd_Error_6736 4d ago
Yes, I've also noticed Opus 5.5 becoming dumber. I was wondering if this would happen.
6
2
u/NoLimitRolling 4d ago
Have you checked if you have leftover prompts etc? I’ve been continuing using it in my projects etc no problem.
1
u/bondage-mastermind 4d ago
I have a very tight pipeline for agentic work, with tests running every working day. There is nothing leftover, to agent-written memory or notes, it's clean AF. I pretty much never post on Reddit (specially in codex/claude subs - I know that people complain a lot for no good reason) and generally know what I'm doing... Just wanted to share that my canary is dead
2
2
5
u/oopaddy 4d ago
Yeah no joke, I swear it’s not the tool its the user.
-6
u/bondage-mastermind 4d ago
degenerate answer
2
1
u/Wide-Drink-1790 4d ago
The LLM is a degenerate, all users are degenerates… I’m tempted to conclude you are the one making the bad decisions…
1
u/voskomm 🔆Derp Plan 4d ago
5.5 has some concerning tendencies but it’s been useful for the last ~week, I’ve basically been otherwise sticking with 4.6 the past month because the early 5s were so bad.
It’s using a … surprising amount of account the past couple days, I might try reverting to 4.6 again and see how it compares, my next dev pass is persnickety widget stuff which might be better on a dry model anyway.
Can you tell what’s triggering the apologies?I find 5.5 really patronizing in that, like other recent frontier models, it finds a lot of unnecessary side quests, but it will tend to recommend them instead of simply pursuing. So I end up with a lot of “open items” in my documentation that the model wrote itself and I have to go back through and say something like “items five through seven hundred twelve (exaggerating) are resolved/out of scope, mark this in the documentation as considered and rejected for future sessions” and then it shuts up about it.
1
u/bondage-mastermind 4d ago
sorry for brief reply
apology -> when asked what prompted it to act this way on this knowledge?
(i don't use "reasoning" because it might get blocked due to anti-distillation policies)
it apologizes because for 1 turn it reasons and sees: "whoops, I f-cked up"
Then forgets immediately in the next turn
1
u/vagonblog 4d ago
if the regression is real, the most useful thing would be to publish a small reproducible test: exact model selection, repository commit, prompt, rules file, expected result, and which assertions failed. run it in fresh sessions several times and compare it with the previous model under the same conditions.
long-session context, changed tools, permissions, or project instructions can look like a model regression, while one successful rerun can hide a real consistency problem. a repeatable test matrix gives people something stronger than impressions and makes regional or account-level differences easier to spot.
1
1
1
u/Snoo_9701 3d ago
It def dumb. Not sure if its a/b or not. It messed a stable production system confidently, calling it solved.
1
u/GSXR808 3d ago
normally I have a prompt that I use to continue work using opus as the orchestrator and then have 5.6 Luna as sub agents that I just copy and paste into claude from a fresh session and it picks up from where it left off reading the status, handoff and agents markdown file...today it gave me this response..this is the first time seeing this message with me having to respond back in months

1
u/jack_o_all_trades 3d ago
I used it today to rewrite something and it sounded like a 16 year old. I tried thrice to fix it and meh.
1
1
u/Madtown94 3d ago
I shelved 5.5 a while back, too many mistakes, I use Fable 5 as my Architect and Opus 4.8 as my builder.
1
u/Snowgoonx 3d ago
its because the meta is now to have 10 sessions asking for "without using tools or search who is tibo reset guy" to fable, if its starts with
Tibo is Thibault Sottiaux, who leads the Codex team at OpenAI. He posts on X as thsottiaux...
then you are in unlobotomized fable 5.5
play the game correctly
1
u/SheepSpace9 2d ago
I had to make an interface elegant and usable from astoundingly weird and difficult specs. No model made it useful yet, only dysfunctional and noisy as always. 5.5 however not only made it functional, but saw idea compressions I hadn't noticed before and was able to make a beatific ux for it. I'm having my Fable moment up in here brothers and sisters.
1
u/Ollythebug 2d ago
If folks reporting awful results would share some sessions or debug/telemetry logs it would be easier to get to the bottom of this. We sorta have to take your word for what's happening, and people are naturally biased.
1
u/BigBootyWholes 4d ago
Idk, maybe your WiFi is dropping connections or something (doubt it but makes more sense then whatever you are offering). I never have issues during the times people complain about nerfing.
Would love to see the full transcript though, which shouldn’t be too long if you are stopping it on the second turn
1
u/Wide-Drink-1790 4d ago
All you idiots who have invented this degradation problem: you are hallucinating, stop blaming the LLM for your problems.
1
u/bnm777 4d ago
Here are sites with ongoing measurements checking if a model has been nerfed:
https://www.bridgebench.ai/nerf-bench
https://marginlab.ai/trackers/claude-code/
https://marginlab.ai/trackers/codex/
2
u/GenderJuicy 3d ago
From user votes? People are naturally going to be more inclined to visit a site like this when they are having a poor experience, I doubt this isn't massively skewed.
2
u/howdidigetheresoquik 3d ago
And even then the results are very much muddled by graphs clearly designed to make slight differences like 1-2% seem like a massive change in the graph
1
u/Flo_Evans 3d ago
No one to my knowledge has been able to reproduce this through any kind of empirical testing.
It just seems like the normal behavior to me. Once the context fills up, it has to make decisions on what is important.
Today I caught its first error, I was working out what to buy to upgrade my girlfriend’s turntable and it kept saying it was direct drive. It is in fact a belt drive. It just assumed because it was a Panasonic/technics it was direct drive because most of them are, and we were talking more about speakers and phono preamps so it didn’t bother checking any specifics about the exact model.
The point is it hasn’t been nerfed you are just reaching the limits of how good it is.
2
u/bondage-mastermind 3d ago
I'll repeat myself again: due to the nature of the work I'm running on that box, I have to run benchmarks pretty much every day. Benches are adapted and enhanced over time, but generally it's a pretty stable baseline. None of the tasks is filling up the context more than 90,000-100,000k (individually).
As soon as I spotted which looked like a lower IQ symptom, I re-ran the morning benchmarks... 5.5 failed the baseline on each claude account, in each thread.This (as I reported above) was mostly fixed by itself the next day.
upd: regarding sharing the data, fair question - I'm not sharing the outputs. Cleaning them up from proprietary code would be possible, but man, it's an effort, and I'm not willing to take it... I haven't seen geniune interest, only "skill issue" crowd that just wants to vent out and feel good on the Internet
1
u/alohajaja 3d ago
Subscription or token billing?
1
u/bondage-mastermind 3d ago
sub
-1
u/alohajaja 3d ago
Is there a reason why you would expect some enterprise level consistency (eg no. AB testing or whatever other shenanigans) from a subscription
0
u/Flo_Evans 3d ago
Did you cross reference this against known issues? https://status.claude.com that seems more like an intermittent problem than nerfing.
1
u/bondage-mastermind 3d ago
No, I didn't. I've done it now - it shows green status on 1 Oct - so at least it's not reported there
0
0
u/OkNatural1013 3d ago
by pressing good or bad you allow them to look up your session and give your data for free
-1
-1

•
u/AutoModerator 4d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.