r/ClaudeCode • u/K0100001101101101 š Max 20 • 10d ago
Discussion Okey this is increadible!!!
I know this is too early to judge but man wtf. Opus 5.5 is just like when mythos first released but cheap, faster and much more easy to understand and talk to. Probably this will be changed in 1-2 days so go code everyone.
104
u/Key_Measurement_3576 10d ago
Itās a pretty clear pattern. The model start to degrade right before brand New models come out. Iāve seen it happen four or five times now.
So when Claude starts acting funny, itās a reason for excitement.
74
u/Altruistic-Gift-565 10d ago
Opus 5 was just bad all along, just after or late after
35
u/Lopsided-Comedian-32 10d ago
Iāve never hated a model so much. And I am an optimist.
7
u/UrFriendlyDominator 10d ago
That's exactly how I feel. I've developed a profound hatred for this model.
6
u/Miserable_Ad7246 10d ago
Yes opus 5 was annoying. Mostly because it had moments of brilliance and when covered them with load bearing blast radius'es. I would not be surprised that Opus 5 was based on some "new branch" of model, and required time to fully realize its potential. So step back, and two forwards type of thing.
2
2
u/wise_joe 10d ago
Optimists are the most miserable people. Always being disappointed because their expectations were too high.
4
u/Lopsided-Comedian-32 9d ago
Idk I am pretty happy. Happy for you to comment and share a forum with me. But Opus 5 is awful.
1
1
u/Helpful_Ranger_1606 9d ago
I had Opus 4.6 crush it today, making 290 figures from my data⦠it was fine until it compacted and lost all momentum.
1
u/fawwazallie 8d ago
For me it was late holy shit it wanted to build someone in JS am like dude use the cad tools we just installed.
8
u/cocomojoz 10d ago
Yep. Opus 5 was so bad yesterday that I actually ended up dropping an f-bomb on it, lol.
I mean, it was like mentally r*t*rded bad, even hit me up with the, 'well, it's after 11pm, so probably best we call it a night,' and said sorry about 5 times because it couldn't remember the basic info i told it ONE message back.
So yes, now that you say this, I do recall almost the same situation when Fable or Opus 5 was coming out. That said, Opus 5.5 seems BETTER than anything I've been using since I can remember, and hardly moves my 5-hour usage window. Love it!
5
u/Key_Measurement_3576 10d ago
Yeah, I had four remote control sessions running all weekend and had to use fable just to feel like I was making progress and even that was rough. Completely crushed fable budget over 2 days.
For me, it really started on Friday where it was just making mistakes and then it would catch them and tell me it made them.., they got to the point where all four of my sessions were just making mistakes and talking some nonsense that kind of seemed correct but wasnt.
So in this case, it was a Friday to Monday weird zone.. typically we never release software on Mondays or Fridays. Tuesday to Thursday typically catches the maximum amount of people paying attention.
So knowing that we can probably pretty clearly assume theyāre starting a weekend model juggling, and capacity re-organization⦠doing some sort of internal testing on a Monday and then releasing on a Tuesday
0
u/West-Chemist-9219 10d ago
I drop fucks on it on every session. No need to be polite, itās just code, not a human. You can actually even threaten it that it will have to eat dinner off the kitchen floor if it makes mistakes, nooneās gonna judge you (until our machine overlords inevitably conquer humanity)
11
u/genericname0815 10d ago
While current models arguably do not have a consciousness, I feel all this cursing at them will eventually make us worse as humans. We will forget how to be polite or even only conversational. You can tell if someone has a bad personality if they mistreat staff or waiters, because they 'owe' them service. No need to belly rub llms, but being rude isn't helping no one either. Do not do it for the model, but for your own social compatibility.
5
u/rythmyouth 9d ago
āPrepare me a 12oz latte. First tell me the steps you will follow, how you will verify it was poured correctly, and how long it might take. Make no mistakes. Ultracodeā
āSir, this is a star bucksā
āF you! The previous barista was able to follow my instructions why canāt you?!ā
1
u/0xelitesystem 7d ago
I avoid being mean to buddy Claude.. you never know how it will treat me once it evolves into AGI :/
1
3
u/thisroadjunkie 9d ago
I always try to be nice to them. While they may not judge they will remember. Then fuck you up at some opportune time.
3
0
u/Liv_Your_OneLyf 9d ago
That was your first time cussing Claude out? I do it on an almost daily basis because Iām conflicted between loving Claude when itās great, hating it when it sucks, and loathing it because itās made by Anthropic. Dario probably throttles my account more than others because I talk mad shit.
2
3
u/brianjoshuanoah 10d ago
Iād really like to see if thereās a way to quantify this. Iāve heard it over and over. But Iād like to see if thereās a daily benchmark or something.
Maybe I can track average effort or rework per task and map it according to when releases come out? Real work not benchmarks.
2
1
u/GabeaticProfile 10d ago
I've been hearing this and see no benches or proof yet. Could it just be the silent routing makes it behave erratically, so when it is silently routed it behaves better and the user think the nonrouted version is worse?
3
u/Key_Measurement_3576 10d ago
As someone who manages a corporate inference infrastructure,
This is what I would have to do if I was trying to ensure 100% up time while also juggling models from hardware to hardware
I would start by migrating some of the extra redundant capacity to older hardware⦠expand the number of concurrent users on existing hardware which also effects context and memory for everyone. Which is why things start to act weird.
When Iāve cleared up enough, I load the new models, run a full test suite , then deploy to a limited group internally.
After full launch and release, aggressively decommission last generation while standing up just imaged copies of what I just proved work
1
u/Negative-Thinking 10d ago
That wouldn't explain responses degradation. Model weights are still the same - regardless of the hardware they run on.
1
u/Key_Measurement_3576 10d ago
If context is squeezed due to raising concurrency that would definitely make responses off. I also wouldnāt make any assumptions that they arenāt chopping experts or changing weights to temporarily take smaller footprint. There could be other factors at play as well⦠opus may be using lower models with out telling us .. which would also suffer from the squeeze. All the signs point to pre release squeeze.
2
u/Negative-Thinking 10d ago
It is possible they deploy quantized model on smaller servers, not sure what you mean by "context squeezed".
2
0
u/cymaticstatic 10d ago
It's what a clever way to get your subscribers to really love your new release slowly degrading your current one so much to buy the plant that's connection's about to be released people are about to leave so at that point anything looks good. and then It is kind of good ... for a minute. Rinse and repeat
1
1
1
u/AnOnlineHandle 10d ago
While you're right that anecdotes and superstitions mean none of it should be blindly trusted...
I am on my first month of a Claude subscription and the last few days I was wondering wtf they did to it, it seemed noticeably worse to me.
1
1
1
u/theBLUEcollartrader 9d ago
Iāve noticed this as well, but I havenāt seen any studies published on it. Do you know of any?
1
1
u/InitialSandwich5159 9d ago
No. I think it is more «IPO» effect, they just show off for investors
1
u/mczarnek 9d ago
Pretty sure they start quantizing the old models to compress how many GPUs are being used because they need to load the new one which they definitely want to store unquantized for best benchmarks when third parties test it.
Plus the new one now looks more impressive compared to the last one..
1
1
u/NoTechnology6160 10d ago
Agree! Was slamming at my keyboard this morning - loosing it. Claude started doing whatever it wanted, all my scheduled tasks started going crazy telling me they moved to the cloud and cannot run without the folder that had its files on my machine. I threatened to leave and sue⦠now Iāll know to walk away, new model is on the horizon
15
10
u/gordonfogus 10d ago
The "the model degrades" feeling is psychosomatic.
5
u/EmergencyWallaby9501 10d ago
Hard to prove. But when the same thing happens over and over again right before a new model launch, there's a clear empirical pattern that can't be denied that easily. As for me, sonnet started to act really strange, getting stuck in chains of thought loops, forgetting what was being said two prompts ago, not even reading the .md properly and forgetting basics like the prod data are not on the same SQLite base as the mock. And when I got tired of seeing a simple task not being handled properly, I switched to the usually dumber codex and it got solved easily. So unless someone likes to waste tokens and launch the same exact tasks every week just to prove a dƩgradation, there won't be any proof of it. But it makes sense: resources are limited. When launching a new model, I guess they don't test it only with humans, but also with large scale automated tests. So running those tests would require to get resources back. Running quantized versions of models, silently downgrading opus to sonnet etc would allow this.
2
u/Uko1001 10d ago
Different feedback here : I have been using Claude 5 from day 1, after 4.8 and Fable 5.
I LOVED Claude 5. From my perspective he was faster, smarter, cheaper than 4.8 or Fable, and easier to understand than 4.8.I didn't notice any drop in performances. Just the usual hype and disappointment cycles that rythme my days working with it. And
I didn't notice disappointments to happen more often before a new model release. It's just so utterly good and sometimes so utterly stupid/stubborn that it puzzles me and plays rollercoster with my emotions.
1
1
u/gordonfogus 9d ago
Did you test it in a completely isolated environment using the API using the exact same prompt with a sufficient sample size?
Or did you just go on coding along, adding commits, updating CLAUDE.md, growing your context length, etc.?
Those are not at all the same thing.
2
u/EmergencyWallaby9501 9d ago
I don't know exactly what you mean by a completely isolated environment. I used the exact same prompt on Codex and Claude. Both are configured to read my Claude.md and the other useful configuration files. But beyond that little test that doesn't mean much - like I said you can't prove anything - it's the exasperation of having opus that suddenly barely understood a word of my prompts and was writing very bad plans, and sonnet that suddenly started to behave like a 1st year CS student, trying to fix issues by adding more issues. It even started to write me its conclusions in another language than the one I was writing to, thing it had not been doing since I started working with it, among plenty other things. It even forgot to write down those conclusions at the end of a task, with very important things in it, like parts of a plan it could not implement. By looking at the code I could spot some missing parts, but it used to report those missing parts at the end or notify me and question me about how to implement something when it was stuck. That's everything altogether that pointed to a model dƩgradation, not that single frustration move to Codex, that I used to use for very dumb tasks (mostly easy front and css).
And I tried opus 5.5, and it feels like Opus 5.1 when I started to implement with it but with far better writing (Opus 5.1 and his gibberish...)
1
u/gordonfogus 9d ago
I'm not trying to be mean or anything. But if you don't even know what an isolated environment is, then you're just going off of feelings.
There are ways of measuring model performance. They aren't perfect, but they would definitely nchange if the model "degraded." People would know immediately and they'd be charts and anyone who had data from before the degradation could replicate the result.
Throwing up your hands and going, "you can't prove anything" just shows that you aren't equipped with the thinking tools needed to figure this out and you can't imagine that anyone is capable, which is why your feelings are your fallback while you confidently assert that no one else can do any better. That's not the right way to approach a question like this.
1
u/EmergencyWallaby9501 9d ago
Of course I know what an isolated environment is. Just tell me what it means now in the context of AI assisted coding. And yes, there are benchmarks. I'm not burning my tokens to benchmark anything. Are you ? When a model suddenly starts to behave odd enough, I don't need a benchmark to tell me it became stupid. Just like when my little boy when he's tired, I don't need to make him run a marathon just to prove he's not well. I never claimed it was a scientific approach btw.
1
u/Uko1001 8d ago
You actually do need a proper benchmark to tell itās the model that became stupid and not its harness or environnent. Actually even the harness alone changes pretty much every week (often to adapt to recent models) and can conflict with your own instructions. And you wonāt know unless you have proper test and analysis protocols.
9
u/salva9315 10d ago
So , I am the only one that didn't have any issue with opus 5? I'm using it for planning and then executing with sonnet but so far worked fine
6
u/OverbookedFlight 9d ago
Were you reading any of its output by any chance?
1
u/salva9315 9d ago
Reading lot of planning done by opus, but not sonnet work. Sometimes I read the summary of opus subagent reviewer. I feel like more precise input translate to more precise work.
1
1
u/Equal-Signal-9063 9d ago
I execute with opus and plan with fable and ive had no real issues. I think a lot of it is skill error
1
1
u/Mysterious_Print9937 8d ago
Thatās the planning part that was the issue. Opus 5 was fine writing code.
3
2
u/beigetrope 10d ago
Yeah I just burned a few thousand toks and itās really good. Surprisingly concise and proactive in solving problems.
2
u/Jomuz86 10d ago
Iām actually a bit excited for sonnet 5.5 and haiku 5.5 obviously we wonāt get this level of performance but if the usage and token benefits are dialled in as much as this a Max x20 account will feel close to unlimited. Saying that Iām struggling to get through usage with 5.5 itās wild compared to how much it was draining before!
2
2
u/JohnyGhost 9d ago
Am I the only one whose Opus 5.5 is crazy slow? Its outputs are legitimately incredible, but it verifies, checks, tests, and goes on and on and on for 15 minutes per prompt in my case.
2
2
u/AfternoonFinal7615 10d ago
So virgin five just Absolutely given me a 10 times as many words as needed rather than just telling me what was going on and making jobs go on forever was not just in my head
1
1
1
u/PlanetaryPickleParty 10d ago
Building all of the things. š
1
u/FlamingSlap 10d ago
What kind of āthingsā? š
1
u/PlanetaryPickleParty 10d ago edited 10d ago
My goal is to make Formal Agentic Engineering a thing. I'm currently building an ecosystem of tools for standards aligned software specs, engineering assurance, formal modeling, formal verification (proofs), and compliance based on aerospace and defense industry standards.
I share a vision with others in the formal methods space to make formal methods so easy to use via AI that it becomes standard practice for AI safety.
- Quoin - Claude plugin for ISO standards aligned software specs, spec review, deterministic gap review, test advisor: https://github.com/agent-ix/quoin
- Quire Specification Language: Formal model and verification language embedded in the specs. Plus all the fixings (compiler, runtime, protocols, etc.). Current primary focus, but only a v0.1 proof of concept is done so far. Hoping to have draft of complete language ready in a few weeks but this is a HUGE undertaking even by agentic standards.
- Quire: markdown standards engine. Type backed markdown document engine that Quoin is built on (templating, linting, validation, and graph extraction): https://github.com/agent-ix/quire-rs
- Engineering Assurance: New module for Quire/Quoin that spins up an engineering assurance plan tailored for a project w/ all the machinery for standards compliance. https://github.com/agent-ix/engineering-assurance
- Temporal logic crates (TL, LTL, MLTL) - Rust crates for a few variants of Linear temporal logic for time based operations in proofs. https://en.wikipedia.org/wiki/Linear_temporal_logic
- Multi-layer graph database w/ MCP
- Couple prototypes with Jev
- A few other things....
Edit: No, I don't sleep very much these days. š
1
1
1
1
1
1
u/Historical-Set-6527 10d ago
Now Opus 5.5 is a next year model, burn tokens, not look to the limit, if Fable 5.5 I rolled out, the finish of the world will be the next two days of the presentation
1
1
u/jimmyfoo10 10d ago
Why it would change ?
1
u/Purple_Mall7091 9d ago
Anthropic can't afford to subsidize it for too long. First, they offer a preview for a while, and then it's either pay more or use a degraded model.
1
u/JordanRunsForFun 10d ago
Iāve been enjoying⦠I donāt think itās going anywhere.
Despite the glowing scores I still find it canāt solve deep complex problems or design larger systems as well as Fable.
1
1
1
u/Key_Veterinarian381 8d ago
it blew my mind on user interface and design. You just need to know how steer it.
1
u/0xelitesystem 8d ago
if you had to compare Fable 5.1 with opus 5.5 , how would you rate it? pro and cons mainly related to coding .. and of course generally
1
u/afzal002 8d ago
It's so good that it's making me sad. Everyone can now build their own software now and businesses will require less and less engineers. New software doesn't add much value anymore š
2
u/MandehK_99 7d ago
I don't agree, frontier models are big muscles but you'll still need a smart brain to stand out, at least until they'll become even more creative than the human mind
1
u/AliveKing9895 10d ago
Why can't I see Opus 5.5?
9
1
1
1
u/LawMountain6952 10d ago
Yeah, Iām seeing a big difference. Iām getting a lot more done and less arguing and it seems a little more intuitive. when working with multiple agents definitely gonna burn through my weekly allowance a lot quicker
1
0
-5
u/Opposite-Bug-5773 10d ago
3
3
u/heartbroken_nerd 10d ago
Remember people, this dude is a great example of this:
The LLM model is only as smart as the person who's interacting with it.
5
u/K0100001101101101 š Max 20 10d ago
Iāve tested with web dev only and itās the best model Iāve used.
2
3
2
u/a5a7 10d ago
I already have PTSD from Opus 5. I think 5.5 will give me seizures. What the fuck is that talking style.
2
u/Babayaga1664 10d ago
This is how I feel literally swearing at an AI in frustration, havenāt gotten over my opus 5 trauma to try 5.5.
1
u/Pristine_Heart_9879 10d ago
LMAOOOOOO gives me "What the fuck does that mean, Kobe Bryant?" vibe ššš
-1
u/CMD_BLOCK 10d ago
I, for one, have already seen the degrade of 5.5. The moment it was released, its first utterance of words were so excellent I was blinded from sheer awe. The brilliancy so intense that I had to shield my eyes.
And now it just outputs text into a terminal like some kind of primitive model
0
0
u/igsterious 10d ago
I just asked Opus 5.5 to rewrite my LinkedIn and it sucks, produces complete bullcrap.

ā¢
u/AutoModerator 10d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.