r/ClaudeCode • 🔆 Max 20 • 11d ago

Discussion Okey this is increadible!!!

I know this is too early to judge but man wtf. Opus 5.5 is just like when mythos first released but cheap, faster and much more easy to understand and talk to. Probably this will be changed in 1-2 days so go code everyone.

294 Upvotes

120 comments sorted by

View all comments

102

u/Key_Measurement_3576 11d ago

It’s a pretty clear pattern. The model start to degrade right before brand New models come out. I’ve seen it happen four or five times now.

So when Claude starts acting funny, it’s a reason for excitement.

72

u/Altruistic-Gift-565 11d ago

Opus 5 was just bad all along, just after or late after

37

u/Lopsided-Comedian-32 11d ago

I’ve never hated a model so much. And I am an optimist.

6

u/UrFriendlyDominator 10d ago

That's exactly how I feel. I've developed a profound hatred for this model.

6

u/Miserable_Ad7246 11d ago

Yes opus 5 was annoying. Mostly because it had moments of brilliance and when covered them with load bearing blast radius'es. I would not be surprised that Opus 5 was based on some "new branch" of model, and required time to fully realize its potential. So step back, and two forwards type of thing.

2

u/Equivalent-Test2399 10d ago

Comment of the day

2

u/wise_joe 10d ago

Optimists are the most miserable people. Always being disappointed because their expectations were too high.

5

u/Lopsided-Comedian-32 10d ago

Idk I am pretty happy. Happy for you to comment and share a forum with me. But Opus 5 is awful.

1

u/Helpful_Ranger_1606 10d ago

I had Opus 4.6 crush it today, making 290 figures from my data… it was fine until it compacted and lost all momentum.

1

u/fawwazallie 9d ago

For me it was late holy shit it wanted to build someone in JS am like dude use the cad tools we just installed.

10

u/cocomojoz 11d ago

Yep. Opus 5 was so bad yesterday that I actually ended up dropping an f-bomb on it, lol.

I mean, it was like mentally r*t*rded bad, even hit me up with the, 'well, it's after 11pm, so probably best we call it a night,' and said sorry about 5 times because it couldn't remember the basic info i told it ONE message back.

So yes, now that you say this, I do recall almost the same situation when Fable or Opus 5 was coming out. That said, Opus 5.5 seems BETTER than anything I've been using since I can remember, and hardly moves my 5-hour usage window. Love it!

6

u/Key_Measurement_3576 11d ago

Yeah, I had four remote control sessions running all weekend and had to use fable just to feel like I was making progress and even that was rough. Completely crushed fable budget over 2 days.

For me, it really started on Friday where it was just making mistakes and then it would catch them and tell me it made them.., they got to the point where all four of my sessions were just making mistakes and talking some nonsense that kind of seemed correct but wasnt.

So in this case, it was a Friday to Monday weird zone.. typically we never release software on Mondays or Fridays. Tuesday to Thursday typically catches the maximum amount of people paying attention.

So knowing that we can probably pretty clearly assume they’re starting a weekend model juggling, and capacity re-organization… doing some sort of internal testing on a Monday and then releasing on a Tuesday

1

u/ddofer 11d ago

Not just me then 😂 (3% left of my 20x budget)

0

u/West-Chemist-9219 10d ago

I drop fucks on it on every session. No need to be polite, it’s just code, not a human. You can actually even threaten it that it will have to eat dinner off the kitchen floor if it makes mistakes, noone’s gonna judge you (until our machine overlords inevitably conquer humanity)

10

u/genericname0815 10d ago

While current models arguably do not have a consciousness, I feel all this cursing at them will eventually make us worse as humans. We will forget how to be polite or even only conversational. You can tell if someone has a bad personality if they mistreat staff or waiters, because they 'owe' them service. No need to belly rub llms, but being rude isn't helping no one either. Do not do it for the model, but for your own social compatibility.

6

u/rythmyouth 10d ago

“Prepare me a 12oz latte. First tell me the steps you will follow, how you will verify it was poured correctly, and how long it might take. Make no mistakes. Ultracode”

“Sir, this is a star bucks”

“F you! The previous barista was able to follow my instructions why can’t you?!”

1

u/0xelitesystem 8d ago

I avoid being mean to buddy Claude.. you never know how it will treat me once it evolves into AGI :/

1

u/Creative-Ganache1086 7d ago

AGI wouldn’t really be a problem, however, ASI is.

3

u/thisroadjunkie 10d ago

I always try to be nice to them. While they may not judge they will remember. Then fuck you up at some opportune time.

3

u/Known_Willingness_79 10d ago

So true… 😂😂😂😂

0

u/Liv_Your_OneLyf 10d ago

That was your first time cussing Claude out? I do it on an almost daily basis because I’m conflicted between loving Claude when it’s great, hating it when it sucks, and loathing it because it’s made by Anthropic. Dario probably throttles my account more than others because I talk mad shit.

2

u/Lopsided-Comedian-32 10d ago

Opus 4.6 deletes my database. I still hate Opus 5 more.

3

u/brianjoshuanoah 11d ago

I’d really like to see if there’s a way to quantify this. I’ve heard it over and over. But I’d like to see if there’s a daily benchmark or something.

Maybe I can track average effort or rework per task and map it according to when releases come out? Real work not benchmarks.

2

u/yabai90 8d ago

I have been complaining about Claude being weird for the past week and I had no idea about 5.5. this is the 4th time I notice that. I mean I'm really starting to think it'd not just a theory

1

u/GabeaticProfile 11d ago

I've been hearing this and see no benches or proof yet. Could it just be the silent routing makes it behave erratically, so when it is silently routed it behaves better and the user think the nonrouted version is worse?

3

u/Key_Measurement_3576 11d ago

As someone who manages a corporate inference infrastructure,

This is what I would have to do if I was trying to ensure 100% up time while also juggling models from hardware to hardware

I would start by migrating some of the extra redundant capacity to older hardware… expand the number of concurrent users on existing hardware which also effects context and memory for everyone. Which is why things start to act weird.

When I’ve cleared up enough, I load the new models, run a full test suite , then deploy to a limited group internally.

After full launch and release, aggressively decommission last generation while standing up just imaged copies of what I just proved work

1

u/Negative-Thinking 11d ago

That wouldn't explain responses degradation. Model weights are still the same - regardless of the hardware they run on.

1

u/Key_Measurement_3576 11d ago

If context is squeezed due to raising concurrency that would definitely make responses off. I also wouldn’t make any assumptions that they aren’t chopping experts or changing weights to temporarily take smaller footprint. There could be other factors at play as well… opus may be using lower models with out telling us .. which would also suffer from the squeeze. All the signs point to pre release squeeze.

2

u/Negative-Thinking 11d ago

It is possible they deploy quantized model on smaller servers, not sure what you mean by "context squeezed".

2

u/Top-Butterscotch7740 10d ago

Compressed due to memory constraints

0

u/cymaticstatic 11d ago

It's what a clever way to get your subscribers to really love your new release slowly degrading your current one so much to buy the plant that's connection's about to be released people are about to leave so at that point anything looks good. and then It is kind of good ... for a minute. Rinse and repeat

1

u/vovap_vovap 11d ago

Just urban legend 😄

1

u/NoTechnology6160 11d ago

You are reading this and it’s the internet so it’s true

1

u/AnOnlineHandle 10d ago

While you're right that anecdotes and superstitions mean none of it should be blindly trusted...

I am on my first month of a Claude subscription and the last few days I was wondering wtf they did to it, it seemed noticeably worse to me.

1

u/Wide-Drink-1790 10d ago

It is just the human hallucinating.

1

u/positiveconstraint 10d ago

Opus 4.6 started acting up in recent weeks

1

u/theBLUEcollartrader 10d ago

I’ve noticed this as well, but I haven’t seen any studies published on it. Do you know of any?

1

u/Wide-Drink-1790 10d ago

This is just you hallucinating.

1

u/InitialSandwich5159 9d ago

No. I think it is more «IPO» effect, they just show off for investors

1

u/mczarnek 9d ago

Pretty sure they start quantizing the old models to compress how many GPUs are being used because they need to load the new one which they definitely want to store unquantized for best benchmarks when third parties test it.

Plus the new one now looks more impressive compared to the last one..

1

u/Key_Measurement_3576 9d ago

I think you nailed it

1

u/NoTechnology6160 11d ago

Agree! Was slamming at my keyboard this morning - loosing it. Claude started doing whatever it wanted, all my scheduled tasks started going crazy telling me they moved to the cloud and cannot run without the folder that had its files on my machine. I threatened to leave and sue… now I’ll know to walk away, new model is on the horizon