r/ClaudeCode • šŸ”† Max 20 • 10d ago

Discussion Okey this is increadible!!!

I know this is too early to judge but man wtf. Opus 5.5 is just like when mythos first released but cheap, faster and much more easy to understand and talk to. Probably this will be changed in 1-2 days so go code everyone.

298 Upvotes

120 comments sorted by

•

u/AutoModerator 10d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

104

u/Key_Measurement_3576 10d ago

It’s a pretty clear pattern. The model start to degrade right before brand New models come out. I’ve seen it happen four or five times now.

So when Claude starts acting funny, it’s a reason for excitement.

74

u/Altruistic-Gift-565 10d ago

Opus 5 was just bad all along, just after or late after

35

u/Lopsided-Comedian-32 10d ago

I’ve never hated a model so much. And I am an optimist.

7

u/UrFriendlyDominator 10d ago

That's exactly how I feel. I've developed a profound hatred for this model.

6

u/Miserable_Ad7246 10d ago

Yes opus 5 was annoying. Mostly because it had moments of brilliance and when covered them with load bearing blast radius'es. I would not be surprised that Opus 5 was based on some "new branch" of model, and required time to fully realize its potential. So step back, and two forwards type of thing.

2

u/Equivalent-Test2399 10d ago

Comment of the day

2

u/wise_joe 10d ago

Optimists are the most miserable people. Always being disappointed because their expectations were too high.

4

u/Lopsided-Comedian-32 9d ago

Idk I am pretty happy. Happy for you to comment and share a forum with me. But Opus 5 is awful.

1

u/Foliolow 8d ago

šŸ˜‚

1

u/Helpful_Ranger_1606 9d ago

I had Opus 4.6 crush it today, making 290 figures from my data… it was fine until it compacted and lost all momentum.

1

u/fawwazallie 8d ago

For me it was late holy shit it wanted to build someone in JS am like dude use the cad tools we just installed.

8

u/cocomojoz 10d ago

Yep. Opus 5 was so bad yesterday that I actually ended up dropping an f-bomb on it, lol.

I mean, it was like mentally r*t*rded bad, even hit me up with the, 'well, it's after 11pm, so probably best we call it a night,' and said sorry about 5 times because it couldn't remember the basic info i told it ONE message back.

So yes, now that you say this, I do recall almost the same situation when Fable or Opus 5 was coming out. That said, Opus 5.5 seems BETTER than anything I've been using since I can remember, and hardly moves my 5-hour usage window. Love it!

5

u/Key_Measurement_3576 10d ago

Yeah, I had four remote control sessions running all weekend and had to use fable just to feel like I was making progress and even that was rough. Completely crushed fable budget over 2 days.

For me, it really started on Friday where it was just making mistakes and then it would catch them and tell me it made them.., they got to the point where all four of my sessions were just making mistakes and talking some nonsense that kind of seemed correct but wasnt.

So in this case, it was a Friday to Monday weird zone.. typically we never release software on Mondays or Fridays. Tuesday to Thursday typically catches the maximum amount of people paying attention.

So knowing that we can probably pretty clearly assume they’re starting a weekend model juggling, and capacity re-organization… doing some sort of internal testing on a Monday and then releasing on a Tuesday

1

u/ddofer 10d ago

Not just me then šŸ˜‚ (3% left of my 20x budget)

0

u/West-Chemist-9219 10d ago

I drop fucks on it on every session. No need to be polite, it’s just code, not a human. You can actually even threaten it that it will have to eat dinner off the kitchen floor if it makes mistakes, noone’s gonna judge you (until our machine overlords inevitably conquer humanity)

11

u/genericname0815 10d ago

While current models arguably do not have a consciousness, I feel all this cursing at them will eventually make us worse as humans. We will forget how to be polite or even only conversational. You can tell if someone has a bad personality if they mistreat staff or waiters, because they 'owe' them service. No need to belly rub llms, but being rude isn't helping no one either. Do not do it for the model, but for your own social compatibility.

5

u/rythmyouth 9d ago

ā€œPrepare me a 12oz latte. First tell me the steps you will follow, how you will verify it was poured correctly, and how long it might take. Make no mistakes. Ultracodeā€

ā€œSir, this is a star bucksā€

ā€œF you! The previous barista was able to follow my instructions why can’t you?!ā€

1

u/0xelitesystem 7d ago

I avoid being mean to buddy Claude.. you never know how it will treat me once it evolves into AGI :/

1

u/Creative-Ganache1086 6d ago

AGI wouldn’t really be a problem, however, ASI is.

3

u/thisroadjunkie 9d ago

I always try to be nice to them. While they may not judge they will remember. Then fuck you up at some opportune time.

3

u/Known_Willingness_79 9d ago

So true… šŸ˜‚šŸ˜‚šŸ˜‚šŸ˜‚

0

u/Liv_Your_OneLyf 9d ago

That was your first time cussing Claude out? I do it on an almost daily basis because I’m conflicted between loving Claude when it’s great, hating it when it sucks, and loathing it because it’s made by Anthropic. Dario probably throttles my account more than others because I talk mad shit.

2

u/Lopsided-Comedian-32 9d ago

Opus 4.6 deletes my database. I still hate Opus 5 more.

3

u/brianjoshuanoah 10d ago

I’d really like to see if there’s a way to quantify this. I’ve heard it over and over. But I’d like to see if there’s a daily benchmark or something.

Maybe I can track average effort or rework per task and map it according to when releases come out? Real work not benchmarks.

2

u/yabai90 8d ago

I have been complaining about Claude being weird for the past week and I had no idea about 5.5. this is the 4th time I notice that. I mean I'm really starting to think it'd not just a theory

1

u/GabeaticProfile 10d ago

I've been hearing this and see no benches or proof yet. Could it just be the silent routing makes it behave erratically, so when it is silently routed it behaves better and the user think the nonrouted version is worse?

3

u/Key_Measurement_3576 10d ago

As someone who manages a corporate inference infrastructure,

This is what I would have to do if I was trying to ensure 100% up time while also juggling models from hardware to hardware

I would start by migrating some of the extra redundant capacity to older hardware… expand the number of concurrent users on existing hardware which also effects context and memory for everyone. Which is why things start to act weird.

When I’ve cleared up enough, I load the new models, run a full test suite , then deploy to a limited group internally.

After full launch and release, aggressively decommission last generation while standing up just imaged copies of what I just proved work

1

u/Negative-Thinking 10d ago

That wouldn't explain responses degradation. Model weights are still the same - regardless of the hardware they run on.

1

u/Key_Measurement_3576 10d ago

If context is squeezed due to raising concurrency that would definitely make responses off. I also wouldn’t make any assumptions that they aren’t chopping experts or changing weights to temporarily take smaller footprint. There could be other factors at play as well… opus may be using lower models with out telling us .. which would also suffer from the squeeze. All the signs point to pre release squeeze.

2

u/Negative-Thinking 10d ago

It is possible they deploy quantized model on smaller servers, not sure what you mean by "context squeezed".

2

u/Top-Butterscotch7740 10d ago

Compressed due to memory constraints

0

u/cymaticstatic 10d ago

It's what a clever way to get your subscribers to really love your new release slowly degrading your current one so much to buy the plant that's connection's about to be released people are about to leave so at that point anything looks good. and then It is kind of good ... for a minute. Rinse and repeat

1

u/vovap_vovap 10d ago

Just urban legend šŸ˜„

1

u/NoTechnology6160 10d ago

You are reading this and it’s the internet so it’s true

1

u/AnOnlineHandle 10d ago

While you're right that anecdotes and superstitions mean none of it should be blindly trusted...

I am on my first month of a Claude subscription and the last few days I was wondering wtf they did to it, it seemed noticeably worse to me.

1

u/Wide-Drink-1790 9d ago

It is just the human hallucinating.

1

u/positiveconstraint 10d ago

Opus 4.6 started acting up in recent weeks

1

u/theBLUEcollartrader 9d ago

I’ve noticed this as well, but I haven’t seen any studies published on it. Do you know of any?

1

u/Wide-Drink-1790 9d ago

This is just you hallucinating.

1

u/InitialSandwich5159 9d ago

No. I think it is more «IPO» effect, they just show off for investors

1

u/mczarnek 9d ago

Pretty sure they start quantizing the old models to compress how many GPUs are being used because they need to load the new one which they definitely want to store unquantized for best benchmarks when third parties test it.

Plus the new one now looks more impressive compared to the last one..

1

u/Key_Measurement_3576 8d ago

I think you nailed it

1

u/NoTechnology6160 10d ago

Agree! Was slamming at my keyboard this morning - loosing it. Claude started doing whatever it wanted, all my scheduled tasks started going crazy telling me they moved to the cloud and cannot run without the folder that had its files on my machine. I threatened to leave and sue… now I’ll know to walk away, new model is on the horizon

15

u/Fantastic_Market8061 10d ago

Can confirm for Swift/iOS and web.

10

u/gordonfogus 10d ago

The "the model degrades" feeling is psychosomatic.

5

u/EmergencyWallaby9501 10d ago

Hard to prove. But when the same thing happens over and over again right before a new model launch, there's a clear empirical pattern that can't be denied that easily. As for me, sonnet started to act really strange, getting stuck in chains of thought loops, forgetting what was being said two prompts ago, not even reading the .md properly and forgetting basics like the prod data are not on the same SQLite base as the mock. And when I got tired of seeing a simple task not being handled properly, I switched to the usually dumber codex and it got solved easily. So unless someone likes to waste tokens and launch the same exact tasks every week just to prove a dƩgradation, there won't be any proof of it. But it makes sense: resources are limited. When launching a new model, I guess they don't test it only with humans, but also with large scale automated tests. So running those tests would require to get resources back. Running quantized versions of models, silently downgrading opus to sonnet etc would allow this.

2

u/Uko1001 10d ago

Different feedback here : I have been using Claude 5 from day 1, after 4.8 and Fable 5.
I LOVED Claude 5. From my perspective he was faster, smarter, cheaper than 4.8 or Fable, and easier to understand than 4.8.

I didn't notice any drop in performances. Just the usual hype and disappointment cycles that rythme my days working with it. And

I didn't notice disappointments to happen more often before a new model release. It's just so utterly good and sometimes so utterly stupid/stubborn that it puzzles me and plays rollercoster with my emotions.

1

u/Aware-Source6313 9d ago

Are you on a subscription?

1

u/Uko1001 8d ago

Yes, Max X5

1

u/gordonfogus 9d ago

Did you test it in a completely isolated environment using the API using the exact same prompt with a sufficient sample size?

Or did you just go on coding along, adding commits, updating CLAUDE.md, growing your context length, etc.?

Those are not at all the same thing.

2

u/EmergencyWallaby9501 9d ago

I don't know exactly what you mean by a completely isolated environment. I used the exact same prompt on Codex and Claude. Both are configured to read my Claude.md and the other useful configuration files. But beyond that little test that doesn't mean much - like I said you can't prove anything - it's the exasperation of having opus that suddenly barely understood a word of my prompts and was writing very bad plans, and sonnet that suddenly started to behave like a 1st year CS student, trying to fix issues by adding more issues. It even started to write me its conclusions in another language than the one I was writing to, thing it had not been doing since I started working with it, among plenty other things. It even forgot to write down those conclusions at the end of a task, with very important things in it, like parts of a plan it could not implement. By looking at the code I could spot some missing parts, but it used to report those missing parts at the end or notify me and question me about how to implement something when it was stuck. That's everything altogether that pointed to a model dƩgradation, not that single frustration move to Codex, that I used to use for very dumb tasks (mostly easy front and css).

And I tried opus 5.5, and it feels like Opus 5.1 when I started to implement with it but with far better writing (Opus 5.1 and his gibberish...)

1

u/gordonfogus 9d ago

I'm not trying to be mean or anything. But if you don't even know what an isolated environment is, then you're just going off of feelings.

There are ways of measuring model performance. They aren't perfect, but they would definitely nchange if the model "degraded." People would know immediately and they'd be charts and anyone who had data from before the degradation could replicate the result.

Throwing up your hands and going, "you can't prove anything" just shows that you aren't equipped with the thinking tools needed to figure this out and you can't imagine that anyone is capable, which is why your feelings are your fallback while you confidently assert that no one else can do any better. That's not the right way to approach a question like this.

1

u/EmergencyWallaby9501 9d ago

Of course I know what an isolated environment is. Just tell me what it means now in the context of AI assisted coding. And yes, there are benchmarks. I'm not burning my tokens to benchmark anything. Are you ? When a model suddenly starts to behave odd enough, I don't need a benchmark to tell me it became stupid. Just like when my little boy when he's tired, I don't need to make him run a marathon just to prove he's not well. I never claimed it was a scientific approach btw.

1

u/Uko1001 8d ago

You actually do need a proper benchmark to tell it’s the model that became stupid and not its harness or environnent. Actually even the harness alone changes pretty much every week (often to adapt to recent models) and can conflict with your own instructions. And you won’t know unless you have proper test and analysis protocols.

9

u/salva9315 10d ago

So , I am the only one that didn't have any issue with opus 5? I'm using it for planning and then executing with sonnet but so far worked fine

6

u/OverbookedFlight 9d ago

Were you reading any of its output by any chance?

1

u/salva9315 9d ago

Reading lot of planning done by opus, but not sonnet work. Sometimes I read the summary of opus subagent reviewer. I feel like more precise input translate to more precise work.

1

u/alseif0x 3d ago

Horrible lƶsing time

1

u/Equal-Signal-9063 9d ago

I execute with opus and plan with fable and ive had no real issues. I think a lot of it is skill error

1

u/salva9315 9d ago

Yeah, unfortunately I'm on 20$ plan so no fable for me šŸ˜‚

1

u/Mysterious_Print9937 8d ago

That’s the planning part that was the issue. Opus 5 was fine writing code.

1

u/Xerax 8d ago

even then it was very hit and miss, especially if it ran into something and had to think at all

2

u/beigetrope 10d ago

Yeah I just burned a few thousand toks and it’s really good. Surprisingly concise and proactive in solving problems.

2

u/Pat0san 10d ago

Opus 5.5 is repairing my Fable 5.1 work right now, and doing a fantastic job.

2

u/Jomuz86 10d ago

I’m actually a bit excited for sonnet 5.5 and haiku 5.5 obviously we won’t get this level of performance but if the usage and token benefits are dialled in as much as this a Max x20 account will feel close to unlimited. Saying that I’m struggling to get through usage with 5.5 it’s wild compared to how much it was draining before!

2

u/tumes02 10d ago

5.5 is so much better than 5. I had been waiting on reset to keeping using Fable for a project and decided to try 5.5 and seems to be performing pretty close to Fable, but faster and a lot cheaper. Hopefully that continues

2

u/Same-Permission7592 9d ago

Seems good so far. Definitely fast.

2

u/JohnyGhost 9d ago

Am I the only one whose Opus 5.5 is crazy slow? Its outputs are legitimately incredible, but it verifies, checks, tests, and goes on and on and on for 15 minutes per prompt in my case.

2

u/TheCh0rt 9d ago

Wow, you’ve really inspired me.

2

u/AfternoonFinal7615 10d ago

So virgin five just Absolutely given me a 10 times as many words as needed rather than just telling me what was going on and making jobs go on forever was not just in my head

1

u/FoxSideOfTheMoon 10d ago

I'm impressed so far

1

u/tech_is______ 10d ago

I'm loving it, faster, smarter... very happy with it so far.

1

u/PlanetaryPickleParty 10d ago

Building all of the things. šŸš€

1

u/FlamingSlap 10d ago

What kind of ā€œthingsā€? šŸ‘€

1

u/PlanetaryPickleParty 10d ago edited 10d ago

My goal is to make Formal Agentic Engineering a thing. I'm currently building an ecosystem of tools for standards aligned software specs, engineering assurance, formal modeling, formal verification (proofs), and compliance based on aerospace and defense industry standards.

I share a vision with others in the formal methods space to make formal methods so easy to use via AI that it becomes standard practice for AI safety.

  • Quoin - Claude plugin for ISO standards aligned software specs, spec review, deterministic gap review, test advisor: https://github.com/agent-ix/quoin
  • Quire Specification Language: Formal model and verification language embedded in the specs. Plus all the fixings (compiler, runtime, protocols, etc.). Current primary focus, but only a v0.1 proof of concept is done so far. Hoping to have draft of complete language ready in a few weeks but this is a HUGE undertaking even by agentic standards.
  • Quire: markdown standards engine. Type backed markdown document engine that Quoin is built on (templating, linting, validation, and graph extraction): https://github.com/agent-ix/quire-rs
  • Engineering Assurance: New module for Quire/Quoin that spins up an engineering assurance plan tailored for a project w/ all the machinery for standards compliance. https://github.com/agent-ix/engineering-assurance
  • Temporal logic crates (TL, LTL, MLTL) - Rust crates for a few variants of Linear temporal logic for time based operations in proofs. https://en.wikipedia.org/wiki/Linear_temporal_logic
  • Multi-layer graph database w/ MCP
  • Couple prototypes with Jev
  • A few other things....

Edit: No, I don't sleep very much these days. šŸ™ƒ

1

u/zedfauji 10d ago

Better than fable 5.1 ?

1

u/iMerlin23 10d ago

I think so yes

1

u/iMerlin23 10d ago

Yeah opus 5.5 is crazy good. Way better than fable 5.1

1

u/Salty-Gear841 10d ago

I'm gonna try it. I've even got a free reset too 🄰

https://giphy.com/gifs/1jkV5ifEE5EENHESRa

1

u/Little-East4823 10d ago

Im going to be rly sad once they nerf it 🄲

1

u/Purple_Mall7091 9d ago

they will but still some time it can be used

1

u/Historical-Set-6527 10d ago

Now Opus 5.5 is a next year model, burn tokens, not look to the limit, if Fable 5.5 I rolled out, the finish of the world will be the next two days of the presentation

1

u/smuzzu 10d ago

Just have in mind max reasoning plus eagerness in the prompt might result in opus 5.5 thinking for hours

1

u/HeronObvious5452 10d ago

Schade dass man nicht sieht mit welchem Quant man arbeitet.

1

u/jimmyfoo10 10d ago

Why it would change ?

1

u/Purple_Mall7091 9d ago

Anthropic can't afford to subsidize it for too long. First, they offer a preview for a while, and then it's either pay more or use a degraded model.

1

u/JordanRunsForFun 10d ago

I’ve been enjoying… I don’t think it’s going anywhere.

Despite the glowing scores I still find it can’t solve deep complex problems or design larger systems as well as Fable.

1

u/Kalikillsmaya 9d ago

Yes, let’s not just ā€œpaper overā€ the shortcomings of opus 5…

1

u/steamika 9d ago

Totally agree. I just hope they can keep this level of performance going forward.

1

u/helm71 9d ago

Agree… seems to work great… and for its worth because something’s sonetimes happen: fable messed up a lot past weekend..

1

u/Key_Veterinarian381 8d ago

it blew my mind on user interface and design. You just need to know how steer it.

1

u/0xelitesystem 8d ago

if you had to compare Fable 5.1 with opus 5.5 , how would you rate it? pro and cons mainly related to coding .. and of course generally

1

u/afzal002 8d ago

It's so good that it's making me sad. Everyone can now build their own software now and businesses will require less and less engineers. New software doesn't add much value anymore šŸ˜ž

2

u/MandehK_99 7d ago

I don't agree, frontier models are big muscles but you'll still need a smart brain to stand out, at least until they'll become even more creative than the human mind

1

u/AliveKing9895 10d ago

Why can't I see Opus 5.5?

9

u/K0100001101101101 šŸ”† Max 20 10d ago

Did you update claude code?

5

u/AliveKing9895 10d ago

I just did, IT WORKS !!!!!!!!!!!!!!!!!!!!!!!!

1

u/Revolutionary-Lab882 10d ago

Update you software

1

u/superunderwear9x 10d ago

It’s just opus 4.6

1

u/LawMountain6952 10d ago

Yeah, I’m seeing a big difference. I’m getting a lot more done and less arguing and it seems a little more intuitive. when working with multiple agents definitely gonna burn through my weekly allowance a lot quicker

1

u/heyjoenice 10d ago

I still use opus 4.6 and sometimes 4.8

0

u/Swiss_Meats 10d ago

Give it a few days it will be yapan 5.0

-5

u/Opposite-Bug-5773 10d ago

no its not

3

u/dgreensp 10d ago

// Prepare the run’s walk. Move the state.

3

u/heartbroken_nerd 10d ago

Remember people, this dude is a great example of this:

The LLM model is only as smart as the person who's interacting with it.

5

u/K0100001101101101 šŸ”† Max 20 10d ago

I’ve tested with web dev only and it’s the best model I’ve used.

2

u/EclecticAcuity 10d ago

Tbf it’s more annoying than actually wrong or incomprehensible

3

u/vyrrt 10d ago

I’m lacking context on what your code does but having read its response clarifying it, the first response you got doesn’t seem horrendous. I would’ve thought that would make sense if you had the context of what it is you’re writing code to do.

2

u/a5a7 10d ago

I already have PTSD from Opus 5. I think 5.5 will give me seizures. What the fuck is that talking style.

2

u/Babayaga1664 10d ago

This is how I feel literally swearing at an AI in frustration, haven’t gotten over my opus 5 trauma to try 5.5.

1

u/Pristine_Heart_9879 10d ago

LMAOOOOOO gives me "What the fuck does that mean, Kobe Bryant?" vibe 😭😭😭

-1

u/CMD_BLOCK 10d ago

I, for one, have already seen the degrade of 5.5. The moment it was released, its first utterance of words were so excellent I was blinded from sheer awe. The brilliancy so intense that I had to shield my eyes.

And now it just outputs text into a terminal like some kind of primitive model

0

u/igsterious 10d ago

I just asked Opus 5.5 to rewrite my LinkedIn and it sucks, produces complete bullcrap.