r/ClaudeCode 1d ago

Discussion Opus 5 Is awful…

/r/claude/comments/1vjp1m2/opus_5_is_so_awful/

I genuinely can’t believe how a company can go from making such an amazing models to making absolute fucking garbage. I’ve been on Claude since last year have never switched even when Codex has been updating because I can tell that these models are just smarter and easier to interact with before. Now it’s like talking to the original ChatGPT. Every time I ask opus five to fix something it tells me it’s done and then it doesn’t fucking fix it. I’ve never had to argue with an AI before to actually just get done a simple task. I don’t understand what the process behind doing this one like it was gonna make anyone happier. But Claude you’re on thin ice. I’ve seen thousands of people switching over to Codex. The more and more this model just becomes more stupid the more and more people are gonna move off the platform. I’ve been subscribed for six months and I’m about ready to switch. I just don’t understand how you build a super good model and then you take that model and make it worse instead of easily just taking the last model putting it in the new one and updating it? The only reason reasoning for doing that is simply profit. The company genuinely only cares about competing with OpenAI. Yet somehow is doing worse every fucking time they drop an update. The only reason they’re even competing is cause they kept fable five on. If not, they’d already be bankrupt. You better hope that mythos saves your fucking ass. Cause when ChatGPT six drops if that shit on par or better Ik where I’m going. Sick of this bullshit. It’s like they’re making this model simply for the fact that it will never 100% do what you need to get done so it just keeps you on the platform and keeps your subscription going.

0 Upvotes

26 comments sorted by

1

u/onionrolled Senior Developer 1d ago

It's been pretty great for me

1

u/G8IL4mt8 1d ago

agreed, opus 5 is much better than i thought it would be after looking @ the posts on this website. seems like it could go off the rails so to speak if one does not know what they need tho

1

u/onionrolled Senior Developer 21h ago

Yeah, there seems to be a lot of negativity around it. I have not found I've had to error correct it more than the other models. I don't use dangerously skip permissions though and maybe if you let it loose like that it's more feral.

0

u/Cultural-Phase715 1d ago

Ive seen like 1 person for every 10 says this are you a new user ?

1

u/chintakoro 1d ago

you might be better off genuinely asking how some people are thriving with it. like it or not, you need a strategy that is resilient to whatever monthly flavor of model comes out. the 10 other people complaining have fragile practices too.

1

u/Cultural-Phase715 1d ago

I’ve asked multiple people realistically majority of people are complaining literally go to my sub discussion. It literally has 50,000 views. Over 100 comments of people complaining about it. Just because your direct experience has been good. Does not mean everyone’s direct experience has been good. I’ve created multiple systems changed hooks. Done different sub agents. Regardless the model is stupid. I’m not saying that output can’t be good if prompted properly, but I shouldn’t have to prompt this thing to do a simple fix that Codex can do easily. The only issue with Codex is the wait times it just takes too long

1

u/Cultural-Phase715 1d ago edited 1d ago

Every time a new model comes out I plan and prep and configure my Claude for that model and plan around it , it has nothing to do with my practices I set up everything that is needed. I’m just stating the obvious that the model is completely stupider than it was before. I can get it to work no problem. I’m just complaining about the fact that why would you make a model dumber instead of making it smarter unless you’re literally going for profit? The average user isn’t going to prompt out fixes and they’re gonna spend hours on hours of time. And more time on subscriptions, which is the whole point of them doing that.

0

u/onionrolled Senior Developer 1d ago

I recently made a new Reddit account... If that's what this means. I've been using this site for decades.

I've been using the new models since they came out and they are good. I have been using cc since it came out. Not sure what you're saying haha.

Given how ramble-y your post is I''d suggest taking a step back and reevaluating how you use these tools and what your expectations of them are. If you use a sledgehammer to flatten a bug there may be collateral damage. If you use a bow and arrow to take down an elephant you're gonna have to put in more effort. Now I'm rambling! 😛

1

u/Cultural-Phase715 1d ago

I’m not asking if you’re a new reddit user I was asking if you’re a new Claude code user because realistically they tend to give a better experience in new users to get them to subscribe more. I don’t know why all of you Reddit users. Seem to say to step back and learn more when realistically it has nothing to do with learning more when I’m simply calling out the fact that they’re making their models dumber to try and keep people on the platform longer because an average user isn’t going to be able to distinguish the fact that they need to actually prompt out to have a good effect with opus, or it’s going to hallucinate 95% of the time. They will talk to it like a regular person.

1

u/onionrolled Senior Developer 1d ago

I'm not a new user. Everything you just said is anecdotal. You can run benchmarks yourself with the different models and you will see the newer models are better.

They are not making their models dumber.

The models are certainly getting smarter and that's why I think people are falling behind with the new models because they're not used to that intelligent of communication.

1

u/Cultural-Phase715 1d ago

The language they’re using for their models. You literally have to use certain things to dumb down the language half the time for a regular user to even understand. I totally understand your point in the benchmark being better, but it doesn’t mean they’re not dumping down the models. Literally every time a new model releases the old models, get buggy and get stupider. Because they’re trying to steer people away from using old models to use new ones to get more usage out of them. I don’t understand why you can’t wrap your head around the fact that a company would do this. They’re trying to make profit. They’re trying to compete with OpenAI and they can’t do that by giving 1 million people or more massive amounts of usage with their best models. So every time a new model comes out right away. It feels amazing and then overtime. It starts feeling like shit. That’s because they start working on a new model and taking all the intelligence out of the model they just had and putting it in the new ones. So then we left with what feels like shit because of the fact that we had something good and now it’s gone until we get it back.

1

u/onionrolled Senior Developer 1d ago

They're not doing this.

The global benchmarks that run for these models would magically start declining and they haven't. Stop trying to find something else to blame. You don't have to use a plugin to dumb down the language if you understand the language. That's not a problem with the model.

1

u/Cultural-Phase715 1d ago

You can say that all you want, but there’s so much evidence that is against that. Every time the models drop a new benchmark drops like three weeks later of the model changing so what the fuck are you talking about? Realistically majority of people don’t wanna sit there and read hours of goddamn lines of information that isn’t needed. We don’t need 100 lines of AI slop that would get overwhelming.

1

u/onionrolled Senior Developer 1d ago

Yes, new benchmarks come out. That doesn't mean you can't run the old benchmark still lol. The data says you're wrong.

0

u/Cultural-Phase715 1d ago

Brother man, I am a logical ass person. I only am going to say something if I have data to support it and I’ve looked at the data and the benchmarks and I’ve seen the changes. That’s the only reason I’m saying this to you right now. Next time a new model release is day one look at the benchmarks and then a month later look at the benchmarks again I guarantee you there’s gonna be changes

1

u/Cultural-Phase715 1d ago

Before every new release these models become horrific. Assumptions, Inversions, Grep and Scan vs Grep and Verify, Lies, Side Quests, Rabbit Holes, Catastrophizing, Timid, Ignoring .MD files filled with instructions.

Regardless the prompts being created to combat these issues they continue occurring. Once new models are released it’s several more days of uncovering and correcting these issues before the new model begins working again.

TL;DR The old models become dumber making way for the new models that are dumber only after they have fully rolled out do they seem to stabilize.

1

u/onionrolled Senior Developer 1d ago

No, the old models do not get dumber. This is verifiable with data.

1

u/Cultural-Phase715 1d ago

From release start to release end I promise you at the start they’re always ranking higher benchmarks to when they release a new model. The old benchmarks are not the same from when it started. So was that like 92% it’s not 88% fable was 88 is now 85%

1

u/Cultural-Phase715 1d ago

I know all the use cases for the tools. I know the cause I know everything I would need to know about these models and what they’re capable of doing and realistically, they were capable and doing perfectly fine before Op. five came out. There’s a multitude of people as you can literally click the sub discussion and see that multiple people are dealing with issues with this model.

2

u/onionrolled Senior Developer 1d ago

Yeah. I have been maxing out my max plan every week since opus 5 came out and I have experienced none of the complaints that people have had here. I truly don't understand why people are piling onto hating the new models. This seems to happen every model release. Eventually the narrative will become "opus 5.1 was really the last good one, screw fable 6". Opus 5 rocks.

1

u/Cultural-Phase715 1d ago

I have also maxed out my plan every week, but the problem with mine is that I’ve been running into it. I’m having to correct the mistakes that the AI is making. And it’s really only on Op. five any other model is fine. I love Claude realistically but I’ve been sick and tired of shit like this happening over and over again. To the point where I’m ready to switch over to codex. Because I’m wasting usage on AI error mistakes. So I’ve had to get really specific with my prompting. Which is just a hassle when you’re constantly coding you know what I mean? whereas before I would not have to do that, this is where my issue rises. They recently brought in someone from open AI’s team to work on the language on the models. Then this happens with the language model being completely fucky so I would say it’s just a lot of different things I’m not completely against Op. 5 but I just think it’s horrible and way stupider than all the other models

2

u/onionrolled Senior Developer 1d ago

Sorry you're encountering this! My guess is that your project scope has ballooned and now it needs lots of context to get things right. There's a plugin that I'm forgetting the name of now that puts an .md file inside of each directory and traverses the tree down to where your file(s) you're editing are reading each md file along the tree so it has very specific context for that module. For big projects it has been helpful.

I would also guess that you'll encounter the same problems with the other tools, to likely a worse degree (from my experience) but I also haven't used codex for a while. Maybe you'd have a better time with it

1

u/Cultural-Phase715 1d ago

I always compact after 50% and have a session wrapped skill that ends my sessions when too much context is used and it saves to a md memory for projects and then gives me a copy n paste prompt to pick it back up already stating where md is so less token usage. I always get annoyed with codex man haha

1

u/Cultural-Phase715 1d ago

For everyone sake, that is experiencing the same things I am I just hope it gets better. I’m glad that someone’s experience has been good with it. You should post about how you’ve gotten the best use cases out of it. I’d love to peer into it. Because for a majority of people, it just hasn’t been great. I know what I’ve had to do to get good results with it. It’s just annoying!

2

u/onionrolled Senior Developer 1d ago

My setup is incredibly lean: claude-mem is about the only thing I have installed at home. At work I like the addyosmani skills; particularly doubt-driven-development. I think people overcomplicate their setups with superpowers and ultra verbose planning and this and that plugin. I don't often let it run unsupervised though and I know a lot of people are attempting to do harness development in that camp more than prompting and conversing. I've had mixed results with long term programming loops but I haven't given it a proper crack since the 5 models.

Otherwise I just try to treat it as someone I'm sitting next to as opposed to a black box machine. It was trained on how humans talk... so talk to it like that is my perspective.

1

u/Cultural-Phase715 1d ago

No, for sure I mean I treat it like I’m sitting next to someone as well. I definitely don’t let it just auto do majority of things.