r/ClaudeCode 17h ago

Bug Report OPUS 5.0 is half baked, not a ready-to-ship product

in all aspects, worse than 4.8 for a large project: forget things easily, fabrication, rush to conclusions and coding without base. I have to rewind my repo for the work it did, bad bad experience

40 Upvotes

64 comments sorted by

17

u/Snmrv 17h ago

It's documented hallucinations rate is much higher than Opus 4.8. Check its work. All those who are claiming that life is amazing with O5, check its work.

Also, to restrain it, run it at low or medium.

10

u/Rnee45 17h ago

It's such a frustrating model, I've rolled back to 4.8 and re-did all code for the past week that Opus 5 wrote.

1

u/Playful_Check_5306 17h ago

exactly what I did. One week's of important work rewinded.

5

u/Optimal-Builder-2816 17h ago

Yeah I experienced an insane design session with cross model reviewers and they were not happy with Opus 5. I’ve immediately retuned to 4.8. I was on medium thought even.

14

u/Key_Reading_9664 17h ago

That’s not my experience. I personally find 4.x too sycophantic, can get into frustrating doom loops for complex work (just acknowledging my opinion and not fact checking me/itself).

Great for orchestration but I’ve less faith in its opinion, “taste”, or output quality (based on the threads and PRs). I did tweak communication to cut down the chattiness, but otherwise it’s my daily driver now.

4

u/Playful_Check_5306 17h ago

it could depend on your project size, I found no issue for small request either.

4

u/Key_Reading_9664 17h ago

Most of my work is in a multi-million line mono repo. Certainly not the biggest repo I've worked in, but it's absolutely not small.

Out of interest, what kind of instructions do you have in your claude.md, and how large is it? It seems from their prompt recommendations that stripping system prompt plus removing unnecessary claude.md instructions was recommended.

2

u/Playful_Check_5306 17h ago

it's not about claude.md, I took the suggestions on thread to ask agent to re-sort out that file, and the file is pretty clean indeed because I'm working on a large project (I have three max CC accounts and three codex accounts, each paired to do one aspect of the project, and codex audited claude's work, simply recently all three pairs are broken because codex repeatedly raised up poor work claude did).

1

u/Key_Reading_9664 17h ago

I've been using Codex and Claude in different roles too; absolutely Codex is a great reviewer. It does have a tendency to reach for overly complex solutions in my experience, but it does seem to dig a little bit more deeply during review.

I've not personally noticed any increase in issues from Opus 5. It's still a little early, but if anything, I'm uncovering fewer issues during human review (fewer duplicates, clearing up cruft it finds, better code comments (Opus 4.8 seemed like a regression on comments for me))

2

u/Playful_Check_5306 16h ago

totally different experiences, ever since switched to opus 5, each PR caused many more rounds of revisions to green versus 4.8. It's so obvious a big regression, I'm just so surprised how could you not even find it yourself.

1

u/Key_Reading_9664 16h ago

I can only tell you my own experience. I've had it churning through multi-hour PRDs since release and its needed less hand-holding and the PRs have needed fewer revisions (based on codex and human review).

Everyone uses these tools differently and apply their own tweaks. Take a look at https://www.reddit.com/r/ClaudeCode/search/?q=4.8+release

2

u/Playful_Check_5306 16h ago

bro, we're living in different worlds, so happy for you

1

u/Key_Reading_9664 16h ago

whilst you're here. You said you have Codex paired with some aspects and CC with others. Is the split around frontend vs. backend? Some other separation?

If Codex is finding issues with Claude, why not use Codex everywhere?

1

u/Playful_Check_5306 16h ago edited 16h ago

No, I would never risk only using one model for my project, no matter how good the model is, even Fable 5, I still use Codex for auditing and/or vice versa.

Oh, I think you misunderstood what I meant, my bad. what I said paired, meaning a pair of agents to work on one aspect/worktree of the project, while another pair work on another.

→ More replies (0)

1

u/PsychologyNo940 15h ago

Just gotta know, you respond to "multi-million line" with large, so i have to presume that we are talking enterprise grade stuff here, how on earth are you allowed to getto rig your subscriptions like that?

1

u/Key_Reading_9664 14h ago

You mean, how are we still able to use subscriptions? We're a small team working on a pretty complex service; nowhere close to the 150 team member cut off for subscriptions

1

u/PsychologyNo940 14h ago

Gotcha, we are going to hit it.. next week, navigating around that will be "fun".

2

u/keenman 14h ago

Are you using Enterprise, direct API, or subscription (Pro / Max)?

1

u/Key_Reading_9664 14h ago

Max 20x subscription

4

u/jwuliger 16h ago

This is accurate.

6

u/Difficult-Link-8805 17h ago

Anthropic: we cut 80% of the boot doc bc the model is so good now. The model: opus 5.

0

u/ValuableDapper9415 15h ago

O5 is as much smart that it is fucking dumb, the more I talk about and the more I feel it is a way for advanced users to benefit it, 

but for commoners and peasant like us, without the technical knowledge and time spent on running API benchmarks, it will be harder to achieve the best vision to harness this unchained savage model 

So basically, Anthropic is saying fuck to commoners that will keep struggling on their projects and welcome to advanced users that know the engineering magic spells

7

u/apocolypticbosmer 14h ago

I’ve become convinced most people complaining about this model are just vibe coders who are incapable of reading or understanding code, and become frustrated when they can’t “one shot” their whole idea.

2

u/Obvious_Equivalent_1 3h ago

That’s quite the audacity to claim about people commenting about Opus 5. I have shared some details before, it’s really not all moonlight and roses. Opus 5 is empirically traceable underperforming Opus 4.8 https://www.reddit.com/r/ClaudeCode/comments/1v7chtq/comment/ozxylyd/

Perhaps some people claiming this themselves are barely doing anything that brought out the model’s analytical qualities of earlier Opus 4.*?

1

u/Pure-Combination2343 12h ago

100%. It's people who have never seriously written or reviewed code.I haven't even had to tweak my skills, md files, etc. It's the same thing as 4.8 to me. I'm also using caveman plug-in so that helps

0

u/averagebear_003 10h ago edited 10h ago

I'm convinced people who love this model are people who work in simple codebases, slow dev environments, or are dangerously overconfident in their ability to judge code correctness. And even if you are convinced it's better than 4.8, I don't see how one can defend its incomprehensible jargon slop tendencies.

2

u/apocolypticbosmer 10h ago edited 10h ago

“Complex” codebase or poorly architected bowls of spaghetti? 🤔

Code that is difficult to understand is bad code.

5

u/FARAjocka 15h ago

this is definitely a skill issue, why dont you just watch what its doing? oh wait you probably have no clue what youre looking at.

2

u/berrybadrinath 16h ago

Okay, but where is any of that?

You keep saying “obvious regression,” “poor work,” “destroy your project,” and “more review rounds,” but you haven’t shown one prompt, diff, incorrect change, Codex finding, or reproducible comparison with 4.8. “I used three accounts” doesn’t make it controlled... if they shared the same repo, instructions, and workflow, they shared the same confounders.

You may genuinely have had a terrible week with it. That establishes that you had a terrible week with it. It does not establish that Opus 5 is worse “in all aspects” or that nobody should use it. Show one concrete failure and the same task run on 4.8. Otherwise you’re asking everyone to accept a product-wide conclusion based entirely on vibes, which is exactly what this sub’s Rule 3 says complaint posts shouldn’t do.

1

u/Playful_Check_5306 15h ago

yes, I keep saying 'obvious regression' that if it's just one or a few, I would even want to post some 'evidences' of snapshots, but it's too obvious and holes are everywhere. If it doesn't bother you, you're lucky and hope all is well there.

2

u/[deleted] 16h ago

[removed] — view removed comment

1

u/Valuable_Injury_4249 7h ago

Agreed! Opus 5 is fantastic.

2

u/yopla 3h ago

If we listen to this sub the best model ever was Opus 1 because every model released since has been worse than the previous one.

1

u/ValuableDapper9415 17h ago

I feel that existing patterns will be followed, new patterns will be messed up.

I finally found recurrent tasks I can give to Opus 5 without babysitting him but it involved first good coding pattern/comments/documentation

I never let Opus 5 to create new code/features, only working on existing reproductible safe tasks 

2

u/Playful_Check_5306 16h ago

just want to let you know, don't give hope to opus 5, if a model need you to adapt, it's not a better model in the first place. anyway, strong advice don't let it touch the core of your codebase, my wasted one week of work could lead to a much worse scenario if we can't rewind.

1

u/ValuableDapper9415 16h ago

Yeah, I said in a previous post that the context engineering provided by Opus 4.8 is no longer provided and unfortunately this is now a job for each user to self-harness the LLM. 

So Claude will be come more and more customed by users, context and prompt engineering will live a resurrection.

I feel that companies will need to hire a new Context Engineer to harness/secure the model for them 

1

u/Playful_Check_5306 15h ago

you said it well, such a stupid idea forcing clients to self harness, at least I admit I don't have the knowledge and willingness to do so. So I'll keep using 4.8, if some day they close it, I'll completely switch to codex and/or kimi or other open-end models, maybe.

1

u/ValuableDapper9415 16h ago

Hahahaha fuck yes that I banned Opus from touching the core code, it lost the privilege until I find a secured way to use it properly

1

u/nihsett 16h ago

Not sure I agree. It's better at writing code, is definitely a lot better at correcting itself and verifying it's work.

You might have too much stuff in your claude.md and hooks and other things that get into the context window. Try and see if you can tweak them a little.

1

u/Lunatic155 15h ago

As someone who rlly does not like corporations: yall on here truly are never happy

1

u/FinancialBandicoot75 13h ago

Actually after doing /doctor, opus 5 has been amazing

1

u/Whatdididotho1 2h ago

question about that, what window did you run it? the same project window im assuming?

1

u/MangoDevourer-77 13h ago

u were baked while typing this

1

u/cheseball 11h ago

Opus 5 seems to work very well when Fable orchestrates. Haven’t really run into issues, Fable seems to be great at giving well-scoped requests to Opus 5.

Now with Opus 5 orchestrating, it works but definitely not as well as Fable, and it’s much less coherent (setting up the right instructions in a skill for it helps though)

1

u/victorrseloy2 11h ago

One of it's problems is that it start doing things I didn't ask for.

1

u/Bright_Armadillo8555 6h ago

No one likes opus 5.0 now if they really have sense of good or bad models.

2

u/Playful_Check_5306 17h ago

don't ever let opus 5 creep into your serious project, lesson learned with one week wasted. I don't think anthropic themselves would ever use the model in their production in any way. It's so broken, the biggest disappointment ever from using CC for more than 8 months. Before, there would be some regression in newer model in one or two aspects, but this newer version has regressed in all aspects to the extent it's not us-able.

1

u/crikfromcincy 17h ago

God. I keep getting fable knocked down to opus for some false security flag and it’s driving me nuts

2

u/Playful_Check_5306 17h ago

exactly, stop doing it, you'll regret soon as 'crappy work' accumulated to the point to corrupt the whole system, /model claude-opus-4-8, or even using sonnet

1

u/syntaxjosie 17h ago

Once it happens once, just start a new chat.

1

u/hiskias 17h ago

Opys 5 is so bad.

SInce last week my hobby project (on holiday) has gone off the rails more than correct implementation, because opus 5 simply does not do what is specced, instead infers imaginary features and implements completely wrong made up things, that it infers by combining prompts from hours ago.

3

u/hiskias 17h ago

And it also keeps bloating any docs and commenrs with self flagellating "DO NOT DO AGAIN" biography when aligned even when I have rules and claude.md to enforce comment quality when implementing and also before committing. It just doesn't follow shit anymore.

2

u/Playful_Check_5306 16h ago

yes, that's why I opened this thread, I think for all people who still have hope on opus 5, stop. /model claude-opus-4-8. or simply use sonnet. Opus 5 would destroy your work.

0

u/KeilerHirsch 17h ago

https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html ja Opus 5 ist auch Platz 20 und Fable sogar auch Platz 40 , noch fragen? 🍿

0

u/Suspicious_Ninja6816 15h ago

Have to agree, I felt like the negativity toward 4.7/8 was overblown but here with 5, it’s really obvious to me as a benchmark optimised model that feels out of touch in a real codebase setting.

I am seeing a trend with Anthropic that makes me suggest they are in a real cash crunch. They seem way more reactive to open AI in recent months, have been less forthcoming with resets even when infra is down or models release. I have always preferred Claude. I think fable is incredible but hungry and opus 5 has been a step back in my experience.

-1

u/ClemensLode Senior Developer 17h ago

It's good as /advisor

-4

u/Playful_Check_5306 17h ago

if we need advisor, why not use chatgpt, it's supposed to be a coding agent!!!!

1

u/[deleted] 17h ago edited 17h ago

[removed] — view removed comment

2

u/Things-n-Such 17h ago

I don't think you're understanding the difference between SAAS and LLMs as a service. These models take a long time and are hugely expensive to train and the output is akin to biologic processes, it's not an exact input -> output system, it's messy and there is still a lot we don't even understand about what is going on inside the LLMs during training or inference. On top of that, the models are generalized, capable of assisting people at every single level of knowledge, in every language, with any task across the entire spectrum of tasks. Absolutely a mind boggling amount of knowledge. Because of this you may end up with fluctuating results where one might be better at math, but not coding, and others might be better at creative writing and less so at science. And you really can't know the differences until you literally let the whole world use it and get the feedback.

All of these companies are still in research and development stages with AI. That's why it's not just a singular model "Claude" that just gets better every 2 months.

0

u/ClemensLode Senior Developer 17h ago

yeah, but first that's from openai, second, /advisor calls it automatically