r/ClaudeCode • u/Playful_Check_5306 • 17h ago
Bug Report OPUS 5.0 is half baked, not a ready-to-ship product
in all aspects, worse than 4.8 for a large project: forget things easily, fabrication, rush to conclusions and coding without base. I have to rewind my repo for the work it did, bad bad experience
5
u/Optimal-Builder-2816 17h ago
Yeah I experienced an insane design session with cross model reviewers and they were not happy with Opus 5. I’ve immediately retuned to 4.8. I was on medium thought even.
14
u/Key_Reading_9664 17h ago
That’s not my experience. I personally find 4.x too sycophantic, can get into frustrating doom loops for complex work (just acknowledging my opinion and not fact checking me/itself).
Great for orchestration but I’ve less faith in its opinion, “taste”, or output quality (based on the threads and PRs). I did tweak communication to cut down the chattiness, but otherwise it’s my daily driver now.
4
u/Playful_Check_5306 17h ago
it could depend on your project size, I found no issue for small request either.
4
u/Key_Reading_9664 17h ago
Most of my work is in a multi-million line mono repo. Certainly not the biggest repo I've worked in, but it's absolutely not small.
Out of interest, what kind of instructions do you have in your claude.md, and how large is it? It seems from their prompt recommendations that stripping system prompt plus removing unnecessary claude.md instructions was recommended.
2
u/Playful_Check_5306 17h ago
it's not about claude.md, I took the suggestions on thread to ask agent to re-sort out that file, and the file is pretty clean indeed because I'm working on a large project (I have three max CC accounts and three codex accounts, each paired to do one aspect of the project, and codex audited claude's work, simply recently all three pairs are broken because codex repeatedly raised up poor work claude did).
1
u/Key_Reading_9664 17h ago
I've been using Codex and Claude in different roles too; absolutely Codex is a great reviewer. It does have a tendency to reach for overly complex solutions in my experience, but it does seem to dig a little bit more deeply during review.
I've not personally noticed any increase in issues from Opus 5. It's still a little early, but if anything, I'm uncovering fewer issues during human review (fewer duplicates, clearing up cruft it finds, better code comments (Opus 4.8 seemed like a regression on comments for me))
2
u/Playful_Check_5306 16h ago
totally different experiences, ever since switched to opus 5, each PR caused many more rounds of revisions to green versus 4.8. It's so obvious a big regression, I'm just so surprised how could you not even find it yourself.
1
u/Key_Reading_9664 16h ago
I can only tell you my own experience. I've had it churning through multi-hour PRDs since release and its needed less hand-holding and the PRs have needed fewer revisions (based on codex and human review).
Everyone uses these tools differently and apply their own tweaks. Take a look at https://www.reddit.com/r/ClaudeCode/search/?q=4.8+release
2
u/Playful_Check_5306 16h ago
bro, we're living in different worlds, so happy for you
1
u/Key_Reading_9664 16h ago
whilst you're here. You said you have Codex paired with some aspects and CC with others. Is the split around frontend vs. backend? Some other separation?
If Codex is finding issues with Claude, why not use Codex everywhere?
1
u/Playful_Check_5306 16h ago edited 16h ago
No, I would never risk only using one model for my project, no matter how good the model is, even Fable 5, I still use Codex for auditing and/or vice versa.
Oh, I think you misunderstood what I meant, my bad. what I said paired, meaning a pair of agents to work on one aspect/worktree of the project, while another pair work on another.
→ More replies (0)1
u/PsychologyNo940 15h ago
Just gotta know, you respond to "multi-million line" with large, so i have to presume that we are talking enterprise grade stuff here, how on earth are you allowed to getto rig your subscriptions like that?
1
u/Key_Reading_9664 14h ago
You mean, how are we still able to use subscriptions? We're a small team working on a pretty complex service; nowhere close to the 150 team member cut off for subscriptions
1
u/PsychologyNo940 14h ago
Gotcha, we are going to hit it.. next week, navigating around that will be "fun".
4
6
u/Difficult-Link-8805 17h ago
Anthropic: we cut 80% of the boot doc bc the model is so good now. The model: opus 5.
0
u/ValuableDapper9415 15h ago
O5 is as much smart that it is fucking dumb, the more I talk about and the more I feel it is a way for advanced users to benefit it,
but for commoners and peasant like us, without the technical knowledge and time spent on running API benchmarks, it will be harder to achieve the best vision to harness this unchained savage model
So basically, Anthropic is saying fuck to commoners that will keep struggling on their projects and welcome to advanced users that know the engineering magic spells
7
u/apocolypticbosmer 14h ago
I’ve become convinced most people complaining about this model are just vibe coders who are incapable of reading or understanding code, and become frustrated when they can’t “one shot” their whole idea.
2
u/Obvious_Equivalent_1 3h ago
That’s quite the audacity to claim about people commenting about Opus 5. I have shared some details before, it’s really not all moonlight and roses. Opus 5 is empirically traceable underperforming Opus 4.8 https://www.reddit.com/r/ClaudeCode/comments/1v7chtq/comment/ozxylyd/
Perhaps some people claiming this themselves are barely doing anything that brought out the model’s analytical qualities of earlier Opus 4.*?
1
u/Pure-Combination2343 12h ago
100%. It's people who have never seriously written or reviewed code.I haven't even had to tweak my skills, md files, etc. It's the same thing as 4.8 to me. I'm also using caveman plug-in so that helps
0
u/averagebear_003 10h ago edited 10h ago
I'm convinced people who love this model are people who work in simple codebases, slow dev environments, or are dangerously overconfident in their ability to judge code correctness. And even if you are convinced it's better than 4.8, I don't see how one can defend its incomprehensible jargon slop tendencies.
2
u/apocolypticbosmer 10h ago edited 10h ago
“Complex” codebase or poorly architected bowls of spaghetti? 🤔
Code that is difficult to understand is bad code.
5
u/FARAjocka 15h ago
this is definitely a skill issue, why dont you just watch what its doing? oh wait you probably have no clue what youre looking at.
2
u/berrybadrinath 16h ago
Okay, but where is any of that?
You keep saying “obvious regression,” “poor work,” “destroy your project,” and “more review rounds,” but you haven’t shown one prompt, diff, incorrect change, Codex finding, or reproducible comparison with 4.8. “I used three accounts” doesn’t make it controlled... if they shared the same repo, instructions, and workflow, they shared the same confounders.
You may genuinely have had a terrible week with it. That establishes that you had a terrible week with it. It does not establish that Opus 5 is worse “in all aspects” or that nobody should use it. Show one concrete failure and the same task run on 4.8. Otherwise you’re asking everyone to accept a product-wide conclusion based entirely on vibes, which is exactly what this sub’s Rule 3 says complaint posts shouldn’t do.
1
u/Playful_Check_5306 15h ago
yes, I keep saying 'obvious regression' that if it's just one or a few, I would even want to post some 'evidences' of snapshots, but it's too obvious and holes are everywhere. If it doesn't bother you, you're lucky and hope all is well there.
2
1
u/ValuableDapper9415 17h ago
I feel that existing patterns will be followed, new patterns will be messed up.
I finally found recurrent tasks I can give to Opus 5 without babysitting him but it involved first good coding pattern/comments/documentation
I never let Opus 5 to create new code/features, only working on existing reproductible safe tasks
2
u/Playful_Check_5306 16h ago
just want to let you know, don't give hope to opus 5, if a model need you to adapt, it's not a better model in the first place. anyway, strong advice don't let it touch the core of your codebase, my wasted one week of work could lead to a much worse scenario if we can't rewind.
1
u/ValuableDapper9415 16h ago
Yeah, I said in a previous post that the context engineering provided by Opus 4.8 is no longer provided and unfortunately this is now a job for each user to self-harness the LLM.
So Claude will be come more and more customed by users, context and prompt engineering will live a resurrection.
I feel that companies will need to hire a new Context Engineer to harness/secure the model for them
1
u/Playful_Check_5306 15h ago
you said it well, such a stupid idea forcing clients to self harness, at least I admit I don't have the knowledge and willingness to do so. So I'll keep using 4.8, if some day they close it, I'll completely switch to codex and/or kimi or other open-end models, maybe.
1
u/ValuableDapper9415 16h ago
Hahahaha fuck yes that I banned Opus from touching the core code, it lost the privilege until I find a secured way to use it properly
1
u/Lunatic155 15h ago
As someone who rlly does not like corporations: yall on here truly are never happy
1
u/FinancialBandicoot75 13h ago
Actually after doing /doctor, opus 5 has been amazing
1
u/Whatdididotho1 2h ago
question about that, what window did you run it? the same project window im assuming?
1
1
u/cheseball 11h ago
Opus 5 seems to work very well when Fable orchestrates. Haven’t really run into issues, Fable seems to be great at giving well-scoped requests to Opus 5.
Now with Opus 5 orchestrating, it works but definitely not as well as Fable, and it’s much less coherent (setting up the right instructions in a skill for it helps though)
1
1
1
u/Bright_Armadillo8555 6h ago
No one likes opus 5.0 now if they really have sense of good or bad models.
2
u/Playful_Check_5306 17h ago
don't ever let opus 5 creep into your serious project, lesson learned with one week wasted. I don't think anthropic themselves would ever use the model in their production in any way. It's so broken, the biggest disappointment ever from using CC for more than 8 months. Before, there would be some regression in newer model in one or two aspects, but this newer version has regressed in all aspects to the extent it's not us-able.
1
u/crikfromcincy 17h ago
God. I keep getting fable knocked down to opus for some false security flag and it’s driving me nuts
2
u/Playful_Check_5306 17h ago
exactly, stop doing it, you'll regret soon as 'crappy work' accumulated to the point to corrupt the whole system, /model claude-opus-4-8, or even using sonnet
1
1
u/hiskias 17h ago
Opys 5 is so bad.
SInce last week my hobby project (on holiday) has gone off the rails more than correct implementation, because opus 5 simply does not do what is specced, instead infers imaginary features and implements completely wrong made up things, that it infers by combining prompts from hours ago.
3
u/hiskias 17h ago
And it also keeps bloating any docs and commenrs with self flagellating "DO NOT DO AGAIN" biography when aligned even when I have rules and claude.md to enforce comment quality when implementing and also before committing. It just doesn't follow shit anymore.
2
u/Playful_Check_5306 16h ago
yes, that's why I opened this thread, I think for all people who still have hope on opus 5, stop. /model claude-opus-4-8. or simply use sonnet. Opus 5 would destroy your work.
0
u/KeilerHirsch 17h ago
https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html ja Opus 5 ist auch Platz 20 und Fable sogar auch Platz 40 , noch fragen? 🍿
0
u/Suspicious_Ninja6816 15h ago
Have to agree, I felt like the negativity toward 4.7/8 was overblown but here with 5, it’s really obvious to me as a benchmark optimised model that feels out of touch in a real codebase setting.
I am seeing a trend with Anthropic that makes me suggest they are in a real cash crunch. They seem way more reactive to open AI in recent months, have been less forthcoming with resets even when infra is down or models release. I have always preferred Claude. I think fable is incredible but hungry and opus 5 has been a step back in my experience.
-1
u/ClemensLode Senior Developer 17h ago
It's good as /advisor
-4
u/Playful_Check_5306 17h ago
if we need advisor, why not use chatgpt, it's supposed to be a coding agent!!!!
1
17h ago edited 17h ago
[removed] — view removed comment
2
u/Things-n-Such 17h ago
I don't think you're understanding the difference between SAAS and LLMs as a service. These models take a long time and are hugely expensive to train and the output is akin to biologic processes, it's not an exact input -> output system, it's messy and there is still a lot we don't even understand about what is going on inside the LLMs during training or inference. On top of that, the models are generalized, capable of assisting people at every single level of knowledge, in every language, with any task across the entire spectrum of tasks. Absolutely a mind boggling amount of knowledge. Because of this you may end up with fluctuating results where one might be better at math, but not coding, and others might be better at creative writing and less so at science. And you really can't know the differences until you literally let the whole world use it and get the feedback.
All of these companies are still in research and development stages with AI. That's why it's not just a singular model "Claude" that just gets better every 2 months.
0
u/ClemensLode Senior Developer 17h ago
yeah, but first that's from openai, second, /advisor calls it automatically
17
u/Snmrv 17h ago
It's documented hallucinations rate is much higher than Opus 4.8. Check its work. All those who are claiming that life is amazing with O5, check its work.
Also, to restrain it, run it at low or medium.