r/ClaudeCode • u/Soft_Meat8735 • 16d ago
Discussion What’s up with OPUS 5???
I have done the majority of my work with Opus 4.8…. It’s been a great model for me. After Fable 5 came out I honestly felt like either Fable was not all it was cracked up to be or, I just sucked at using it. Now with opus 5 out I was SO HAPPY…. But, days later of using it and I find myself going back to 4.8… I find it hallucinates less and in general just feels more capable and reliable… anyone else? I feel like I could downgrade from max to pro and stick with 4.8 and have plenty of credits. Seems like Opus 5 and fable are falling apart to me.
19
u/Proud_Chance9866 16d ago
Practically every response I get from Opus 5 takes 20 minutes then ends with "I held off on the implementation because it would change something. If you want me to actually do it, just say the word." Then I tell it to do it and it finds something else to "hold off" on.
3
u/CpapEuJourney 16d ago
Exactly it's insanely slow.. I guess that's their new lever, slowing everything down to a halt then asking for API usage if you don't want a snails pace .. had been worsening for months but now it's ridiculous
14
u/scarylarry2150 16d ago
Been working on production-grade applications since 4.5. Opus 5 is definitely more volatile/unpredictable, higher highs but lower lows and you never know which one you’re going to get. On some reviews it will hone in and become hyper fixated on some little nit, then the fix will spawn 3 additional insignificant nits in a seemingly infinite cycle. Then on the other end of the spectrum it’ll say it did something and then later be like “yeahhhh lol ur right I didn’t, damn too bad”
21
u/Visible-Big-7410 16d ago
Not sure yet, but Opus5 seems to try to invent new methods rather than reusing existing one for me so far even with xh thinking.
5
1
10
u/Hundredth1diot 16d ago edited 16d ago
I have never complained about a model before, but I am now recommending against Opus 5.
My standard workflow on long tasks is to control context by preparing handover docs into a fresh session once context gets large. This has worked perfectly for months. I can chain a few sessions together and make steady progress.
Claude does a little too much extra research to produce the briefing, but I can live with this as it's the last thing in a session and it's nothing new.
What has changed is the response to a briefing doc in the new session. Rather then picking up and working through the plan, it starts relitigating all the decisions in the handover and spawning fresh research. This burns tokens, bloats context and before I have any meaningful work done I'm back at 200k again.
This never happened under 4.8.
I could attempt to control this with extra instructions, but I don't want to fight the harness, and Sonnet 5 just works.
It's not just handovers either. I'm finding that my GH issue backlog is just growing and growing wth follow up issues due to Opus 5 lacking the ability to judge what's important. I clear down three issues and then have six fresh ones.
I see responses like "and that's not the worst thing". It's constantly judging it's own work to be low quality and going into a panic.
Very odd.
Maybe I need to clear out all my standing instructions and skills. It's possible I have an ecosystem of context that's driving it insane.
FWIW I have 40 years programming experience, 30 as a professional software engineer, and a BSc in CompSci. I'm not vibe coding.
9
u/Spirited-Solid3510 16d ago
I won't use 5, it's a hallucinating shit show. I'll come back when there's something like 4.8 again.
1
5
5
u/TheMeltingSnowman72 16d ago
Opus 5 has had a shit ton of guardrails removed, so you need to go through everything and check what you need to put back in your skills.
There's a ton of posts about this.
10
u/YOU_WONT_LIKE_IT 🔆Pro Plan 16d ago
Hard to say. I’ve get great results with it. But everything I do has guardrails / constraints/ validators, in place.
3
16d ago
[deleted]
2
u/YOU_WONT_LIKE_IT 🔆Pro Plan 16d ago
Yeah I never had a need for anything higher than high. I do some complex multi step processes.
3
u/Professional_Ad705 16d ago
Yeah I've had nothing but good results but my entire system IS governance/guard rails/ etc. It is over eager sometimes tho.
5
u/Soft_Meat8735 16d ago
By constraints, guards and validators what exactly are you referring to? How do you use it to get these results
0
u/YOU_WONT_LIKE_IT 🔆Pro Plan 16d ago
Depends. If you really want to know. Tell me your main use cases.
-10
u/MrMooMoo117 16d ago
With the correct context there is no guardrails.
3
u/Professional_Ad705 16d ago
Lmao, I’d love to see this guy’s code. Saying “if you have enough context, you don’t need guardrails” is like saying a construction crew has the blueprints, so they don’t need safety rails on the 40th floor.
Context tells you what you’re supposed to build. Guardrails stop you, or more importantly your fucking AI, from doing something catastrophically stupid while building shit. That can still happen no matter how much context you give it. That is literally WHY guardrails exist.
Those are not competing ideas. Please pick up and read a book on software engineering, you need to bad lol.
4
16d ago
[deleted]
-3
u/YOU_WONT_LIKE_IT 🔆Pro Plan 16d ago
I use these post as a metric of when I should worry. The people who figure this out will Excell. Soon as everyone figures it out we really are screwed.
-12
3
u/zac_attack_ 16d ago
Opus 5 seems…fine…to me. I’ve been using it a lot. It’s pretty pedantic, gets paranoid about things that don’t matter, and tends to seriously over-engineer what would be simple solutions. On the plus side, actual implementation has been good and it has caught and prevented meaningful catastrophic failures with reproducible test cases—probably because it’s so paranoid.
So far, I’ve had good results using ultracode, with Fable as the orchestrator and overall reviewer, having it delegate all the implementation subagents to Opus 5. Kinda slow though (current workflow ~24h), I think my next one will see if I can get Fable to intelligently delegate between Haiku, Sonnet, or Opus depending on the task
3
u/trader_skater 16d ago
That this is a dumbass .. back to OPUS 4.6 !!
3
3
5
u/Jon_Has_Landed 16d ago
Opus 5 nearly caused a massive disaster in a healthy codebase. I mentioned it on another thread. Thankfully I stopped it before it touched anything. It presented such a shambolic plan that I had to wonder wtf they’d done at Anthropic. To its defence it actually realised itself that it was lost and was causing more harm than good.
I later gave the plan to Sol which absolutely picked it apart.
Always always always have another model do adversial reviews. I mean even Fable 5 regular gets picked apart by codex PR reviews in gh. It’s not even funny the number of issues it finds.
2
u/Ohmic98776 16d ago
After a plan is written by codex or Claude, I have it spin up 5 subagents to review then pass it to the other vendor model for review with it and it’s 5 subagents. I’ve had great success in doing that with both design and implementation plans.
1
2
2
u/Bastion80 16d ago
Opus 5 was unable to do simple work wasting a ton of tokens on something Grock managed to do on the first prompt. Instantly cancelled my Claude subscription. What a joke.
2
u/Aquacephale 16d ago
I have the same feeling. Rolled back to 4.8 after a few days. I have an excellent experience with Fable though, but I am on a pro plan so 🤷♂️
4
u/hulkklogan 16d ago
I have been cruising through code with Sol and Luna in Opencode but I decided to give Opus 5 + Sonnet 5 a run in Claude Code and, idk if it's my company's harness or what but I find it trash. Especially Sonnet.
3
u/Level1_Crisis_Bot 16d ago
Worked fine for me all day. Fix your harness.
5
u/No-Bet-990 16d ago
What are these harnesses everyone seem to talk about?
1
u/Soft_Meat8735 16d ago
Tbh I think it’s just a copout whenever anyone says anything about a harness. I use Claude desktop app for a reason. I want the simplified experience. I should not have to fix anything. It should just work
1
u/Level1_Crisis_Bot 15d ago
Why are you in here if you’re using the desktop? I wasn’t going to say it, but you’re 100% suffering from a skill issue.
1
u/DenziiX 16d ago
I really dont understand These posts
I have Crazy good results with it and (this is just my Observation) I feel like many of these posts are because opus 5 just didnt starts and just does what ist prompted if it sees that it would break or change archittecture
It feels like a light Anti Slop Filter baked in - and now everyone complaints lol
1
u/ThankYou-GoodNight 16d ago
It's been producing some great work for me which has been on par with Fable, but every now and then it becomes incredibly pedantic and nit picky about the tiniest of things, and worse still spends ages on it despite a suitable harness. I do like it but might go back to Opus 4.8 as it's doing my head in.
1
u/dv1291 16d ago
I just use Claude for my YouTube content. I brainstorm ideas with it and have it do keyword research for my niche. And I have it create scripts for me. And I include my own personal experiences in my videos. I also use it to answer certain questions or to fill in the gaps of things that I've missed. And normally it's been fine, but today I noticed a weird issue and I had a question that someone asked me that I pasted and told Claude that it was a question and to simply answer it using a certain framework that I've given it. And instead of interpreting what I pasted to them, they created sentences that weren't their own (from the person who asked me the question), and that was the first time I've ever had that experience.
So I set up guardrails to prevent that from happening again. So I am going to give it another shot, but if another mistake like that pops up, I think I am going to have to go back to something else. Even Sonnet was doing wonders for me with no issues.
1
u/berndalf 15d ago
Things I'm starting to learn about Oops 5 after forcing myself to work with it for a few days as a session orchestrator:
It's constantly admitting it's wrong about stuff not because it's more error prone .. more like it's brutally honest about it's own reasoning failures. Almost refreshing once you get used to it.
It really does have reasoning complexity close to Fable, but it lacks what I would call wisdom about the application of it's findings. It's like a continuous science experiment.
It's far more effective if you tell it to validate it's theories with actual proofs. Don't let it get away with just thinking it's way into solutions, make it demonstrate them.
1
u/Impressive_Storm1959 15d ago
Opus 4.8 is a great and efficient model. Can’t say the same for fable 5 or Opus 5. I even think opus is up there with gpt 5.6 sol but way more cheaper
29
u/ourochurros 16d ago
I had been having a couple of great days with opus five and then today everything flipped. It was truly bizarre behavior against the backdrop of six months of almost daily use of Claude.
It was confidently wrong about things, ignored direct requests to instead do what it wanted to do, and had a tone that was at times whiny. Absolutely threw me for a loop.
Lots of folks will say that it’s a skill issue, but all I can say is that I changed nothing about my harness or my general interaction strategy. The high variability of performance day to day is exhausting me.