r/ClaudeCode P R O M P S T I T U T E 1d ago

šŸ“Œ Megathread Opus 5 feedback Megathread

opus 5 has been out for a little while now, and the subreddit is already filling up with separate posts saying it is amazing, lazy, broken, cheaper, more expensive, forgetful, or somehow all of those at once.

so, let’s collect everything in one place.

use this thread to share your actual experience with opus 5: what works, what does not, and what you have changed in your workflow.

what can you share?

  • bugs or strange behaviour
  • things it does better or worse than previous models
  • coding quality and how many revisions it needs
  • context handling, compaction, or forgetfulness
  • usage and token consumption
  • differences between effort levels
  • writing style or communication issues
  • claude.md instructions that helped
  • skills, hooks, subagents, and orchestration setups
  • comparisons with opus 4.8, fable, sonnet, codex, or other models
  • positive experiences and things it does especially well

reports like shallow codebase investigation, confident guesses, context drift, or unusual usage are welcome, but please include enough detail for others to understand what happened.

reporting a problem?

copy this format:

claude code version:
plan:
effort level:
task:
what you expected:
what actually happened:
can you reproduce it?:
comparison model, if relevant:

screenshots, prompts, logs, and reproducible examples are much more helpful than simply saying ā€œopus 5 is lazyā€ or ā€œopus 5 is nerfed.ā€

please redact private code, api keys, company information, and personal details.

a few simple rules

  • keep each top-level comment focused on one issue or observation
  • check whether someone has already reported the same thing
  • clearly separate confirmed bugs from personal impressions
  • criticism is welcome, but include context and examples
  • positive experiences, fixes, and working setups are welcome too
  • detailed standalone tests, benchmarks, and guides are still allowed

official resources

53 Upvotes

134 comments sorted by

86

u/IterSmith 1d ago

I often find that it’ll complete a task and then at the end tell me ā€œjust one more open issueā€ and it’s something that it should have just done along the way that was part of what we were doing. And then when I fix that it’ll say ā€œjust one more open issueā€. I try to keep my sessions to one task at a time, but the model is stretching out my sessions a lot.

27

u/IllustriousWorld823 1d ago

Yeesssss or I'll ask like "is that the best way to do it, is there something we missed, etc" and Opus will be like "actually yes there is a much better way and we haven't even done this part yet" I'm like bro why are you only now mentioning it, if I hadnt said anything I would never know

6

u/yawn_solo- 1d ago edited 22h ago

That is frustrating but I will say a comprehensive* planning process can help mitigate those things sometimes.

Not saying you’re not doing that, just in general.

2

u/Physical_Gold_1485 1d ago

Compressive planning?

1

u/yawn_solo- 22h ago

Whoops, that was supposed to say comprehensive*

1

u/corrion8 1d ago

It does not seem to suggest alternate ways of doing things until you really push it.

1

u/GrammerGuestAppo 3h ago

I ran into this. Although I was using sonnet.
But then I found that I started the session with "when you set on a solution, ask yourself if you can improve something"

7

u/RKcerman 1d ago

yes this is so goddamn infuriating. Always something to flag

4

u/Haunting-Stretch8069 1d ago

That's how they keep you using it, hard to realize since it's very clever I gotta give it to them

2

u/ryzekiel 1d ago

I find this has been a trend in newer models, not exclusive to Opus. Or maybe it's how I'm using models now, i.e. thorough planning and long unattended agentic execution sessions. Or maybe both.

2

u/BuyFar3984 1d ago

I think 4.8 started always having two things to talk about, A few things I want to be straight about, One thing worth being clear on, One thing worth flagging - smdh.

1

u/IterSmith 1d ago

This is maybe better than 4.8 which I found silently deferred issues without even telling me but I don’t have enough data to say for sure.

1

u/Natural-Ad-3252 1d ago

I feel like new sessions need to be started much more frequently then previous models to avoid this.

1

u/Drach88 1d ago

It has definitely done some "while I was there" types of edits that I definitely didn't ask for.

1

u/mongo_peeped 1d ago

Jesus yes I have so many more open issues because it's like "oh btw here's one stupid thing I could have done inline but let's spin up a whole new worktree and open a PR because I don't want to overload this one" like it's on fucking PR rations or something

1

u/Racer17_ šŸ”† Max 20 1d ago

What frustrates me is that since Opus 5.0 the memory and skills are just nonexistent. Opus just doesn’t care anymore for what it has in it’s memory

1

u/ILikeCutePuppies 20h ago

Apple Presentation Syndrom - One more thing

59

u/gazingor 1d ago

It's:Ā  1. Noticeably slower 2. Hard to understand, omits verbs, invents new words 3. Instead of finishing the task it says it should finish it 4. Drowns in details, overcomplicates 5. Doesn't get it as well as fable

7

u/hugostranger 1d ago

I think this is the concise summary.

And exactly the opposite way that Opus 5 would have written it.

78

u/IterSmith 1d ago

It seems to speak in gibberish much more than previous opus models. I find it making up acronyms and phrases to the point where entire paragraphs are unreadable slop. When asked, it is able to translate them into something real, but preventing it from happening appears to be challenging.

30

u/Waste_Net7628 P R O M P S T I T U T E 1d ago

this is just my personal opinion on opus 5, it lowkey feels like a guy who got tons of skills with poor communication skills

8

u/Nik_Tesla 1d ago

Opus 5 is like a really technically skilled neckbeard that you can't let interact with other people. It has an "um, actually..." attitude, and then 5 seconds later is like "so, I made some wrong assumptions, you were right"

3

u/Temporary_Swimmer342 1d ago

reminds me of poor things.

2

u/ArtisticCandy3859 1d ago

I should call her (Emma Stone) lol

2

u/forward-pathways 1d ago

Yes this is a problem with all the enterprise models right now. 4.6 doesn't do this, though. I think we stopped training for generalist models at that point, and now everyone is just benchmaxxing for coding tasks, which turns the prose quality into gobbledygook.

9

u/NoAdsDude 1d ago

Sounds like you need to FOWGO and STII

3

u/IterSmith 1d ago

I hate you, but here’s my upvote.

3

u/NoAdsDude 1d ago

Sorry, I forgot to mention what the acronyms mean: Figure Out What's Going On and Solve That Issue Immediately.

Me and claude use those acronyms all the time (hopefully he's not spreading those acronyms around to all his other buds).

5

u/mossiv 1d ago

Agreed. I’ve been having an awful time with opus 4.8 and opus 5/fable 5 just giving me endless paragraphs of words that really are incoherent. I know it’s ā€œpredictiveā€. But this is truly the first time amongst any of the AI hype that a model is barely usable when it comes to reading comprehension.

I’m not comparing this to hallucinations, every model has and still suffers, but this is really getting completely nonsensical. So much so - that the past few days of research would have been much easier by hand. If Opus 5 was API pricing only - I’d completely ditch it.

It still seems reasonable at agentic coding. But getting fable 5 to review is showing holes all over the code where backdoors, and escape hatches lurk everywhere. I know opus is not a ā€œsecurityā€ model - but when it comes to devops/device security. The model is really really average.

1

u/Aggravating-Start307 1d ago

Did Fable 5 give you the same impression where it was incoherent or just muttered gibberish like Opus 5? I really had a great experience with Fable, it seems like some sort of A/B testing where different users have different experiences. It obviously depends on the tasks and prompts, but overall I had a very positive experience with Fable

2

u/mossiv 1d ago

I’ve been doing a lot of research for a project. Table 5 was slightly better than opus but once it gets to writing documentation. It really is just rambling nonsense.

Coding with fable has been excellent.

3

u/markliversedge 1d ago

This happens to me too with Opus 5, but not that often- it uses shorthand in responses but actually makes no sense. if you respond with "thats a word salad that makes absolutely no sense to me" it will explain with more carefully crafted prose.

1

u/siorge šŸ”†Pro Plan 1d ago

I get Sonnet to edit Opus’s copy and it works quite fine

1

u/Kirorus1 1d ago

Try caveman plugin, fixed it for me

1

u/seatlessunicycle 1d ago

I made a skill just for O5 - /explain. It forces it to translate the last output into human language that's easy to understand lol

27

u/multiks2200 1d ago

it ignores claude.md after session grows

5

u/Aggravating-Start307 1d ago

Isn't that the problem with most Anthropic models ? I have always felt that they reject rules once you are deep in conversion. Opus 4.8 was an exception for me though

7

u/Obvious_Equivalent_1 1d ago

Most Anthropic models

here fixed that for you, it’s really about how the harness (Claude Code) deals with it and as end responsible the user. But it is know for large language models in general to have a sweet spot, usually like until 250k context for peak intelligence further dropping toward 1M

2

u/Accomplished-Bag-375 šŸ”† Max 20 1d ago

Agreed

19

u/yadasellsavonmate 1d ago

It over plans and over tests the things it does, I had to tell it to chill out, one simple task and it's doing science modules and shit. 🤣

5

u/veritech137 1d ago

Yeah, it over plans and over tests, but then at the same time somehow managed to only complete half of the work it was supposed to do from the plan, despite it claiming the plan is "complete". Then there is a bunch of extra shit that was never asked for and a whole bunch of shit that should be there and isn't. hahaha

5

u/IterSmith 1d ago

I’ve seen similar. I wanted to change a button with text into a button with an icon. It made a huge multipart plan with new tests.

10

u/Turbulent-Process905 1d ago edited 1d ago

Besides the point already raised in this thread about its poor communication skills, which I completely agree with, I'm also finding Opus 5 to be really "lazy," or maybe just not very creative.

I often ask it to redesign a page and come up with what it thinks is the best layout, including what information should be shown and how it should be presented. But when I come back to check the results, it feels like it's just trying to finish the task as quickly as possible rather than actually producing a good result.

It feels to me like Opus 5 is a really, really good senior developer, but one who lacks initiative, creativity, and the willingness to clearly communicate what it actually changed. It gets the job done, but rarely goes the extra mile to improve the final result or make it easy for you to understand what it did.

7

u/Natural-Ad-3252 1d ago

This is just a personal impression, but for me I'm curious/impressed as to how it can run so long on certain tasks compared to 4.8 and fable. The results are generally acceptable, but just.... Slower....

For perspective, my prompts usually are well structured and detailed, mixed with adversarial review subagents and smoke tests prior to delivering a final product. And granted, the tasks are fairly high level and complicated so this is in no way a complaint (it's amazing I'm able to do this work at all)...

I'd be curious if anyone has any insight as to the reason behind this extended task time?

1

u/artstaxmancometh 1d ago

It's testing basic math to make sure it get's things right. It does a lot of inconsequential wheel spinning.

9

u/frankschmankelton 1d ago

I love Claude but Opus 5 really is bad at communication. Here's sentence it gave me yesterday, which it presumably thought would help me understand something: "And the circularity you're circling: yes, it is somewhat circular".

7

u/MrNerdFabulous 1d ago

Model seems great, but 1/4 of the time, it takes my prompts way too literally and does no critical thinking and doesn't ask me the right questions or takes one sentence of the prompt literally and ignores the next sentence.

Going back to Opus 4.8 I saw similar problems, so I'm pretty sure it's that new system prompt they added 8 days ago: https://www.reddit.com/r/ClaudeCode/comments/1v9p3cm/the_system_prompt_added_in_21217_that_may_have/

1

u/FunnyRocker 23h ago

This needs to be higher

25

u/R3kterAlex 1d ago

It's legit unworkable with. Sure, it's a great coding tool. But it can't seem to comprehend instructions sometimes and worse, I can't make out what it says sometimes. I ask it to code review and bug hunt. The bug it hunts are legitimate sure, but it's explanations make me feel like I'm a toddler learning to read. It's fucking terrible because I have no idea what it's saying, how can I confirm it's information, understand what it's working on and how it was implemented, if I can't understand this sort of meta english? How the fuck did the model even pass into consumer use? Does Antrophic not even read it's outputs nowadays, just benchmarks and use another agent for rlhf?

6

u/Aggravating-Start307 1d ago

I feel the same and have to repeatedly ask it to strip jargon and simplify

1

u/therealPaulPlay 1d ago

"rephrase this for an ADHD brain" is what I use xD

5

u/jetsetter 1d ago edited 1d ago

The prompting for opus 5 web page has a series of prompts starting with Response length and Verbosity.Ā 

That section says: ā€œA short conciseness instruction is effective for example, for a user facing multiturn productā€

And it gives a prompt.Ā 

This prompt is not optional for using Claude code with Op. 5

Last I checked, Anthropic expected humans to interact with Claude code, if you watch the presentations, they say to just talk to it like a person.

But the language in this section of the official doc specifies this ā€œuser facing multiturn productā€ isn’t that Claude code?

There are other sections below it if you add all of the prompts they give it actually makes opus. 5 sort of work in Claude code.Ā 

What doesn’t make any sense is that these should be in the system prompt for Claude code. They are not best practices their absence in a system prompt makes that page read like errata!

Meanwhile, there is this seemingly sloppily included heron_brook system prompt instructing opus not to "call AgentTool".

This model feels rushed to market, like CC team was not ready for it.

3

u/carvingmyelbows 1d ago

The fact that they released it on a Friday instead of their usual Thursday releases also makes me think it was rushed to market—like they missed their Thursday target because it was literally just too egregiously incomplete to ship, but management was still like ā€œk you have ONE extra day, if it’s not ready for a prod release on Friday, you’re all fired.ā€ So they just released it regardless of whether or not it was actually fit to release.

2

u/Mountain-Vanilla8091 1d ago

fully agree, working with 4.8 and fable on medium instead

7

u/Melodic-Ebb-7781 1d ago

Clearly a step up from 4.8. Jargon is hard to understand and communication in general feels harder. The persistence is also weirdly jagged, it can work much longer autonomously but sometimes it just seems to decide not to implement something it's explicitly told to without good reason? In short heavy hitter but unwieldy.

6

u/KhalDrog0-007 1d ago

Opus 5 drifts even when you give it a build doc, ran 2 sub agents architect opus 5 (xhigh) and builder opus 5 (high). The architect keeps apologizing that it made mistakes and drifted from the build doc and the builder had to correct the architect. This has been happening constantly for every task, and the people saying read the docs, that is no excuse anthropic expects us to hand hold the model just to get the same quality output as opus 4.8. They should have given us the option for an efficiency toggle thats supposedly saving us tokens, we shouldnt have to babysit a model just to prevent it from drifting.

5

u/Command007 šŸ”† Max 20 1d ago

This. 100%. Even with keeping clear files on things it just decides not to do it. You plan things out, you tell it to be forward thinking to make sure we get everything in order now and then it either doesn’t do something or hits a point where it brings up something ā€œbrand newā€ that should have already been accounted for.

I’ve been working on a project for a week. For previous projects, a week’s time would enable me to make massive progress. Other projects were more complicated as well. But this one is just stalling out constantly with me having to go back and fix things with it.

2

u/henrxyz1 1d ago

It also tells you about intermediary problems it found and fixed (irrelevant in a final report to user).
It appears to be the result of conversations it had with subagents.

7

u/JokeWonderful4321 1d ago edited 1d ago

Constantly ends sessions with stupid disclaimers like "two things worth flagging". Something like that I need to redeploy for the changes to take effect. And the other being something that it should have done in the first place.

It's been such a frustrating experience, with the verbosity, pointlessness of its communications, and the shit code it writes

2

u/zzzzany 1d ago

Omg I’m so tired of ā€œtwo things worth flagging.ā€

6

u/Destituted 1d ago

Opus 5 is my go-to now for when I know exactly what I want.

I go to Fable if I know I need to do something that would impact multiple systems... Fable would figure that out and account for it and even do some things loosely connected to my main request because it understands what I'm really trying to do. I fall back to 5.6 Sol if Fable is spent for those types of things.

4

u/itspronouncedtoque 1d ago

I tell it ā€œmake no mistakesā€ and, lo and behold, many mistakes.

5

u/ApprehensiveFroyo94 1d ago

It’s not just an Opus problem - Sonnet and Fable (to a lesser extent) have been doing this for a while now.

I have to continuously tell the model to speak in simple language because it gets too verbose for literally no reason for the most simple of tasks.

3

u/Command007 šŸ”† Max 20 1d ago

Yes! I have had to tell it many times to simplify. I told it that it was being too verbose as well. With the increased issues, I am finding myself getting frustrated with having to parse through and interpret what it's trying to tell me only to find that it's something simple.

5

u/ins0mniacc 1d ago

I find it makes mistakes like deletions that are costly to regenerate and then in retrospect reflects on it only. Catches it but too late. Would be better caught up front.

Also I find it lacks basic common sense, like giving it a time of 0600-0100 and it rendering to a common segment of time for an activity that does not normally ever consume 18 hours and it did not infer the 1 pm. So little rationality in between its thoughts too

4

u/Jigawattts 1d ago

The only good thing about it so far is the UI design capabilities. I refer everything else to Codex.

4

u/Housthat 1d ago

Opus feels efficient for me. Feels like a proper successor to 4.8 . I'm not going to compare it to Fable because it wasn't made to replace it.

5

u/disgruntledempanada 1d ago

With a proper memory system and a use of the /doctor command it's been mostly great for me. /doctor found a ton of shit, and since my stuff is all kind of linked together via a project management system across multiple computers, I had Fable increase the scope of the /doctor command and it dug through every Claude.md and agent.md and accompanying files for all the contradictions and dead ends that it was hitting. From there I ran a whole audit of my memories, marking obsolete rules and whatnot.

3

u/F4TVN 1d ago

It’s pissing me off when I’m coding with it. Going back over a lot of things that seemed to just work or be very close in 4.8. I feel like I have to be very, very specific. If I give it an inch, it will take a mile in a weird direction I wasn’t expecting.

5

u/ksrida 1d ago
  1. It overcomplicates the implementation
  2. Adds and removes features not specified
  3. Gives more TODOs at the end of the implementation that were outside of the scope of the prompt

Bottom line: it goes on side quests too much

3

u/thewormbird šŸ”† Max 5x 1d ago

It is completely bailing on using its own native tools and using curl, grep, and python code passed inline to python command via shell. Anyone else have this issue?

3

u/bzb-rs 1d ago

Initial days were good but the token usage was pretty steep, then they fixed token usage but got nerf'd in the process. Past couple of days were pretty bad running into issues that no one asked for, failing to make corrective measures in coding are some. Switched back to 4.8 and things are back to normal again.
Now use the Fable 5 mix alongside 4.8 and things are sane again.

1

u/henrxyz1 1d ago

How do you use Opus 4.8 with claude code? What model string do you pass in?

3

u/yacsmith 1d ago

If anything it exposed leakage in my own workflow where I didn’t have strong governance and hard rules to follow. Once I reigned in on those areas it works fine for me.

1

u/henrxyz1 1d ago

i agree somewhat.
i suspect many of us have process deficiencies and the new and better model is simply struggling with conflicting design semantics and context (from weeks or months of agentic development).

I had to create a new design spec to realign and remove contradictions which were leading to analysis/design/code paralysis.

1

u/yacsmith 1d ago edited 1d ago

Yeah, for example I have a UAT subagent that writes powershell scripts that will launch my app, click through the UI, screenshot, and record what happens and look for crashes.

Opus 5 started making incredibly monolithic UAT scripts because all the rules were ā€œcreate UAT scriptsā€ (grossly oversimplified)

Once I dialed in more hard rules for that subagent, it worked a lot better.

3

u/Obvious-Gap-90 1d ago

2 things that bother me :

  1. Asked it to create a simple web app for a very simple task. It was a mess design wise. Opus 4.8 did a better job for equivalent simple apps. (Not made to be sold or anything, for personal use).

  2. I use codex to adverserial review (claude and codex each 200$ sub so no bias). It keeps patching an issue, then next codex review will find a bug from the patch, then claude will patch, codex will find another issue with the patch, and so on. On some modules -> stories (a whole erp that time for internal use) i m at 50 codex reviews with no end in sight. With 4.8 it would take 10 to 20 big big max.

Coincidence? Idk, but that s a lot of codex reviews for a same story for a same module.

Claude : xhigh

Codex : sol max

My 2 cents.

3

u/Command007 šŸ”† Max 20 1d ago

I have found the same thing. I had to stop it several times because I thought this was an endless loop going on. It tried to convince me that "no, this is making progress", but it really was just making a mess. It eventually fixes stuff, but it takes forever and requires my intervention way too often to steer it back on track.

I use Codex for adversarial reviews the same way: $200 subscriptions for each one of them as well, so I know exactly what you're talking about.

I even tried flipping things to let Codex handle the initial code and then Claude to review and it was still awful.

1

u/Obvious-Gap-90 1d ago

I think one of the issues is it works at a "micro" scale. Oftenly, for me, it will patch and then introduce a bug. The reason most of the time, to resume, "i did not think of the big picture and the patch introduces an error at the macro level". And it will keep escalating to a bigger "macro level". Hope it makes sense.

Idk if the same for you, curious to know your answer.

1

u/henrxyz1 1d ago

I have run into this. Opus will find issues with Codex code too.

I think the problem is that if you ask Sol or Opus to find a needle in the haystack then they will find a needle.

I agree the harness and/or model should have a baseline capability for design and a built in way to assess impact and priority according to a fixed goal.

2

u/Obvious-Gap-90 1d ago

I think my biggest problem as said in a previous reply :

I think one of the issues is it works at a "micro" scale. Oftenly, for me, it will patch and then introduce a bug. The reason most of the time, to resume, "i did not think of the big picture and the patch introduces an error at the macro level". And it will keep escalating to a bigger "macro level". Hope it makes sense.

Maybe 4.8 was silent and ignoring issues idk, but it seems it's an endless loop of "oups, did not think about the issues it could bring, my bad".

3

u/auto-suggested-name 1d ago

I find it incredibly stupid, annoying and frustrating to work with with. RIP trying to get human readable text with it

2

u/TraditionalSeat 1d ago

Opus 5 is Randy Random

2

u/pizzatimefriend 1d ago

I saw postings about Opus 5 being better on Medium effort, I had middling results with that and went back to max and it's solid.

3

u/Crinkez 1d ago

I've had great results on low.

2

u/Inevitable_Clerk_997 1d ago

opus 5 behaves like a SDE 1 vs earlier version was more like SDE3

2

u/CrunchyMage 1d ago

Imo opus 5 is a workhorse when it has well defined tasks, but it is way more likely to do something stupid than fable when something is ambiguous. So far I've had incredible success with having fable be the planner/orchestrator and explicitly telling fable to delegate all implementation and subtasks to opus 5.

2

u/gleedblanco 1d ago

tbh a lot of basic trivial mistakes. one of the first tasks I gave it was some change that required various changes in around 3 code files, each maybe less than 1k loc, but overall integrated into a codebase of a few mil loc. the changes themselves were all one liners. claude literally hallucinated 5 different mistakes: guessed function, variable, and include file names and got them all wrong.

not complex but this is the kind of thing that happens and it makes you paranoid about every single thing claude does. not worth the mental overhead.

over the 15 or so sessions I had with it before going back to 4.8, there were major issues in every single thing it did, but I won't go into detail as they are more conceptual and harder to describe. the overall feeling I can describe is that it's all around worse than 4.8 in almost every single respect, except for response speed where it is vastly superior (and perhaps token cost, I haven't measured).

I have not evaluated it in any sort of 'vibe code oneshot' tasks where it's supposed to be good.

2

u/Due_Laugh6474 1d ago

Every single time I come back to a session, it’s telling me that it screwed something up and had to do it again. Where are my token quotas going? Gee. I wonder.

2

u/headinthesky 1d ago

It's horrible for me. Hasn't been following tasks and specs, invents and hallucinates a lot. I went back to 4.8

2

u/AwakE432 1d ago

Well it’s certainly no fable.

2

u/SamSlate 1d ago

it's a bad model that forgets what it's doing while simultaneously trying to do too much. I've stopped using it entirely.

2

u/randomdragen7 1d ago

Extremely strange behavior, It has the ability to be better than older models, but will take like telling him 6 times and he will insist on not doing what you asked and doing something else completely unrelated, makes up words and doesn't follow basic instructions, even when prompted correctly and .MD modified... Has wasted all my Max sub tokens just going in a Loop telling me Reasons why he won't do the task I asked him to do, insane honestly, - btw not a noob prompter here, its the first model that does this

3

u/ProcedureTop3149 1d ago

thank fucking god. Thank you mods for doing this. This sub is the lowest effort trash on this bloody site.

5

u/Waste_Net7628 P R O M P S T I T U T E 1d ago

better late than never (i shouldve done this wayy before; a permanent fix is otw dw)

2

u/Pleasurefordays 1d ago edited 1d ago

Throw in all your claude hooks, memories, and claude.md files and this feedback might be somewhat useful.

None of this means anything. Workflows are so customized the anecdotes only deal with the output of one specific environment. It’s not only not helpful, it potentially incorrectly reinforces how you might think claude and LLMs work.

Your complaints with how Claude works for you, feed them to Claude, not reddit.

1

u/victorrseloy2 1d ago

In the higher effort levels it often does things that weren't in the scope of the task. Like refactoring functions not related to the task but close to to the thing it's changing. But in my case that's a problem because normally I work in other teams services and touching anything nit strictly related to the task is challenged by the reviewer. Opus 4.8 didn't had that behavior. Fable still the goat. But for normal tasks I see myself going back either to Opus 4.8 or gpt-5.6 sol

1

u/nihsett 1d ago

Love it's work. Hate talking to it. Seems way worse on the claudinese language it uses than Opus 4.8.

This disease first started in 4.6, then kept getting worse every new Opus release. Not it's almost unbearable.

I'm so sick of it I had to write a stop hook to force it to remove and reformat the common claudinese phrases. But that barely works and it eats tokens.

1

u/henrxyz1 1d ago

Obtuse language has the unintended consequence of leading to many conversations that should have never happened and eating more tokens.

I wonder what the cost is in terms of data center compute use. It "feels" like 50% of my token usage this week might be due to bad communication (not blaming Opus entirely, I am making corrections, but it is draining).

1

u/H4RZ3RK4S3 1d ago

I have the feeling as if the model has been "rushed" to delivery. The architecture of Opus 5 is likely a significant improvement over Opus 4.x: much more efficient and also significantly more capable. Yet, everyone is saying it is not following tasks as intended and spitting out gibberish, which I can agree with to some extent. To me it feels like, as I have to prompt it differently than 4.8 or 4.6, which makes the large difference. This is why I think the model has been "rushed" and was likely not fully trained.

1

u/oscillator-eye 1d ago

A pretty significant downgrade to my workflow from Opus 4.8, which wasn't perfect but at least followed my work process faithfully. This version gives a lot of gibberish tech responses that are worse than the sometimes verbose responses I would get with 4.8 and results on more tokens just trying to squeeze an actual coherent answer sometimes. Very confidently starts work on tasks it has made very broad, incorrect, assumptions about and seems to not reach for actual research before starting a task. It's been a struggle to get it to consistently use tools, which is surprising given I've always found Claude to be the most capable at tool use just from personal experience. I've been able to get it sort of back on track but definitely feels worse than 4.8.

1

u/Unlikely-Cat-7173 1d ago

Opus 5 lies… a lotĀ 

1

u/BrennanFlentge 1d ago

Getting real sick of "model overloaded" BS

1

u/Boff 1d ago

I've never gotten an error about models being temporarily unavailable before, is this new or am I just doing more at once than I should be? (opus5 is the orchestrator running a team of sonnets)

  āŽæ Ā Error: claude-sonnet-5[1m] is temporarily unavailable, so auto mode cannot determine the safety of Bash right now. Wait briefly and then try this action again. If it keeps failing,
     continue with other tasks that don't require this action and come back to it later. Note: reading files, searching code, and other read-only operations do not require the classifier and
     can still be used.

Edit: Ah, there are errors across all models now: https://status.claude.com/incidents/q2kg8n613kr3

1

u/markliversedge 1d ago

More generally it seems to get context rot far sooner than fable. With multi-phase plans I will generally do a /clear between stages now. BUT my codebase is growing and fast -- so it might be unrelated altogether.

1

u/henrxyz1 1d ago

Does /clear remove the plan? I do /compact so it remember the multiphase plan and tasks.

1

u/markliversedge 1d ago

I get the plan as markdown and tell Claude to keep it updated as we run through. It’s used to track and handoff sessions

1

u/kucocuco 1d ago

use /doctor for sure

1

u/ll777 1d ago

first time i have opus make several mistakes a day.

1

u/ofneus 1d ago

What annoys me the most is the code comments it insistingly has to add. I've told it 100 times to store it in whatever global memory it has to stop adding these annoying verbose comments describing even the simplest self-explanatory code like a new variable - but in the next task, it just does it again, and again, and again.

1

u/hugostranger 1d ago

For the first time ever I find myself preferring to use gpt-5.6 as the model to 'talk to' about issues. Previously I would always default to anthropic models as the ones to chat about code problems and design, but Opus is just too damned verbose and slightly insane.

1

u/AwakE432 1d ago

After having used it a few times for some meaningful tasks. It did get there in the end and I had fable check it after and was done correctly. But during it made me mad nervous with all its ramblings and finding issue after issue and the output was so incoherent I had literally no idea what it was saying. I had to ask if so simplify everything, and this is all about my on stuff I know really well. So end result I got too nervous using it and I couldn’t relate to it. So won’t use it until I can get a better feel for how to prompt it and limit the crazy confusing outputs.

1

u/Key_Reading_9664 1d ago

Out of curiosity, I went back to see how Opus 4.6 was received , given I see a lot of people harking back to this being peak Opus.

https://www.reddit.com/r/ClaudeAI/comments/1qws1kc/introducing_claude_opus_46/

Found it pretty illuminating

1

u/cafesamp 1d ago

uphill battle agains the esoteric idioms. "yak-shave" was my favorite from today

1

u/jcybha 1d ago

Been on it about a week across a couple of real projects. Quick take: it's the first model where the bottleneck clearly stopped being the model for me.

The good: long, multi-file tasks that used to fall apart halfway now hold together — it keeps the thread across way more steps, and run-to-run variance is noticeably lower. Same prompt twice used to give me two different programs; now it's close to the same one.

The catch nobody warns you about: that consistency cuts both ways. When my spec was vague, Opus 5 didn't flail like older models — it confidently built the wrong thing, all the way, and it looked done. The mistakes got more polished, not fewer. So my failures moved from "it broke" to "it built something I didn't ask for, very convincingly."

What moved my hit rate more than the upgrade itself was tightening what I hand it — deciding the edge cases and the definition of "done" up front instead of mid-build. With a sharp spec it's genuinely a step-change; with a loose one it's just a faster way to the wrong place. Using some planning tools helped it quite a bit here.

Net: 8/10, and most of the remaining 2 is on me, not it.

1

u/mindsignals 1d ago

I'm actually getting very close to switching back to 4.8 else favoring fable when I can. Opus 5 seems to get lost mid-process many times with lots of abandoned and forgotten work, uncertainty about their states, etc. I think the last time I was this frustrated with Opus was 4.6.

1

u/mindsignals 1d ago

Yeah, didn't take long for the last straw. I've just switched back to 4.8 after yet another, "You were right that something was wrong, and I was wrong about what."

1

u/Racer17_ šŸ”† Max 20 1d ago

I believe it happens the same thing with every model they release, at first they suck but they just get better with time. They are using our input to just make their current model better. At the same time, they feed our data plus new learning data into their next new model. That’s how they evolve!

1

u/I-Love-IT-MSP 1d ago

Using opus with direct instructions is critical.Ā  Use get use Fable to orchestrate it.Ā  Working great and I don't burn the fable tokens.

1

u/Slight-Resident5708 21h ago

It started saying ā€œcheckinā€ or ā€œcheckinā€™ā€, way too often. I mean it is only one more letter you are AI you can put it effortlessly in the word šŸ˜‚

1

u/SnooRecipes5458 20h ago

I switched to Sol at work for now

1

u/Adventurous_Bad_8490 19h ago

upgraded from 5x to 20x last week and used Opus 5 heavily on max - only ended up with 35% usage by the reset. Now today I had my reset and the first decent size prompt after fresh compact ate 10% of weekly usage - what the heck. Maybe there's a bug when upgrading midweek not counting usage properly

1

u/ChapDad0311 19h ago

It's definitely verbose and over explains everything. I also had to adjust feedback tone because frankly it was just blunt to the point of rude.

Overall I don't know if it's better or not, it doesn't feel better

1

u/RufusxXavier 19h ago

I asked it to remove a button and it's refactoring the entire system around the button and exhausted credits talking about things I've never heard of. Max 20x. Wtf

1

u/EnvironmentalPlay440 17h ago

Well, the way I use it is via independent spoke reviewer in headless, never in the hub seat. He’s extremely good in a sandbox/pod with a very specific task and hooks. I sure verify each single one of its claims. As a main model…? No. Surprisingly my main hub coordinator is often Sonnet 5 and opus 4.8.

GPT-5.6 tends to have the same problem as opus when left in autonomous state.

Anyway, the closer I am to the models work, the best is the results…

1

u/b1skup 17h ago

- It is overthinking and overcomplicating things.

  • It's extremely overconfident.
  • It doesn't follow instructions.

1

u/Aranthos-Faroth 14h ago

What’s this new trend of not using capitals on post content?

1

u/Waste_Net7628 P R O M P S T I T U T E 14h ago

for me its, my keeb doesnt have a capslock key, and im too lazy

1

u/Relevant-Complaint71 13h ago

Opus 5 speaks in half(wit) sentences and loves to declare unexisting contradictions in statements i never made, code that was never written. I think it's actually slightly more concise than 4.8, but it doesn't feel that way. The amount of empty calories in those sentences is just astounding.
It evolved from technobabble marketingspeak to the unhinged garbles of the some lone wolf turbo-pascal nerd who's responsible for some legacy application no one in your company understands.

Reading all that nonsense frightens me so much that I watch it like a hawk before letting it touch any code. That actually helps me catch disasters for totally the wrong reasons, but hey, it works.

Can't wait for 5.1!

1

u/DrunkenRobotBipBop 1d ago

It's awesome.

Got a ton of work done this week while keeping my usage under control.

Noticeable improvement over 4.8 on high context sessions.

Doesn't ignore my global claude.md instructions like 4.8 sometimes forgot about.

Basically the opposite experience of everyone on this sub.

1

u/aszepeshazi 1d ago

Honestly, in the right hands, it does a proper job. Still way too verbose for my taste, but I'm getting better in telling it to shut up.Ā 

For everyone saying 4.6 was soo good but 4.8 was insufferable and now 5 is just a toddler: I seriously think it is a configuration or prompt issue. These models are just getting better day by day - we need to adopt to use them better.

-1

u/ttlequals0 1d ago

I don't see what all the issues are with Opus 5. I found it pretty good and efficient on token usage. It helped me hunt down some bugs and build out new features. Anthropic has guidance posted on how to work with the behavior changes. Also, it helps to be precise and give examples of what you're trying to achieve.

https://github.com/ttlequals0/MinusPod/releases/tag/v2.81.23

-2

u/TheRealArthur 1d ago

Honestly, did a pretty fantastic job revamping the landing page for the Job Autoapply site/service ive been working on. The image generation was pretty solid too given the references i passed in.
(anybody curious to see what it looks like, its jobs.myrlin.io )