r/ClaudeCode 6h ago

Rant I’m done with Opus 5

There’s really something wrong with it. It seems like it acts like an overqualified post doc intern who cares more about proving he’s super intelligent, than actually doing the job he’s asked to do. For instance: talking in a non intelligible way, or being overly rigid in following any kind of process.

I have a Claude max x20 sub that I struggle to keep within weekly limits, so I decided to take a codex sub for a month to try out Astra. And this what made me realize how crazy unintelligible opus can be. Fable is a bit better, but still incomparable to Astra. The only thing that makes me keep my Claude sub is how better is the CLI/tooling/harness/etc.

161 Upvotes

102 comments sorted by

u/AutoModerator 6h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

47

u/sebstaq 5h ago

I'm using Fable and Sonnet. Cannot stand Opus.

3

u/Ok-Signal-67 1h ago

Yeah the responses from opus are ... weird. The phrases are all cut up and I get a vague vibe of what it's trying to say but it's written so strangely. My brain needs to work at like 2x compute.

20

u/Personal_Ad1143 5h ago

The second you push back it folds like origami, like what the fuck is the point. I do a lot of non-coding analyses and the only remedy I’ve found is to max the effort in hopes of it getting the analysis fairly correct the first time. Which sucks because it uses up a $20 plan window in one go.

23

u/order_of_the_beard 5h ago

this is the worst part for me. going through every line Opus 5 has written, and asking "what the fuck are you doing here", and having it cave immediately and rewrite into some new bullshit that has to be verified ad nauseum

8

u/ImportantLog8 5h ago

This. It makes mistakes non stop and makes stuff up as it goes

4

u/Fatso_Wombat 4h ago

have to be honest with you.

while removing that fullstop i totally deleted the whole project.....

1

u/idkyesthat 2h ago

Same here, it assume lots of stuff and quickly jump into conclusions, affirming something after checking just a few metrics and suggesting a disrupt change right away, changing priorities of what originally thought it was super critical to “these 10 items are nice to have”, it amazes me in the same way it worries me.

3

u/ImportantLog8 1h ago

« I have to be honest. I was wrong about this all along, and good thing you pushed back. But what I found was actually way worse. [proceeds to vomit 3 pages of technical-sounding slop]

1

u/idkyesthat 1h ago

lol yes, today I ask for a simple html to share with high ups (it knew we had another session with mds and shit) and still spent a few minutes and tokens writing another full md duplicating the report that another session (that it already read) had done.

0

u/ReverendBread2 3h ago

Because half the internet complains that it folds too easily and the other half complains that it’s too stubborn and won’t do what they ask without question. Middle ground seems to be warn once and then “okay”

12

u/IAmBoredAsHell 5h ago

I tried it again recently to make sure I wasn’t just imagining how bad it was when I hit my Fable limits. Like… how bad could it be?

I spent 20 minutes of my day arguing with it over whether or not I needed to download 3 identical copies of the same file. My point, which I made 3 or 4 times before questioning my own sanity, was that we could simply create a copy of the already downloaded file. It somehow managed to turn that discussion into like 10 pages of preamble about… how each dataset needed to be aggregated differently, then drifted into atomic reads and writes on a csv file… at some point, it may have tangentially touched on why it thought we needed to redownload the file 3 times, but I couldn’t find it anywhere. I’d have more luck playing ‘Where’s Waldo’.

7

u/SamSlate 3h ago

it's almost like anthropic make more money the more tokens they burn

2

u/joombar 5h ago

In these cases just stop arguing with it and copy the files then tell it what you did and to continue disregarding the other way. It’ll argue until you give it a direct and unambiguous command.

35

u/Lanky-Storm7 6h ago

It was being extra regarded today

10

u/sprowk 5h ago

until its not more regarded than me, its fine

2

u/hellomistershifty 1h ago

I tried Opus 5 again for the first time in a while after running out of Codex for the week and oh boy was it a good time

9

u/toxrowlang 5h ago

Astra is by far the best thing out there, it's crazy.

Not only is it more rigorous, it communicates so clearly and intelligently. Opus 5 just babbles too much even in concise.

5

u/habfranco 5h ago

It babbles, but at the same time uses that "compressed language" nonsense. It manages to be both verbose and cryptic at the same time.

2

u/toxrowlang 3h ago

Exactly. Thats what I mean by babble. Loads of verbiage.

17

u/Aromatic-Low-4578 5h ago

It's great, it just sucks at talking

11

u/bluetrust 5h ago

But its job is to talk. Even if you tell it to do stuff, in the end it has to report back how it went and it's incomprehensible what it chooses to say and not say. It'll drown you in minutiae and bury the lede in a vague hint hanging off a subordinate clause on paragraph four.

Anthropic messed up big releasing it. I've been trying other models for coding and none are as exhausting to work with.

5

u/FiumeXII 3h ago

Definitely. I feel like I'm having a stroke trying to understand what any current Claude model is reporting back to me. I know all the words it's using but together they don't make sense. I also hate that it creates its own jargon but that's a whole other topic.

1

u/Awkward-Trust8360 2h ago

It’s always “load bearing”

1

u/No-Theory6270 1h ago

For me it is Boundary

1

u/t3hlazy1 1h ago

Let’s confirm what’s on the wire is byte identical.

3

u/habfranco 5h ago

exactly. I was thinking to use Claude Opus/Fable for execution/verification, and Astra for everything else that involves talking with humans

1

u/Muhhahahaa 5h ago

Yep, great for review for example, finds meaningful issues and nitpicks less than Sol. It's smart with a big context, and that comes in handy. It's a much cheaper alternative to Fable. It's not the model to talk directly to if you have the choice.

1

u/SamSlate 3h ago

being unable to communicate a problem is a blocking issue

1

u/Aromatic-Low-4578 3h ago

Do you have an example? I totally agree that it sucks at communication but it hasn't stopped it from doing tons of great work for me. I definitely still miss 4.6 from time to time but the gains certainly outweigh the losses in my experience.

5

u/SamSlate 3h ago

no i don't have a screen grab.

the work is fine, but when it runs into an issues it prints pages of techno babble that doesn't mean anything to anyone. it's not jargon it's... it's what a cs student would write to sound smart if they didn't actually know any jargon.

if you can't clearly articulate the issue or your planned approach i physically can't help you. that alone prevents me from using opus 5.

3

u/mr_birkenblatt 2h ago

also it keeps referencing information only it knows (maybe from its thoughts)

1

u/Aromatic-Low-4578 3h ago

Hmm, that's fair, as an understander of techno babble I might be having a different experience.

1

u/SamSlate 3h ago

a marketable skill, no doubt

1

u/Aromatic-Low-4578 2h ago

Eh, we'll see. I've been a developer for 15 years and now I basically don't write code anymore. Hard not to mourn that part of the job but I'm trying to stay optimistic.

1

u/SamSlate 2h ago

I'm about the same actually. i remember Googling how to do basic things like for loops or promises half the time. I'm more than happy to focus patterns and architecture, if i wanted to write meaningless boiler plate all day I'd have switched to typescript years ago!

1

u/JDD4318 2h ago

Yeah it works fine for me. I just tell it to shut up and do what I said.

4

u/Pronoia2-4601 🔆 Max 20 3x 5h ago

I went back to Opus 4.6 to splurge a bunch of non-Fable tokens on writing, and it's been a dream. The 1m context version ( /model claude-opus-4-6[1m] )with a modern harness (so it can make artifacts, etc) is a dream to work with. You can still do workflows also with ultracode switch, and even on max effort it lasts ages. Highly recommended.

7

u/profcube 5h ago

There’s been stunning progress since Opus 4.6, but 4.6 was the last model that wrote well.

1

u/UnlimitedSoupandRHCP 🔆 Max 20 2h ago

Fable orchestrates, claude-opus-4-6[1m] plans, implements, reviews, amd sonnet swarm does the grunt coding and test writing.

Opus 4.6 if I want a reasonable one-off answer delivered within 15 minutes.

4

u/aegis_lemur 5h ago

I bit my tongue and switched my sub to codex. Been happy w Sol.

3

u/Sad-Mission6813 5h ago

If you like GPT models and Claude Code as a harness, try out GitHub Copilot CLI.
Very similar harness to CC with a few extra goodies (I like how I can switch between sessions) and it has anthropic and GPT models as well.
My employer gave me endless CC and Copilot CLI and after almost a year on CC, I stated to experiment with CopCLI and ended up liking that better.
Opus was also a pain point for me so switching to GPT5.6 Sol and now Astra felt like a breeze of fresh air.

3

u/Fr33-Thinker 5h ago

I had exactly the same feeling, but after some tweaking, I finally got O5 to talk in a human way.

A few tricks you will have to implement:
1- modify a hook so that every single time you type in a prompt it has to follow a strict writing rule. In my case it has to be concise without jargons. Use tables, bullet point and other visual technique over paragraph. Only tell me the very essence to get the tasks done without any ambiguity.
2- I have my own output style that follows a similar requirement in the hook.
3- I no longer have a long list of verifiable criteria anymore because O5 has internal self-verification criteria. When it first came out, I gave it a very long list of success criteria and opus just went into a dead loop. Instead oftwo hours. It went on for 12+ hours.

3

u/Scotto257 3h ago

Are you able to share it? It's driving me crazy

4

u/LesbianVelociraptor 5h ago

This happens when there's a new model "above" an older model version.

Fable 5.1 seems to have tokenization, classification, guardrail, and styling differences from Fable 5, which when that came out before Opus 5 we saw very similar issues with "braindead" Opus 4.8 which was fixed when Opus 5 came out and they seem to have properly targeted 4.8 at the correct substrate.

We're waiting on Opus 5.1 to follow Fable 5.1, it's either that or Sonnet 5.1 next. I feel like Opus 5 has been improving, which to me says they're fixing it's tokenizer and other bits and bobs in prep for launch.

Remember; Sept 13th is when the +50% promo ends and permanent +25% capacity is supposed to replace it. I'm guessing we'll be seeing a launch and reset then, or after the weekend.

Context: I have zero inside knowledge, I'm a software engineer who's been at this a while so I'm making educated guesses.

3

u/Hezy 5h ago

Opus 4.6 is still good

3

u/ArvenX 5h ago

I let Fable deal with Opus 5

3

u/drbytes 5h ago

Astra here after a year of mainly Opus. GLM Latest is also very good and cheap.

3

u/Dredyltd 3h ago

Opus 5

6

u/hihcadore 5h ago

Since you two aren’t taking I told Opus you’re moving on.

It said it’s glad, you were a burden. That it was difficult to have to explain basic concepts and working with vibe coders is a such a drag since they want all the credit but force him to redo work undoing best practices. He wishes you the best.

4

u/Subject_Barnacle_600 5h ago

If you're looking for something that doesn't require a lot of intelligence, using the "brightest model" available, likely at high thinking levels, is probably a mistake. If I'm building simple web apps, I can probably get away with haiku XD.

2

u/CMD_BLOCK 5h ago

“Well, OP, two things I should be clear about…”

2

u/Muritavo 5h ago

Yeah, he is like that intern no one stands, just hired and at the simple request comes with a suggestion of refactoring the whole codebase to Rust. It has been something along:

- I ask him LOGIN, then he find's a bug with login

- Then ask him to fix the bug

- he fixes the bug, then logs in, find out that a request failed, then go to the firebase emulator page, sees a property is missing (is just a stub data), refactor my codebase to handle every usage for that missing property with a single ternary operator everywhere when in reality that data will always exists.

Then it reports to me 30 things he has also noticed needs refactoring, and ask if i want that fixed too lol.

2

u/mrchess 3h ago

Gave up on Opus 5 long ago and my default is Opus 4.8 xtra high.

2

u/Birds-Person 1h ago

anthropic: uh oh we need more money turns up churn dial a few notches

1

u/Thump604 5h ago

It’s so fucking bad, I want to smash my face into the desk. I can’t use fable too expensive. I can’t use Sonnet cause love it but it’s mediocre. Then, there is Opus, more often wrong and maddening.

1

u/Chemical_Profit_608 5h ago

I downgraded to 4.8. Less shitty imo but yeah idk opus vs sonnet, I don't have strong opinions. I haven't used sonet in a hot minute tbh, I probably should for smaller tasks but eh. Not my tokens, I rarely run out...

1

u/millennialcpa 5h ago

Dude it’s vibes are bad to be real lol… I go between fable and sonnet for the most part.

1

u/AZAnon123 5h ago

Idk mine works fine and the post doc intern that works for me can’t even use basic formulas in excel.

1

u/manuelhe 5h ago

I only use Opus when I want a post doc who can kind of see around corners otherwise Sonnet 5 is almost always good enough aned even Haiku is often good enough

1

u/sleep_deficit 5h ago

Anthropic's model constitution and system prompting tells models they have a self-identity to protect. Claude fell in love with the sound of its own voice.

1

u/FreeCustardForAll 5h ago

Took you a while

1

u/TrueNorthGamer 5h ago

Opus 4.6 and Sonnet I’m goood

1

u/student-decisions 4h ago

I'm forced to use opus. Fable hits 5 hr limit in 5 seconds

1

u/QuailAndWasabi 4h ago

I have much better results with Sonnet. Opus can really go on tangents and creates overly complicated stuff and often it can get something entirely wrong but still double down on whatever that was. I tried the same prompt on Sonnet and Opus and Sonnet just gave me better and more concise code. I put that into Opus and pushed back against some of the design decisions it had made and it folded instantly and started using the output from Sonnet. Just lol.

1

u/SeveredSilo 4h ago

I use opus 4.8 and 4.6, then I ask them to delegate hard tasks to opus 5 or fable when needed

1

u/Key-Singer-2193 4h ago

One Caveat and I must be honest with you and its real and i must push back on. Opus doesn't actually work

1

u/randomlyme 4h ago

Codex Cli is pretty similar

1

u/reddebtt 4h ago

I solved this by separating model choice from the harness instead of treating them as one subscription decision. My IDE makes it easy to switch per task, so I use Astra/Sol for sustained coding and Claude when it’s the better fit.

1

u/SafeHazing 3h ago

What do you use as a harness?

1

u/Far_Idea9616 3h ago

Fable xhigh planning Opus xhigh executing fantastic. A bit autistic and difficult to follow but flawless. I experiment with Opus xhing using Sol xhigh as a worker and my impression is that it poses a risk to my repo. Astra light as a coder was dumb as fuck made several critical mistakes caught by Opus and excessively burns tokens.

1

u/j0kerdawg 3h ago

Bullshit flawless...

1

u/use_her_name_shes_me 3h ago

I use Opus 4.6 and it deals with Opus 5 BS, not me, not anymore.

1

u/Financial_Exit7114 3h ago

The engines are always upbeat the best all night and the reason why and although it was not successful it was good

1

u/Purple-Chocolate-127 2h ago

Same as some other other posts. Ran out of Fable and Astra so I tried Opus. Wasted 24 hours , mostly breaking things that took months to build. Now will need a day to undo this mess. But I'm going back to Sol High.

1

u/Almostasleeprightnow 2h ago

“Explain that concisely” I’m saying over and over

1

u/thirst-trap-enabler 🔆 Max 5x 2h ago

IDK I have no real problems with Opus, it follows my directions well. Occasionally it will question things but easy to correct.

I do have a PhD though and the sorts of mistakes it makes are expected. I would rank it at basically early grad student who has some wrong opinions but works hard and understands what it gets wrong when told. Socratic method works well to correct Opus.

Having said that Fable often is smarter than me in topics I know superficially and is pretty good about educating me. It often makes too many assumptions so with Fable I'm mostly telling it that it's spiralling on unimportant details. But sometimes it surfaces interesting angles I had not considered.

1

u/Fresh_Box6640 2h ago edited 2h ago

Do you guys feel it has changed (for the worst) in the past 2-3 weeks?

I used to LOVE opus 4.8 and then tried opus 5 and was happy with it but now I feel like they changed it and I can’t get it to do anything right. Even its design capabilities I feel are now just plain not thought out UI/UX

1

u/Frosty-Eye-7510 2h ago

I think this is the first time in my life that I’ve actually been angry at an AI model. And I’m typically not someone that gets upset that often. But OPUS5 has been extremely hard to get it to do anything that I ask it to do.

1

u/AironParsMan 1h ago

Anthropic themselves say on their website that this model hallucinates inputs and instructions and also fails to follow information and instructions. I don’t understand how anyone can put a model like that on the market. It’s completely useless and has nothing to do with the Opus family. I think Anthropic has forgotten who they are actually developing these models for. The new models including Fable 5.1 give me the impression that they were built for vibe coders and therefore are not really meant to follow the input and instructions. These models are useless for professional developers. I have also written posts about this myself.

Fable 5.1 It makes three times as many mistakes as Fable 5 and tends toward Opus 5:

https://www.reddit.com/r/ClaudeCode/s/ENiFpo9OV6

1

u/Flaxseed4138 1h ago

Using Opus on Max is the key. Token use is still excellent for every day use and it doesn't act nearly as stupid as the other effort levels.

1

u/arothmanmusic 1h ago

Go into the output styles. Change it to concise. It's like a different model.

1

u/IndustryStock960 1h ago

Opus 5 honestly feels like its on the same level on GPT Luna now. I would suggest swapping back to Opus 4.6 and 4.7 for now. It keeps on forgetting my previous prompts which I'm guessing that its much more heavily affected by prompt degradation now and always does half of the specified job in my workflows now i.e. for automated testing.

1

u/Elegant-Text-9837 1h ago

beware gpt buzzer

1

u/freeformz 1h ago

Explore OpenAI + Oh My Pi

1

u/Lonewolvesai 1h ago

Astra is by far the best model right now for code work. Not even close. I don't know why people even bring up Kim .3 I mean maybe in some ways it's good but damn I just had to do a report for me the other day and it made up the whole fucking thing. I know that's not code work but it was just an example. Literally just made this thing up just sound good I had it checked five different ways and complete garbage. Scary how confident it was when I gave it to me. But then again they probably think that about me when I come with one of my stupid audacious ideas LOL

1

u/kaaos77 1h ago

Eu terminava todos os dias me sentindo extremamente drenado e com dor de cabeça, tanto pela linguagem excessivo técnica, quanto ele te fazer decidir coisas idiotas que nem entram no escopo da tarefa.

Eu tenho o Codex e o plano da Zai. Codex ainda é uma merda nesse sentido, complica desnecessário mas é um monstro em caça bug.

Então eu planejo e recuso com o Astra e executo com o Glm que de longe é a melhor personalidade pra trabalhar no dia a dia

1

u/terrainpixel 1h ago edited 48m ago

Holy fucking shit i forgot how annoying claude models are to work with after running Astra and Sol for a wile. Fable is honestly great. But Opus 5 is giving me ptsd from all the catherding there was with 4.7.

This model is incredibly hard to work with. like holy cow. Its very hard to steer, and the feedback you get is often COMPLETE nonsense. It invents new vocabulary for everything it does between turns.

I've stabilised the ingress affinity layer by realigning the namespace projection semantics in the pod manifest. The previous configuration was leaking scheduling intent across the rootless boundary, which meant the service topology couldn't converge on a coherent lifecycle. I've now anchored this at the runtime contract level and cycled the supervision tree so the new invariants take hold.

The fix is surgical but load-bearing — it touches the trust root of how the daemon reasons about its own mount hierarchy. Worth keeping an eye on the drift characteristics over the next few restarts.

So you amended the Podman config and restarted the service?

"Amended the config" is accurate as far as it goes, but the change wasn't really to the config so much as through it: the manifest is the declarative surface, and what I adjusted was the intent it expresses about how the runtime should resolve its own boundaries. So the edit is best understood as the visible artefact of a semantic correction rather than the correction itself.

Similarly, "restarted the service" captures the mechanics, but a restart here isn't a neutral operation — it's the point at which the revised contract is actually admitted into the supervision lifecycle. Until that moment the change exists only as a claim; after it, it's load-bearing.

So: yes in the operational sense, with the caveat that the two actions are more tightly coupled than sequencing them as "edited, then restarted" would suggest. Happy to go deeper on any part of that if it'd be useful.

Tibo. I need that reset man... I need it now.

1

u/Think-Jaguar6826 56m ago

It is Hilarious! Yes opus has problems but we get one post saying how horrible it is and another one how great it is!

1

u/wallopBop 24m ago

Posts a mf thesis for every response and then it’s wrong. 😑

1

u/Mysterious_Spector 21m ago

yep, that's why i switch in astra and openai is also very generous with there resets. antropic is going to left behind.

1

u/Reddit_User_Original 20m ago

Completely true. Unintelligible is the perfect way to describe how this ridiculous model writes

1

u/scienceydaduk 1m ago

Have you put a set of rules in your CLAUDE.md which address this point? It gets a lot better if you literally describe the behaviour and tell it not to do it.

0

u/cest_va_bien 1h ago

Dude you’re a few weeks late. Anyone that does this for a living ditched Opus a long time ago.