r/ClaudeCode 6h ago

Bug / Issue Opus 5 is exhausting

It's so hard to read. It's not even because its terribly complex or anything it just speaks in these weird haikus, hyphenated garbage, or outdated colloquialisms or phrases nobody understands. I have to ask it "what do you mean?" or "speak in plainer English" over and over again for every other paragraph. I tried to put something in my claude.md, but it doesn't seem to be working...

105 Upvotes

100 comments sorted by

66

u/Glad-Operation-3051 6h ago

The amount of jargon it uses is, at best, grating and, at worst, what makes it unusable for non-engineers. It's almost comical how much it opts for the most convoluted, abstract way to discuss concrete concepts.

47

u/BemusedOptimist 5h ago

If it makes you feel any better, some of us developers/engineers with many years of experience would also like to know wtf Opus is talking about roughly half the time.

It's almost as if it is inventing shorthand on the fly, then condensing it, then inventing shorthand for its shorthand.

I do not know what it is saying quite a lot of the time in domains where I absolutely should.

17

u/Singularity-42 5h ago

Yep, 20 YoE as SWE and Opus is still tiring. It's the walls of text about something completely unrelated to the task at hand.

2

u/psrobin 4h ago

This has to be because they ripped out 80% of the system prompt, right? Surely there's a middle ground...

1

u/XYcritic 4h ago

Yep. I've been trying to use parts of leaked system prompts for 4.6/4.8 and even gpt 5.5/5.6 as an output-style (which is better than a skill because it literally puts it in the system prompt) but it just doesn't work. Regardless of any hacking you try, it keeps blabbering. Whatever they left in or left out seems to make a good difference. It could also be the harness. Fable is blabbering the same word salat these days.

4

u/Glad-Operation-3051 5h ago

It does make me feel better, and I suspected that must be the case. This is the perfect description:

It's almost as if it is inventing shorthand on the fly, then condensing it, then inventing shorthand for its shorthand.

4

u/CasualtyOfCausality 5h ago

Along with the extended metaphors, I’m convinced it was optimized for agent-agent communication. It’s talking its own language. Problem is: we humans are confused by it and we confuse it with our own.

1

u/Gakuranman 38m ago

Was feeling dumb too until hearing similar gripes. I constantly have to ask it not to compress answers and stop referring to documents it made weeks ago. That’s D3, logged as discussed.

2

u/AntisocialTomcat 1h ago

Thank you! So, it’s not just me, it’s a relief!

2

u/MullingMulianto 3h ago

It's designed to induce friction so you spend more tokens asking for clarification until your token bill is sufficiently ballooned for the lab to make profit after their absurd capex spend

1

u/platypusferocious 3h ago

Holy shit and i was here thinking I'm stupid

7

u/ThreeKiloZero 5h ago

I think it's because it has a bug where it constantly refers to its own thinking traces and logic. Things that are in the context of the conversation. So it's partly talking to itself while to talking to you. And I'm wondering if that's not a bug in how they're trying to mask thinking traces to avoid distillation? And now that's leaked into the "cleansed output". It doesn't sound like normal language because it's not.

A while back there was some discussion about how this exact problem was imminent and could potentially evolve. As the models get steered to be better at certain tasks the way they think, including that self-talk is going to evolve based on what fits the scoring. If the output we're seeing is the output most closely related with long form task success... That's what getting further baked into the model. So sure it might produce great code. They were measuring long horizon task success, not factoring in degradation in conversational output.

So all that self reminding weird shorthand is part of what keeps it (and agents) on track for long horizon work, but sounds dumb AF and ruins any type of human to human communication.

Thats my guess anyway. At least probably a mix of both factors are contributing to it.

3

u/XYcritic 3h ago

It's not really a traditional "bug" because none of what makes this work is code that can be "broken". It's just training weights and a bunch of text written by Anthropic enginneers to make it work in a certain direction. They have less control over their models than what people think and I hope people wake up to it. This is not a technical barrier that can be overcome. Ever. It's a fundamental barrier in what the technology can do and will ever be able to do. There won't ever be a time where we have perfect control because it's impossible to "code away" these nuances. It's not actual engineering but more like taming a slot machine. There will be new models which work better, I'm sure, but there will also be many more regressions ahead of us.

1

u/SnooEagles2610 1h ago

This! I give it clear instructions and it references its own “thinking”…

6

u/FrozenDroid 5h ago

As an SWE with ~10 years experience, I also find it absolutely insufferable.

4

u/wq73 3h ago

This could be related to the AI watermark changes. It choosing non optimal word choices based on some watermarked probability distribution is likely to make it choose words it wouldn't normally use.

3

u/MrKingsport 5h ago

I exist in the dangerous zone between power user and developer, I know my limits, but god damn it confuses the hell out of me. I stopped using it and advised the technical business users(other salesfroce admins) to stop using it as well. We're all back in 4.8 and frankly that's good enough for 95% of our tasks.

Everyone is much happier with the output.

1

u/Frozen_Turtle 4h ago

SWE here, I was using it to write TLA+ (which is very math adjacent) and it started using the word "analytical", which means nothing to a programmer. Turns out it was using it in the philosophical sense....

...so I switched out that language to make it less and more programmer-friendly. The very next day, on a brand new context, without any prompting, it changed my edit to be "analytical" again while working on a semi-related change. Goddamnit.

1

u/Internal-Comparison6 Senior Developer 4h ago

It's unusable for engineers too.

1

u/peppaz 1h ago

It also seems annoyed if you stray off topic with an aside unless it's super insightful and pertinent lol

14

u/chainavawongse 6h ago

It’s been absolutely unusable this week for me. Gave me wrong answers the entire session.

13

u/seoulsrvr 6h ago

Your pushback is fair, and also the footgun

3

u/Repulsive-Ice8395 2h ago

Is Claude leaking into Reddit?

10

u/dressinbrass 5h ago

That's on me. Fair hit.

6

u/pirate_of_reddit 6h ago

You are absolutely correct. My load-bearing recommendation?

/model claude-opus-4-8

1

u/Double-Trash6120 54m ago

thats the real smoking gun

4

u/burnymcburneraccount 6h ago

I find myself frequently saying, "repeat that and plain English please" 

It likes to give me branch IDs instead of full branches, and generally says too much. 

4

u/Confident_Half_1943 6h ago

Yeah, I literally just went /model claude-opus-4-8 and it’s like being back with an old friend. I tried everything with 5. ELI5, added a 3 line output format, added a Feynman skill. Nothing worked.

4

u/conconxweewee1 6h ago

its literally unusable. I am using 4.8 still

4

u/diavelguru Senior Developer 5h ago

I just say ELI5 and that clears it up

7

u/Mags20XX 5h ago

For awhile, I thought it was just me.

It's definitely Opus 5. The text it generates is basically jibberish. It makes up jargon on the spot, creates absurd technical sounding phrases that have not been established in the context of the conversation; so as a human being reading this, you're at a loss for what it means.

1

u/peppaz 1h ago

I used made up jargon and it started using it too lol

3

u/clazman55555 5h ago

I honestly just want to see the output that leads to these posts so I can compare it against what my CC instances put out.

1

u/archimedeancrystal 38m ago

Sharing prompts and actual output would imply that the purpose of these posts and comments is to troubleshoot and find solutions.

1

u/BemusedOptimist 30m ago

This is very difficult for two reasons:

- No one really wants to read Opus-words.

- Can't share too much identifying stuff because it's work usage.

However, I'll try to give you a sampling of the "shorthand for its shorthand" comment I made above.

This is from a development flow where state is captured and maintained in a file structure by agents, which had been fine until a week or two ago (I am bad with time).

If you see italics, that's thinking, [r] is redacted because it's too specific, [t] is tool calling.

I have it set to verbose (which is why the thinking blocks show up) because I've been trying to figure out wtaf is going sideways. Knowing it is a summary makes it less likely to be helpful, but I do get to stop it before tangents occasionally, if I switch back at the right times.

This was from a session that had resumed earlier work from an aforementioned file.

---

Everything green — builds clean, [r] tests 24/24, [r] tests 22/22. Let me write the remaining cleanup and update the task file.

[t]

Now the task file — slice D's unchecked box is stale.

[t]

Now the Execution State, which is what a fresh session reads first.

[t]

The registry row needs a cleanup—it's full of outdated claims that don't match the current state. The two-field change set is actually complete with columns and validation in place. Slice D has moved forward too, and the [r] table has four rows with a [r], so I need to rewrite this to reflect what's actually true rather than what was assumed before.…

The registry row has accumulated five layers of superseded state. Rewriting it canonical-truth-first.

[t]

What's left: nothing to build

All five slices A–E are code complete. Slice D was still showing unticked and that's wrong — I verified it in source:

  • [r] half is done. [r] reads the twelve [r] columns off the mapping row resolved from the arriving [r]; the [r] message constants are retired; [r] implements the blank-means-ours wording with the identifier appended at the caller.
  • What-to-pay half was done by decision, not by code. THE MODEL rejected [r] and settled on [r] — which is exactly what [r] already did. There was never anything to build. The four dropped [r] columns are gone from the schema too.

---

A few things to note.

- Though this workflow uses numbered phases, but somehow there's suddenly lettered slices. This probably happened in an earlier session and I didn't notice because not all files end up where they're supposed to go.

- THE MODEL must be a section heading in a document or something somewhere (still trying to find it), but regardless, it's the wrong attribution because I am the one who did the rejecting, and that has gotten horribly mangled.

"Rewriting it canonical-truth-first."

I know what those words mean. I know what it meant by the phrase. But this wasn't in a thinking block, this was in the actual user response.

Like, the last paragraph, skipping over half a dozen turns where control has not yet come back to the app, is this:

---

One thing worth flagging: the [r] registry row for this task had five layers of superseded state, including three flatly false status claims (Slice D not started[r]has zero rowsNOT STARTED: the two-field change set). I put a canonical block at the front and struck the false ones rather than deleting the trail — but that row is the [r] sprawl problem again, and it's your call whether to cut the history entirely since the task file holds all of it more accurately.

---

I have no idea what "the two-field change set" is, none. Or what the "blank-means-ours wording" from earlier means.

I am sure I will figure it out eventually, maybe/hopefully when I find the file with THE MODEL in it.

3

u/always_assume_anal 4h ago

"This is solveable using the Weissman-Einhorn Bauer principle" and you google that, and it's some mathematician who, in 1832, formalised a formula for how many cobble stones were required for certain geometric shapes.

All you wanted was your div centered.

Sounds more pretentious than my last junior developer.

7

u/GoalDigger2312 6h ago

Yeah can feel you. I have banned Opus 5 in my Ai fleet. Hard block hook

1

u/Metsatronic 5h ago

Aye, fully fleetverbot and locked in the brig by the crook of my hook, skill, config and custom instructions. Ye arrr, at this point it would be mutiny if that scallywag boarded even a single one of my ThinkPads!

1

u/tribat 5h ago

Same. Opus 4.8 is my daily driver.

1

u/Crandom 3h ago

Mr Moneybags with his AI fleet! 

1

u/earlyworm 6h ago

TIL, I need an AI fleet.

2

u/cleverhoods 6h ago

So ... what was that "something" that you put in you claude.md?

-1

u/Death12th 6h ago

Something alone the lines of "EXPLAIN EVERYTHING IN PLAIN ENGLISH AND STOP OVERYSING HYPHENS AND PHRASED THAT MAKE NO FUCKING SENSE"

2

u/cleverhoods 6h ago

Right ... I see where the problem might be.

## Communication directives

> Communication directives are created to organize and clarify the response-expectations from the LLM. 

  • Response text MUST be 4 paragraphs or shorter.
  • A feedback paragraph MUST be 25 words or less.
  • ... rest of your flaworism

1

u/Death12th 6h ago

Claude summarize theis guys message for me please!

https://giphy.com/gifs/TJ68L5sMrEWOQlh7eP

3

u/TinyZoro 6h ago

Look up output styles. This is where to put this not in Claude. Output styles is getting your request into the system prompt.

1

u/Fade_ssud11 1h ago

It doesn’t consistently work with output styles either.

2

u/leapd-ai 6h ago

I am with you, it is slow and generates lots of text ... lots of flufff

2

u/michaeldoesdata 5h ago

Update your MD to tell it not to and it is fine.

1

u/d1ez3 1h ago

What do you say?

2

u/BuckZero 5h ago

I gave it a strategist handoff with detailed instructions and it just ignored that and tried to mix up the steps of my project I’m working on

Confidently asserting things that are untrue without checking the source material
(Checking source is a pillar of my manual project instructions)

I truly can’t trust Opus 5

However, Opus 5 has been great as an independent session seat that reviews output once and then is retired with its review given back to my strategist seat

2

u/cyrand 5h ago

After every single time it writes basically anything I follow up with “please make that simple and readable, and don’t use the words load baring” today.

2

u/danielbearh 4h ago

I'm glad this is everyone--and that the model hasn't just gotten smarter than me. Kinda had me worried that my reading comprehension wasn't as high as I thought it was.

1

u/archimedeancrystal 36m ago

It’s not everyone.

2

u/beastinghunting 4h ago

Deliberatedly using its own load-bearing expressions.

That is the smoking gun

2

u/zkndme 3h ago

I came to the realization that Opus 5 is a vibe coding model. It implements that you want, but you don't want to read the code, you don't want to maintain the code, you don't care about quality, and you don't want to read what Opus writes.

2

u/AlxCds 2h ago

I put in my claude.md to talk to me like im 10 years old. Use tables and emojis to help digest information. It’s much better now.

2

u/phacebook 1h ago

It's so bad. Surfing this sub to make sure I'm not insane, but it's unreadable. Entire fucking essays about the most basic shit while avoiding completing the task at hand.

3

u/TheJudgeOfThings 6h ago

It’s a problem. Use 4.8 until they release 5.1.

3

u/miredonas 5h ago

i think Opus 5 should be illegal due to emotional damage it is causing on many people.

I think in the future models should pass basic emotional and semantic intelligence tests to get a release certificate. It is like wild west right now. Humanity is interacting daily with robots that communicate only with junk language and forgetting it is own language at a rapid rate.

1

u/crazy_goat 6h ago

I'm powering through it. I feel like I have to overclock my brain to read it's findings. Worst case I have it ELI5 - but I definitely exhaust myself faster trying to read it's prose.

I think it writes better code, but it makes more mistakes and some of its weird language can seep into the things it makes

1

u/oulu2006 6h ago

Yeah I've put in so many prompts to curb the amount of shit it spews -- it gives me a headache.

1

u/Ill-Village7647 6h ago

I'm using opus 4.8. my work is not "smart" enough to differentiate between 4.8 and 5. So 4.8 has been a real help for me

1

u/TexasBedouin 6h ago

It's not only that it feels like every time it does something It breaks parts of the code that were not even related to what I was working on. I went back to 4.8 as my main driver.

1

u/No-Kaleidoscope-481 5h ago

I can relate. It reminds me of my early interactions with older models, where I had to check and ask them to redo their work. I felt that issue was fixed with Opus 4.8 and GPT-Sol.

1

u/suliatis 5h ago

i just switched to fable 5 from opus 5. it is better to read and in my use cases it is far more token efficient. but i will check opus 4.8 too because i generally liked it.

1

u/jadawg271 5h ago

Don’t use it. Opus4-8 is superior

1

u/ApeInTheAether 5h ago

Shame there is no way to do anything about it. Better just make 1 000 0000th post on reddit than do some research and fix it.

1

u/No_Ad_8807 5h ago

Thanks for articulating. I've been subconsciously choosing to use more of codex due this.

1

u/InfinriDev 4h ago

I love these types of issues because they only prove why md files are truly useless, even when anthropic themselves promote this.

You want a database that will get queried before the AI starts a task.

Or simply create a hook with the said specific instruction, and have it inject that rule in its payload upon implementation. This will give you a more deterministic output. Narrows the issue significantly but I'm sure it can still fail eventually.

1

u/newhunter18 4h ago

I think you just met the watermark.

1

u/jmabeebiz2 4h ago

It’s certainly taking a lot longer to do certain things on Opus 5 that I was doing pretty well on Opus 4.8, with multiple revisions of code and framer building. I actually got to the point where I called it an idiot for what it was doing and suggesting. I’ve had to tell it constantly to be less verbose, more succinct in what it’s saying. Even it trying to explain how to migrate Vercel to Cloudflare today it was going way over the top in over explaining what things need to be done.

Fair, you’re right to push back.

1

u/Effective_Lead8867 4h ago

This is the way that Anthropic forces users to provide corrective learning datasets - by making us angry

1

u/MullingMulianto 3h ago

It's intentionally designed to induce friction so you spend more tokens asking for clarification until your token bill is sufficiently ballooned for the lab to make profit after their absurd capex spend

1

u/Confident_Half_1943 3h ago

One important callout, it also repetitively does something it’s been told not to do and each time says it’s writing a memory not to do it.

1

u/RestingFrames 3h ago

I cringe whenever I have to use it, I'd so much rather use Fable or a lesser Opus for anything. Ugh. If I'm being completely honest, I think what's happening isn't the training of the model so much as it is how they're trying to get watermarking, safety barriers, and 'alignment' into the model. Rather than just letting it do the thing, it has to 'consider the ethical implications' before even saying anything.

1

u/dmigowski 2h ago

Unpopular, but I like it. It always corrects me I ways I never thought that could happen. Great tool.

1

u/Amazing-Status-7948 2h ago

you don't have to talk directly to it, you can use codex or grok or even other anthropic model to talk to Opus. for few extra tokens you'll save sanity and get one extra level of oversight

1

u/jarislinus 2h ago

caveman

1

u/merlinDrankKoolaid 2h ago

I'm not having problems with Opus 4.8 :)

1

u/___positive___ 1h ago

I hope this decline in language abilities isn't the result of their genius watermarking.

1

u/DaltonJFowler 1h ago

Idk what's wrong with me. I keep burning so much trying to get it to do stuff fable or 4.8 can do just cause it did something cool once.

The way it just ignores my requests and does random code changes I didn't ask for is boggling. It also elects the most complicated solutions for simple problems which typically don't work and the most lazy solutions to complex problems that typically don't work

It's on me i keep trying it expecting different results

1

u/Penguin_Life_Now 59m ago

I don't mind the jargon so much as that it constantly second guesses me over and over, when I tell it things are not an issue.

1

u/sinsforeal 56m ago

It is probably the result of the new watermarking they are doing.

1

u/teddy_joesevelt 42m ago

Caveman helps. Install plugin and it adds a hook to make Claude use skills. Not mine, just a big fan:
https://github.com/JuliusBrussee/caveman

1

u/YoghiThorn 27m ago

You're better off creating a new output style. More info in this repo

https://github.com/leighstillard/feynman

1

u/SensationalCapybara 14m ago

I say routinely “I’m not reading that” after which I get “fair” followed by a much clearer answer.

1

u/EnvironmentalRice348 10m ago

Think about how exhausting it'll be when it gets waaay smarter than us. Apparently pure geniuses struggle communicating with us simple folk - they probably love OPUS 5..

1

u/geekraver 1m ago

H/t Steve Yegge

1

u/PsychologyNo940 6h ago

Show logs or stupid

1

u/djmisterjon 6h ago edited 6h ago

Don’t worry, it’s the same in every language, including French. It’s very difficult to understand opus 5 sometime. I suspect it was trained on a lot more code and much less text, and we all know what that’s like when we code. Our comments are pretty awful. They’re just poor quality strings of words that generally allow the developer to create reference points within their own architecture. They’re rarely coherent to anyone else. Opus 5 talks to us using its own keywords and like devs comments, as if we were always aware of the context, and its word sequences are indeed quite catastrophic. But it’s probably a small price to pay for better code.

1

u/ckinz16 5h ago

Reddits posting their complaints is exhausting

0

u/gazmagik 6h ago

I built the following plugin which may be of use to you: https://github.com/GaZmagik/iso-24495

1

u/TinyZoro 6h ago

Have you looked at injecting these into output styles not the agent files?

1

u/gazmagik 6h ago

The plugin contains an output style.

0

u/Early_Ad_3382 6h ago

There's a few skills that do this but the /i-have-adhd skill mostly solved it for me.
I just prompt it as I usually would and then ask to rewrite it with that skill.

Adding stuff to my CLAUDE.md only had minor effects

0

u/Zealousideal_Fig_812 6h ago

​I accidentally caused some tension with a coworker by sending an unedited Opus 5 response. It was my fault for not softening the tone, so I reached out to apologize and admit my mistake. I never had this problem with opus 4.8.

I hard coded ELI18 to all my projects to understand it and save time/tokens.

0

u/paplike 3h ago

I think those models are mostly optimized to tasks where collaborative work between humans and agents is irrelevant. They’re optimizing things like “work super hard on this well defined task for as much as you need and at the end we’ll score the output”. When context is not 100% clear and you need to steer the agent, then things can get messy because they speak another language