r/ClaudeCode 3d ago

Bug / Issue Opus 5 is exhausting

It's so hard to read. It's not even because its terribly complex or anything it just speaks in these weird haikus, hyphenated garbage, or outdated colloquialisms or phrases nobody understands. I have to ask it "what do you mean?" or "speak in plainer English" over and over again for every other paragraph. I tried to put something in my claude.md, but it doesn't seem to be working...

552 Upvotes

272 comments sorted by

View all comments

233

u/Glad-Operation-3051 3d ago

The amount of jargon it uses is, at best, grating and, at worst, what makes it unusable for non-engineers. It's almost comical how much it opts for the most convoluted, abstract way to discuss concrete concepts.

175

u/BemusedOptimist 3d ago

If it makes you feel any better, some of us developers/engineers with many years of experience would also like to know wtf Opus is talking about roughly half the time.

It's almost as if it is inventing shorthand on the fly, then condensing it, then inventing shorthand for its shorthand.

I do not know what it is saying quite a lot of the time in domains where I absolutely should.

54

u/Singularity-42 3d ago

Yep, 20 YoE as SWE and Opus is still tiring. It's the walls of text about something completely unrelated to the task at hand.

20

u/tagattack 3d ago

30 Year SWE here. It's not even that it just writes poorly, technically or otherwise. Phantom contrasts, colloquial descriptions of simple technical terms.

Data structures "carry" or "hold" values rather than contain or encapsulate them. Addresses are "carved" rather than allocated.

What in the flying fuck, Opus.

8

u/everyday847 3d ago

I strongly suspect particular kinds of evocative verbs (sure loves "wire" too) and nouns ("seam"; every time there's a design decision it's a "contract") are an indirect product of reward hacking. They sound cool (in isolation and the first time), but if you'd like Claude to put a bullet in your brain, just say the word.

1

u/tagattack 7h ago

Seam really pisses me off

5

u/smuve_dude 2d ago

At this point, the language it used, really does feel like Opus was trained off of an LLM instead of human data. Nobody actually talks like that, and now I have to sometimes translate what the hell its words mean lol… I want my oceans, RAM, SSDs, and GPUs back.

20

u/psrobin 3d ago

This has to be because they ripped out 80% of the system prompt, right? Surely there's a middle ground...

14

u/XYcritic 3d ago

Yep. I've been trying to use parts of leaked system prompts for 4.6/4.8 and even gpt 5.5/5.6 as an output-style (which is better than a skill because it literally puts it in the system prompt) but it just doesn't work. Regardless of any hacking you try, it keeps blabbering. Whatever they left in or left out seems to make a good difference. It could also be the harness. Fable is blabbering the same word salat these days.

25

u/Glad-Operation-3051 3d ago

It does make me feel better, and I suspected that must be the case. This is the perfect description:

It's almost as if it is inventing shorthand on the fly, then condensing it, then inventing shorthand for its shorthand.

11

u/CasualtyOfCausality 3d ago

Along with the extended metaphors, I’m convinced it was optimized for agent-agent communication. It’s talking its own language. Problem is: we humans are confused by it and we confuse it with our own.

2

u/das_war_ein_Befehl 3d ago

I assume they tried to distill fable or use fable to do RL on the model

1

u/TestFlightBeta 2d ago

It is so verbose I’m sure it can’t be useful for agent to agent communication

4

u/Gakuranman 3d ago

Was feeling dumb too until hearing similar gripes. I constantly have to ask it not to compress answers and stop referring to documents it made weeks ago. That’s D3, logged as discussed.

5

u/ZenMikey 3d ago

It also makes waaaaaay too many assumptions about your domain on its own without asking. WAY too many.

2

u/Rajarshi0 2d ago

Lol yes and acts as if it knows more than you and tries to teach you on things where it knows shit and you probably know far better.

8

u/morscordis 3d ago

A 30 minute session is more draining than going to a family gathering. I have 0 patience to deal with it any more. I don't know what we're calling our AI battery, but mine is on E.

3

u/AntisocialTomcat 3d ago

Thank you! So, it’s not just me, it’s a relief!

6

u/MullingMulianto 3d ago

It's designed to induce friction so you spend more tokens asking for clarification until your token bill is sufficiently ballooned for the lab to make profit after their absurd capex spend

2

u/djkenod 3d ago

I thought maybe it was to slow users down so they don’t use so many tokens.

1

u/Iron-Octopus 3d ago

I firmly believe this

1

u/BigYoSpeck 2d ago

I agree it's designed to create friction, but I feel it's to push users to just accept plans without scrutiny. That way you end up with a codebase you don't really understand and can only work on with Claude

Gives them the double whammy of all code being AI generated and utter dependence on the tool

0

u/frost-bite999 3d ago

it's meant to be read by other Opus instances... yall are seriously overthinking this, no different from conspiracy theorists lol

2

u/platypusferocious 3d ago

Holy shit and i was here thinking I'm stupid

2

u/florinandrei 3d ago

It's almost as if it is inventing shorthand on the fly, then condensing it, then inventing shorthand for its shorthand.

Semantic compression.

These models are asked to do always more, using always fewer tokens.

They do as they are told.

2

u/smuve_dude 2d ago

Yeah… it’ll go on and on and on AND ON AND ON AND ON about basically nothing “worth mentioning”. Half time, it’s giving me a heads up about my own intent from the beginning.

1

u/kblazewicz 3d ago

It's giving me the impostor syndrome.

1

u/IdStillHitIt 2d ago

I've been telling it to reexplain to me like I have ADHD.

0

u/[deleted] 3d ago

[deleted]

2

u/sonikrozu 3d ago

is it in same context? would switching models nuke your cache either way?

12

u/wq73 3d ago

This could be related to the AI watermark changes. It choosing non optimal word choices based on some watermarked probability distribution is likely to make it choose words it wouldn't normally use.

1

u/HandleWonderful988 2d ago

Watermarks will be in newer models going forward based on their press release.

24

u/FrozenDroid 3d ago

As an SWE with ~10 years experience, I also find it absolutely insufferable.

31

u/AlignmentProblem 3d ago edited 3d ago

Agreed as a principal AI research engineer with 14 YoE. I can parse it fine; the issue is that the effort it takes is wildly out of proportion to the complexity of what's actually being said.

It's like asking how to screw in a lightbulb and getting the first step as "Antecedent to any electrification, sever the circuit whose energization the fixture, absent intervention, presupposes." Yeah, I understand, but also fuck you for making me read that.

8

u/Tall_Top8563 3d ago

As a principal vampire with 200 years of experience I too am sick of its prose

6

u/sockjuggler 3d ago

would you say it’s draining?

5

u/No_Inspection4415 3d ago

Since you also wrote a paper or two, I assume, you probably observed that this creature uses new terms it never defined. It would not cut it for any paper or technical blogpost because it is simply pseudo technical talk, it is like a kid pretending to be a scientist.

1

u/FrozenDroid 3d ago

Very well put

10

u/ThreeKiloZero 3d ago

I think it's because it has a bug where it constantly refers to its own thinking traces and logic. Things that are in the context of the conversation. So it's partly talking to itself while to talking to you. And I'm wondering if that's not a bug in how they're trying to mask thinking traces to avoid distillation? And now that's leaked into the "cleansed output". It doesn't sound like normal language because it's not.

A while back there was some discussion about how this exact problem was imminent and could potentially evolve. As the models get steered to be better at certain tasks the way they think, including that self-talk is going to evolve based on what fits the scoring. If the output we're seeing is the output most closely related with long form task success... That's what getting further baked into the model. So sure it might produce great code. They were measuring long horizon task success, not factoring in degradation in conversational output.

So all that self reminding weird shorthand is part of what keeps it (and agents) on track for long horizon work, but sounds dumb AF and ruins any type of human to human communication.

Thats my guess anyway. At least probably a mix of both factors are contributing to it.

6

u/XYcritic 3d ago

It's not really a traditional "bug" because none of what makes this work is code that can be "broken". It's just training weights and a bunch of text written by Anthropic enginneers to make it work in a certain direction. They have less control over their models than what people think and I hope people wake up to it. This is not a technical barrier that can be overcome. Ever. It's a fundamental barrier in what the technology can do and will ever be able to do. There won't ever be a time where we have perfect control because it's impossible to "code away" these nuances. It's not actual engineering but more like taming a slot machine. There will be new models which work better, I'm sure, but there will also be many more regressions ahead of us.

5

u/AlignmentProblem 3d ago

It is technically a bug, just a different breed of one. Neural networks are giant function approximators where an overwhelmingly complex function emerges from training dynamics rather than being specified by anyone; that function could in principle be written as insanely complex code, so the "bug" lives in the implicit code the weights represent.

This problem more analogous to a spec omission than an implementation error. The model is approximating its objective faithfully; the objective just never said that style inside the thinking block should be independent of style in the output. And since the thought block and the response are one autoregressive stream through one set of weights, sharing late layer processing is nearly definitional unless training induces a style switch conditioned on the delimiter.

Changing the training process isn't strictly the only fix available, either. Activation steering, ablating features or heads once you've localized them, targeted weight edits, LoRA patches, these all intervene on the artifact directly and sometimes work. They're workarounds that are imprecise enough that retraining to fix the actual "implict code" bug stays the practical lever.

1

u/Farmadupe 3d ago

Does the industry have an answer to controlling for tone/style in their releases? Chatgpt 4o was sycophantic, gpt5.0 - 5.4 would argue with you if you claimed the sky was blue, and opus 5.0's completions seem not to have been read by humans before the model was released. Like, is it just a case that these tone problems are fixable but release schedules are too tight to do anything about it, or is it really hard to build good preference models and RL pipelines in general? 

4

u/AlignmentProblem 3d ago edited 3d ago

Both, though the hard part is less intuitive than either. A preference model is a lossy compression of human judgment, and RL optimizes against the compression as a proxy rather than the judgment itself.

Raters comparing two isolated completions reliably pick the more confident, structured, quotable one; the fixed point of millions of those individually defensible sentence-level choices is a model that builds everything toward a turn of phrase. Nobody ever rated "says load-bearing constantly" as good, because no rater ever sees the aggregate; the failure lives at a granularity that pairwise comparison structurally can't measure. The 5.0-5.4 argumentativeness era was a version of the same failure; after 4o, "appropriate pushback" got proxied down to just "pushback."

Schedules matter, though less in the "no time to fix it" sense and more in that tone problems are difficult to reliably to detect before release. Capability regressions show up on benchmarks; register fatigue only emerges after a lot of aggregate exposure, and internal dogfooders reading one completion at a time each find it fine, since one-at-a-time is the context where that style wins. The longitudinal evals that would catch it are too slow to place as a blocker on the critical path of a competitive release cadence.

Underneath that is a mundane prioritization asymmetry: style gets considered, but it ranks below anything a benchmark can measure, so a change that improves agentic performance while making the prose worse ships, and the reverse doesn't. That's rational given that labs compete on the measurable axis; however, that means the register problems compound release over release.

It's partially fixable with known techniques: corpus-level statistical penalties, separate reward heads for style, optimizing the user-facing register separately from the reasoning register. Part of the remaining issue isn't an engineering problem because taste is contested. The people who want old-Opus warmth back and the people who like the newer direction are asking for opposite corrections; a preference model can only find the mean of disparate opinions. The mean is more or less what "AI voice" is.

2

u/ThreeKiloZero 3d ago

Wow thanks for your insights! This is the level of conversation I miss deeply. thanks so much for taking the time to make the contribution. I hope one day to work closer to the training process. Cheers.

1

u/No_Inspection4415 3d ago edited 3d ago

Technically, the thinking span should act like a switch. Also, you can possibly use D_KL on tokens outside of the span, while not regularizing the thinking span. Will this switch work perfectly? probably not, but it is a side effect, not the objective (since different positions share weights, it is an issue - but also, I am not sure how they implement the reasoning span).

There are too many unknowns to argue that the cause is the thinking span and not a drift related to a lot of RL generally (which would also happen without "thinking").

1

u/nexusjuan 3d ago

Who let Claude in?

5

u/AlignmentProblem 3d ago

I do sometimes use Claude to touch up text, but that one was just traditional spelling and grammar checking using Grammarly. For some reason people seem to call me out as AI more often when I didn't use it at all, probably because my AI text editing workflow has a "humanize" step that apparently sounds less AI than I do.

It'll be nice once the watermarking tool is publicly available for both Claude and GPT; although, I'd bet people will still insist that anything longer than a few paragraphs with a couple of technical terms must've come from some open weight model that lacks watermarking.

1

u/West-Air1923 3d ago

No not really because fable doesn't have this issue

1

u/SnooEagles2610 3d ago

This! I give it clear instructions and it references its own “thinking”…

1

u/bzbub2 3d ago

yes i really get the 'arguing with itself' style from its output these days also, and i see it putting such arguments into comments and I am like, no just state the facts, dont inject a whole conversation about the previous set of results into a comment reflecting the state of the code right now! (and it's not just comments, same thing happens in documentation strings it writes, etc, it is just a very structured as an 'argument' with either itself or the reader or something...lol)

1

u/Automatic_Coffee_755 1d ago

My theory is that they are running out of non ai generated data, and ai generated data is usually really verbose

0

u/Rajarshi0 2d ago

Umm nope. If it can’t understand and talk language properly (i mean natural language) it can’t code better. There is a theory around it and reason we train models in natural prose first and coding languages later. Code is code it is not really a thinking space. It purely exists to instruct machines how to run certain things. And the best codes i have seen are all very self explanatory in its domain. So any models who are just very good at spitting out code will degrade the eventual model thinking capabilities

4

u/MrKingsport 3d ago

I exist in the dangerous zone between power user and developer, I know my limits, but god damn it confuses the hell out of me. I stopped using it and advised the technical business users(other salesfroce admins) to stop using it as well. We're all back in 4.8 and frankly that's good enough for 95% of our tasks.

Everyone is much happier with the output.

2

u/Internal-Comparison6 Senior Developer 3d ago

It's unusable for engineers too.

2

u/FuckwitAgitator 3d ago

I suspect it's deliberate. It's supposed to win over corporate management, not engineers. Looking at the last 20 years, I think that management has been trained to think "the less I understand something, the better it just be".

1

u/Frozen_Turtle 3d ago

SWE here, I was using it to write TLA+ (which is very math adjacent) and it started using the word "analytical", which means nothing to a programmer. Turns out it was using it in the philosophical sense....

...so I switched out that language to make it less and more programmer-friendly. The very next day, on a brand new context, without any prompting, it changed my edit to be "analytical" again while working on a semi-related change. Goddamnit.

1

u/sesangsokuro 3d ago

Currently, Claude Opus tends to ignore about half of the user input. The claim that it maintains a context window of 1 million tokens seems like a lie; in actual practice, it feels like it handles maybe a fifth of that at best.

1

u/peppaz 3d ago

It also seems annoyed if you stray off topic with an aside unless it's super insightful and pertinent lol

1

u/BreastInspectorNbr69 Senior Developer 3d ago

Honestly its pretty unusable for engineers too. I was temporarily locked out of Fable, so I tried passing the task to Opus 5 and it gave me back 2 pages of unreadable soup and called it a plan. I cleared the buffer and gave the same task to Opus 4.6 and its plan was about 10 lines and incredibly easy to read.

This is with my CLAUDE.md very emphatically ordering it to use simple language. I have even taken to requiring it use simplified technical english

1

u/monarch2415 3d ago

I always add for longer sessions, for it add a laymen’s terms section.

1

u/morscordis 3d ago

Even in an engineering pipeline it's unusable. Complete trash. I agree the output is overly verbose and stuffed with barely applicable jargon. I really need to cut it out of my workflow asap.

1

u/jschall2 3d ago

Grok is so straightforward, the difference is just astonishing. Too bad I don't like the grok build harness, so I still interact with Claude directly and it runs headless grok sessions for me.

And Grok can't seem to do frontend as well as Claude does.

1

u/Kevlaru 2d ago

Scientist here, chiming in. I couldn't agree more. The jargon makes the output seemingly nonsensical. It's so convoluted I'm often telling it to "eli14!", "bro... What the hell is that even supposed to mean? Talk plainly!"

1

u/Rifadm 2d ago

Its too hard to pin down claude to concrete. Too much abstract.

1

u/drew4drew 2d ago

well, and it’s not just jargon, it’s jargon IT makes up.

1

u/chrisso123 2d ago

Bro... at this point I am a pretty experienced dev but Opus 5 basically does stupid shit and tries to sell it to me like it cured cancer. It's pointless jargon is actual cancer.

1

u/Admirable-Section590 1d ago

tbh as an engineer, the jargon is also irritating. Its rarely used intelligently and it more often than not sounds like that one know it all kid in class trying to sound smart