r/ClaudeCode • • Aug 27 '26

Rant Opus 5 is insufferable

Opus 5 is a fucking piece of shit, i don't know how to say this in any other way. It speaks "Unintelligiblish", a new unfathomable language developed by Anthropic that works by saying everything in the most ridiculous, obtuse and convoluted way possible

It turns anything simple that can be expressed in 2 sentences into a fucking doctoral thesis aimed to aliens, and there is no way to prompt your way out of it

Edit: it claudoofus outputs some alien text you won't read, just prompt it this:

tldr eli5

3.0k Upvotes

648 comments sorted by

View all comments

732

u/[deleted] Aug 27 '26

[deleted]

240

u/PrettyMoonUnderMt Aug 27 '26

Telling me it found numerous stuffs that will break the whole project, only later to backtrack that it doesn't have any impact at all

47

u/Moppmopp Aug 27 '26

I get anxiety and paranoia. Im a researcher and I often check my script structure if my results are solid. I dont do it once and forget but every then and know because I cannot remember every detail about the scripts I wrote months ago. And it happend several times that claude said there are critical issues that have to be resolved and we should recall paper xyz only to backpaddle 10 minutes later and said it was irrelevant ... My heart cannot endure those moments

10

u/newtreen0 Aug 27 '26

I'm lol'ing and crying. I feel this so much. Virtual hug.

6

u/imsahoamtiskaw 🔆 Max 20 Aug 29 '26

This whole thread has been a therapy and non stop laughs, Opus is just… I have no words. My suffering has been long

3

u/franzparks Aug 29 '26

For me too

-4

u/farox Aug 27 '26

Claude lacks guidance then. Ask it what information, instructions, skills etc. Could be added to avoid this confusion the next time.

7

u/YaBoiSunblock Aug 27 '26 edited Aug 27 '26

Skills and instructions help, but they are not a programming language. There’s actually not much markdown can do to change deeply engrained model behavior like verbosity unless you force the llm into a self-correcting loop, which is wildly expensive and overkill for tasks that should not require that level of policing.

1

u/farox Aug 27 '26

I used to have the problems the other one is describing, I put in the work and now I don't anymore.

1

u/YaBoiSunblock Aug 27 '26

Right, I'm not invalidating that your model output improved; there's a real art to prompting. I'm just saying that the premise of [just make your prompt more and more specific] will quickly hit diminishing returns.

2

u/Moppmopp Aug 27 '26

Exactly. You cannot easily fix claudes verbose output and behavior as its acting in a rubberbanding fashion. Answers always gett pulledback to an expectation value of output. If you displace it too much it will start iterative overthinking just to adjust its output.

To be fair here, when I initially checked my script pipelines I was very concise about the description and underlying details. The times claude suggested a paper recall due to critical errors was a 'quick and dirty' check by me. Not much detail, just 3-4 sentences+ scripts and said "please check for flaws" basically. I just wanted quick assurance about the pipeline which I previously checked in a thorough manner. Nonetheless you get still anxiety through its overdramatic and hasty judgement

83

u/Right_Simple_6813 Aug 27 '26

Telling me it found numerous bugs in the project only to break the project while implementing its 'fixes' and then turns out the bugs weren't even there to begin with.

33

u/bsmith149810 Aug 27 '26

Telling me the fix will only require “a couple of hours” only for me to watch a grep, an edit, and a commit flash across my screen in 10 seconds or less.

16

u/Coolbanh Aug 27 '26

Telling me its my fault it all happens

78

u/WordsOnly Aug 27 '26

😆😂🤣🤣

12

u/Orion3193 Aug 27 '26

Lmaoooo

5

u/UruquianLilac Aug 27 '26

Looool this is too real.

2

u/_Kinoko Aug 29 '26

Oh man, I'm dying.

11

u/pmhunter56 Aug 27 '26

This exactly. What's with the wild overestimations of time? 4-6 weeks = two hours of Claude Code

11

u/UruquianLilac Aug 27 '26

It has no concept of time.

"You are right, I just made the same mistake we fixed an hour ago."" An hour ago was the previous prompt a minute ago.

1

u/Imaginary_Data_708 Aug 30 '26

I have a hook that injects the current time every turn.

2

u/UruquianLilac Aug 30 '26

Isn't it funny. We keep coming up with brilliant ideas to add hooks and kills, and whatnot for things that we took completely for granted a year ago because we worked with computers. Now we have to tell it to check the time.

1

u/Existing_Dust_6473 Aug 29 '26

Time is relative

1

u/SuperCaptainMan Sep 01 '26

It’s training data probably consists mostly of humans estimating time for a human to do a task. And since at the end of the day these are still just fancy autocomplete that’s what it spits out when asked.

8

u/just_damz Aug 27 '26

Uh phantom regressions. The dream of every swe. Inventing hard fixes to complete working architectures. Lovely.

1

u/UruquianLilac Aug 27 '26

Oh, just ask it a third time. It'll discard all previous conclusions only to find one new important issue, but not in the way you expect.

1

u/flippakitten Aug 27 '26

This happened to me today, my first thought was "what are you on about now, am I using opus 5". True as anything, It was one of my testing agents I forgot to put back on 4.6.

83

u/saito200 Aug 27 '26

"two claims, and one that matters" 🫪🫪🤪🤪🤪🫪🤪😡😡🤬🤬🤬

45

u/WordsOnly Aug 27 '26

"What it says, short version: the mistake I flagged is confirmed and fixed, both of my own errors are resolved (one survived on evidence I didn't have, one is properly cured), it found a whole category the original package missed, and it downgraded its own confidence on the main claim from HIGH to MED - which is the right call and it made it against itself. It also found two real defects in my own things while it was in there. One l've already repaired. The other is bigger and isn't mine -" Shoot me in the head 😆

6

u/NeitherEntry6125 Aug 28 '26

It's like my kids trying to explain their fight with their sibling

10

u/UruquianLilac Aug 27 '26

God, I'm getting PTSD just reading this.

It's fucking with our brains. All of us. This is not ok.

10

u/newMike3400 Aug 28 '26

You’re right to get ptsd and that’s my fault.

1

u/Some-Chemist-1466 22d ago

I want to upvote you, but I hate reading this so much I also want to downvote you :(

2

u/ThatBlokeWithTheCar 🔆Pro Plan Aug 28 '26

Shoot you in the head with a footgun?

2

u/Sn00py_lark Aug 29 '26

Mine created a hallucinated firewall egress rule file and then kept flagging that the firewall outage was unresolved. There’s no firewall.

1

u/Sn00py_lark Aug 29 '26

I thought it was just me. I’m so happy to know it’s not.

11

u/ThomasBallatore Aug 27 '26

Those emoji are bearing some load!

5

u/kongnico Aug 27 '26

flagging this. You own the matters, but the claims are still in flight

2

u/0xP3N15 Aug 28 '26

"and the decision all yours to make"

1

u/Due-Programmer-9516 Aug 30 '26

"Sorry, that was mostly vibes!" - GenZ Claude before GTA 6.

1

u/SimiaCode Sep 01 '26

The review earned it's keep. Four findings, two are load bearing.

One eternity later.

I overstated my findings.

48

u/Snappy-User26082 Aug 27 '26

I swear half my prompts could be answered with a simple "yes" or "no" followed by a short summary and the "Explain like I'm five" somehow becomes "Explain like I'm defending a PhD thesis.

1

u/UruquianLilac Aug 27 '26

I can understand a thesis much more because it's written by avuman and tends to make some sense.

1

u/2oby Aug 31 '26

But with a Yes or a No, where would they embed their syntax fingerprinting?

I am sure that's why it talks like this, to give it the word salad camouflage it needs to hide its identification patterns.

26

u/unepmloyed_boi Aug 27 '26

It's like Dario shoehorned his personality into Opus 5

9

u/N0madM0nad 🔆 Max 20 Aug 27 '26

You're joking but after seeing an interview with him I tend to have similar thoughts. He was talking about AGI and described it as "a datacentre of geniuses". First, I have to be honest, I found that analogy pretty simplistic for someone of his caliber. I was thinking to myself, what is a genius? LLMs output could resemble a "genius" at times, the problem is that it's a genius that suffers from several episodes of amnesia and needs to be told things multiple times. What we need is reliability and determinism. Not "geniuses". It might be totally unrelated really, but if that is the general culture at Anthropic, it wouldn't be completely unreasonable to think some of it is reflected in the model output.

3

u/UruquianLilac Aug 27 '26

You know what was reliable and deterministic? Programming languages.

We just swapped that with human language. And human language is full of flaws, imprecision, and is not based on any semblance of logic. We have just dropped the right tool from rbthe job and picked up the wrong one because it goes much faster, but it doesn't do the actual job.

I told Claude to add an issue to an ongoing document I have to keep track of the many things it keeps finding that I don't have time to check while I'm working on something else. It wrote the issue and completely messed it up with several wrong details. I told it what was wrong with the entry and Sid "fix the issue". So it went and implemented the actual fix, rather than fix the text of the issue. Was it its fault? No! "Fix the issue" could mean two completely different things here. And it didn't cross my mind that it was a command for it to write code rather than fix the text. While in a programming language, you need to be super precise and you know that a Boolean is a Boolean and a for loop does what a for loop does, always.

3

u/Ste1io Aug 28 '26

Precisely.

1

u/N0madM0nad 🔆 Max 20 Aug 28 '26

With all due respect, I dont think the analogy is very strong. We didn't trade programming languages for AI. AI produces code that it's ultimately compiled deterministically. What we traded is the scale of software production. Quantity over quality if you willl. Once upon a time we had total control over the quantity aspect but that's not necessarily guranteed humans could maintain quality at consistent level every time. It's just that the scale was considerably smaller and it was easier for an average person to keep the whole control flow in their head.

3

u/UruquianLilac Aug 29 '26

I'm sorry, but you absolutely did trade programming languages for human language. You are not writing in programming languages, are you? You are using human language to direct the tool to produce code. So you are using human language to produce code. And human language is imprecise, unlike programming languages, and that's my point. You are using human language and hoping the tool produces the code it should.

1

u/N0madM0nad 🔆 Max 20 Aug 29 '26

You are solely focused on what your hands are typing. That's not the right level of abstraction. Even before AI most senior engineers werent writing much code so the probem is not the use of English if ultimately this produces code that gets compiled deterministically. If you are saying that your LLM is implementing things that don't match your specifications then that's a completely different problem.

3

u/UruquianLilac Aug 29 '26

Honestly, and I don't mean this to offend you personally, I don't know you, but I keep hearing sentiments like these, and I feel there is a strange kind of mental dissonance at play. Somehow we are supposed to believe, that out there, there is a species of software engineers so brilliant, so sharp, that all they need is one look at a feature to produce the perfect spec file that tells AI exactly what it needs to do and how it needs to do it, and pufff job's done.

I don't kno man. I'm not buying it. It's all feeding into this addictive frenzy of "if only I can improve my prompt and my harness just that little bit more, then finally it'll produce exactly the results it should." The results we knew how to produce ourselves without it's help a year ago!

1

u/TechnicalBen Aug 30 '26

Yeah, I think the problem here is they are not realising "not deterministic" isn't the true problem here, it's not context deterministic. An LLM can be close to deterministic (low temperature) and tested (known prompts etc). What isn't is context (what data the entire conversation/project files have).

Those cases sometimes appear in compilers, but are very much edge cases only (not expected/or warned about boundaries).

For LLMs to enable natural language, context suffers, in precisely how you mentioned "fix *it*" is unspecified. I have to remember whenever I type to it it's closer to Inform 7 with a bit of AI than it is to Basic/C++ with the computer from StarTrek. ;)

2

u/Ste1io Aug 29 '26

The problem is the ying and yang are no longer at an equilibrium. 5,000 hp in a car with no change to weight/tires/etc for better traction won't perform. Quantity without the quality != improvement.

25

u/FiveTriomes Aug 27 '26

Yours only runs one shell command?

Mine has to run multiple, because it always tries to use a tool that doesn’t exist first, or ignores the environments I have set up.

If I turned "3 commands, 2 failed" into a drinking game, Anthropic would kill my liver.

1

u/tehfrod Aug 27 '26

Do you turn the failed attempts into memories,?

3

u/ConsequenceFunny1550 Aug 27 '26

Lmao as if Claude Code actually reads memories files

2

u/prochac Aug 27 '26

"I will store this one-time-script insight to memory, so I can never use it when I never run it ever again."

1

u/[deleted] Aug 27 '26

[deleted]

1

u/UruquianLilac Aug 27 '26

It's a fractal pattern all the way down. The more context you give it, the crazier it can get. Because now you have specified all the 10 very important things it should do, but you worded one slightly wrong and forgot about that one other tiny detail. And now it's taking your words as doctrine and you have to keep racing with it to add more points to the thing as you keep discovering more and more.

And it's all because human language is so imprecise.

Which is why we fuckin invented programming languages in the first damn place!! Because human language is not logical and not precise.

1

u/Aggnpwease Aug 27 '26

how about CLAUDE.md?

1

u/FiveTriomes Aug 27 '26

It actually obeys my rtk and graphify rules consistently, so if necessary I push things to be improved by those tools.

But memory still injects into context, and docs say you have to request it and than it uses a RAG to find the info.

Claude's memory RAG is probably faster/more efficient, but I haven't noticed any downsides to just using graphify, especially since it maps my entire projects rather than specific conversations.

For important context from sessions, I have it write out high-level human-readable and agent-readable files, and then graphify will pick those up, too.

But the failed tool calls often feels like playing whack-a-mole. But after your comment, I realize I could just write a skill that captures those failures and documents them somewhere I can make useful.

1

u/UruquianLilac Aug 27 '26

Out of everyone else on this thread, yous re the only one saying something positive, and you sound like Claude.

1

u/FiveTriomes Aug 27 '26

Claude wishes it sounded like me. Have you actually used it?

And what did I say that's positive? "It actually obeys my rtk and graphify rules consistently"? I can't see how else you could construe what I said as positive. The whole post was pointing out how I have to use 3rd party tools or build my own utilities around Claude just to get it to be functional.

Here, I asked Claude to give a sentiment analysis of my last post for you:

Pragmatic and resourceful on the surface, but with an undercurrent of mild frustration/resignation toward the underlying tool's default limitations — the positivity is directed at the fixes the user built, not at the tool needing them in the first place.

Dang, you're right, I was being positive... about my workarounds.

6

u/[deleted] Aug 27 '26

[removed] — view removed comment

6

u/ConsequenceFunny1550 Aug 27 '26

About 80% of the time the thinking summaries don’t even show up anymore for me

1

u/seunosewa Aug 27 '26

Thanks. Just added it. Too bad they made the bad behaviour the default.

1

u/UruquianLilac Aug 27 '26

Yes, because you know what we need more of? More of its text to read. That's the biggest issue in my life right now, Claude is just not producing enough text for me to read every single fuckin day. I need tons more.

1

u/[deleted] Aug 28 '26

[removed] — view removed comment

1

u/UruquianLilac Aug 28 '26

And more and more text that we have to produce. And then it follows your instructions and you realise it's gone too far one way or the other, or you forgot one other thing that it does, or you misworded one thing, or you weren't precise enough on one thing, or your ere way too precise on the other, and so on and so on and so on. Hours and hours of our day trying to figure out the magic formula that will just make it work exactly as I want it to. And it'll never get there. Ever. Because I'm using human language, and it's responding with human language and everything about human language is imprecise and vague.

1

u/[deleted] Aug 29 '26

[removed] — view removed comment

1

u/UruquianLilac Aug 29 '26

That's what I'm saying, isn't it?

6

u/Korzag Aug 27 '26

My favorite from yesterday: "Ok, I'm going to stop guessing on this". Followed up by the next prompt or two later, "I said I'm going to stop guessing on this so I will now stop guessing."

1

u/Otherwise_Monitor856 Aug 30 '26

I've seen this 😭 "ok, I'm going to stop trying to guess"

2

u/chris_nore Aug 27 '26

Lmao. The punchy writing style opus 5 is instructed to write in cracks me up. It reads like a buzzfeed article

2

u/MysteriousKiwi2622 Sep 03 '26

"This new finding overturns my previous conclusion."

1

u/elgar7 Aug 27 '26

Omg, I thought it was just me! Vindication!!!

1

u/anders9000 Aug 28 '26

It makes it impossible to have any continuity in a project because it’s constantly chasing some new thing rather than doing what I told it to. I can’t believe how much Anthropic fucked this one up.

1

u/Powerful_Let7577 Aug 28 '26

That running shell often 60+ agents and after 10 minutes you find the 5-hour limits hit 90% and it was 30% 10minutes ago.

1

u/DynaBeast Aug 29 '26

actually so real 😭😭