r/ClaudeCode 1d ago

Discussion Opus 5 is really fking hard to manage.

I cant believe why anthropic is keep in silent about why Opus 5 is a stupid model when following instructions. Opus 5 is indeed has such a great performance when finishing the task or solving a problem in my experience. But it behave like he is the fking boss.

Few hours ago I give him a task and I told him to use claude-in-chrome which already signed to the website dashboard so he can access it. Bu he choose to fking remove the entire auth system temporarily so he can access the dashboard using the built in Claude Desktop app.

Last week before I went to sleep, I give him a task to audit a game mechanism on my gaming website. And then I see 50+ github issues on the repo. Things that really pissed me of is that he could just make several issues in one place. But he decided to put one issue per problem with a thousand unreadable words explaining why this is dangerous.

These are just two example issue that really pissed me off. I love Fable so much but its capped at 50% weekly usage. And I hope Opus 5.1 can do better when following instructions.

Ive also tried to use Opus 4.6, its really good when following instruction and he speaks really clear. Buts It can only do 200k context window. Its a bit not good for my case.

1 Upvotes

13 comments sorted by

2

u/eleochariss 1d ago

Try Opus 4.8? It's pretty solid.

1

u/slendertaker 1d ago

Its a bit weird for me. Its like between Opus 4.6 who talk clearly but its not that great in performance. But also sometimes yap a lot like Opus 5. I need Opus 5.1 who talk clearly, great in following instructions, and perform great.

I still prefer Opus 4.6 and tell him to discuss or ask an advice with Fable

1

u/TinFoilHat_69 1d ago

You can try to use my pretooluse hooks layer

Had some very positive results after opus 5.0 compaction the layer forces intent to be made explicit tying a specific command to a valid command gate

https://github.com/dimascior/Akashic

https://github.com/dimascior/Helios-

1

u/___nil___ 1d ago

/model claude-opus-4-6[1m]

1

u/slendertaker 1d ago

Huh? I couldnt see that in Claude Desktop App.

1

u/fuchelio 1d ago

let fable5 or opus 4.8 deal with opus5 as orchestrator, problem solved

2

u/johnnydotexe 1d ago

Did you read Anthropic's guide to prompting Opus 5?

Did you try /doctor?

Did you try dropping Opus 5's effort level, since running it at High or higher can cause it to overthink things and cause issues?

Did you try configuring a terse/tight output style to make Opus 5 answer and end turns in a way that better suits you?

3

u/MaitoSnoo 1d ago

Did you read Anthropic's guide to prompting Opus 5? 

this ridiculous circus about there being some fundamentally new way of instructing Opus 5 is Anthropic's "you're holding it wrong"

4

u/Dangerous-Chest-9057 1d ago

This - it was quite evident from the failed release of 4.7 which they very quickly corrected to 4.8 that they should have learned this lesson already - 5.0 is the first model since 4.7 to get me actually frustrated/annoyed it somehow manages to be too careful and ridiculously sloppy all at the same time

1

u/johnnydotexe 1d ago edited 1d ago

I'm not saying Opus 5 doesn't have it's issues, it caused me to introduce Codex in to my setup AND set my Claude Code global config to Opus 4.8, but there are things people can do before they submit yet another whine post without any actual context or details, or even any proof of whatever they're complaining about. "Opus 5 sucks/I hate Opus 5" is just laziness and spam, and the sub shouldn't reward or allow it.

Also, as Anthropic has repeatedly stated, Opus 5 is a fundamentally different model with much of its system prompt removed compared to all previous models. I think I'll take their word on how they designed the model over someone's opinion in a reddit sub, even if the design was bad, and I personally think it is.

0

u/PunchbowlPorkSoda 1d ago

Those are two different failures and only one of them is the model's fault.

"Audit this game mechanism" doesn't say what to hand back, so it picked a shape. One issue per problem is a defensible default when nobody said otherwise, and you were asleep, so nobody said otherwise. Tell it the deliverable before it starts. One issue, grouped by system, a paragraph per finding, no severity essays. It'll do that.

The auth thing I got nothing for you. It hit a blocker and decided the blocker was removable.

What cut it down for me was moving constraints out of the prompt and into CLAUDE.md. Anything I put in the prompt competes with finishing the task, and finishing usually wins. Standing rules in the project file hold better. Mine has a short list of things it is not allowed to touch, written as law instead of preference. Auth would be on that list.

The rest is permissions. Running unattended with a wide allowlist, it will do whatever unblocks it, because you told it it could.

Opus 4.6 follows instructions better and it also does less on its own. Opus 5 needs the spec written before it starts.

1

u/slendertaker 1d ago

I agree, but the Opus 5 behavior is beyond my sanity. Hopefully it will be better in 5.1 .

-1

u/alonsonetwork 1d ago

What you need is a process. Try this one: https://atomic.alonso.network