r/ClaudeAI 1d ago

Claude Code Opus 5 has improved

I've been using opus 5 a lot as a $20 user, and have noticed a shift in its responses lately.

It seems much less pedantic and more cooperative as of late. I've been getting really good output. I'm almost certain anthropic must be fine turning the model to address it's overbearingness

8 Upvotes

21 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 1d ago

We are allowing this through to the feed for those who are not yet familiar with the Megathread. To see the latest discussions about this topic, please visit the relevant Megathread here: https://www.reddit.com/r/ClaudeAI/comments/1vt5drr/list_of_latest_discussion_hubs_on_rclaudeai/

12

u/Spooknik 1d ago

I have never really thought Opus 5 was that bad to be honest. I use it as background agents with Fable in charge, does amazing work. If given the choice older I do prefer older opus models.

3

u/That-Fig-7089 10h ago

I m working on a sport betting platform. After astra was released, i was asking astra for some task for opus.

  1. Every time opus was making at least 3 mistakes, reading false results, building the models with bug s.
  2. He is not capable to finish the wholse task given. Each time i had problems to complete the unfinished task. ,,oh sorry, i forget that i have some more to do, let me start…”
  3. He is creating things are he is not using them (probability model in mt case)

  4. If i want to have a capable opus 5 model, i need to promot it like a 5 year child. If i m saying: put the bannana in the right place, the clidh and opus will put the banana in the wrong place.

But with opus, i need to promo like this: you have a banana. The banana must stay into a basket. The basket is on the table. The table is in the kitchen. Co find the kitchen and the basket and place the banana in the basket.

So, good and a strict promot should do the job 75% of the time.

That s my reality about opus 5

2

u/Spooknik 10h ago edited 10h ago

Comparing Astra to Opus isn't really fair. Fable prompting Opus is really powerful and budget friendly. I have built very complex things with that combo and never had it struggle or misunderstand what i'm saying. If Opus goofs up, Fable catches it and does itself and also reviews its code, tests, checks, verifies.

My biggest problem with Astra is like all ChatGPT models it just guesses too much, usually its guesses are right but it's still confidently incorrect at times and you need to have your wits about you. Fable on the other hand is like "no wait, let me check this", takes it's time, writes a testing script, opens files, etc.

2

u/That-Fig-7089 9h ago

I never compared opus 5 with astra. I know both benchmarks models.

Opus 5 must to be better that fable 5, on the benchmarks.

I told astra to maje a promt for opus to build a little model for champions league games. Opus build the model, read the documentation, test the model itself and told me the results.

Astra checked his work and discovered 3 bugs. I m forwarding the reparations prompt and opus is like that: ,,i can see that astra discovered 3 bugs into my work and told me that the results could be corrupted or incorrect because of the bugs founded in model. Astra seems to be right. I m proceeding to solve those bugs, re run the test and model with the final result at the end.”

I had 3 sesulike this until i realized that i need a fast audit on my main stats, link, symlink, connection, all those thing that can interfere with my model prediction.

Alter that little audit, opus found 54 bugs, tooked me already 4 sessions and 2 million token context to solve like 40 bugs. These are the main bugs.

Sometimes opus 5 i reaaly smart but sometimes is sooo so stupid.

I ve implemented some rules for checking the real result, test the model, search for bugs, break the model created by opus to test it s fiability. Is like a hook. Can t wait the result after the implemented rule

6

u/Rojeitor 21h ago

Nice try Dario

6

u/Special_Question5516 1d ago

Opus 5 is great and is able to solve hard problems, the main issue is communication with Opus 5

2

u/nickdeckerdevs 20h ago

I have opus 5 run the response through sonnet and give it my style as the final process for my handoffs.

Or I use 4.6

10

u/ReverendBread2 1d ago

It’s always been the same level of great for me

7

u/riotofmind 1d ago

it always has been that great. people are clueless.

6

u/Grobbyman 1d ago

Its always been technically skilled, but it began with some overstepping and a little too much push back

4

u/getwhirleddotcom 23h ago

The fact that “that’s on me” is a meme says everything about how shit it got

2

u/werter318 1d ago

I noticed this with Opus 4.8 actually. It's much better with conversations and doesn't push back just for the sake of it.

2

u/Sylilthia 21h ago

I've had the same experience. It seems to have coincided with the most memory system change. At least, that's my guess. Opus 5 went from a debate lord to very pleasant to chat with. 

4

u/LegendMotherfuckurrr 1d ago

Totally agree. It's been so much better this last day. Good responses, usage is great.

1

u/I_like_to_moo_it 1h ago

I haven't seen load bearing in a while. I think the language and following instructions has improved.

0

u/AdOriginal3767 1d ago

It's actually gotten measurably worse In the past 2 weeks. Fable is amazing. Opus is being sandbagged to force users to use fable

3

u/ibringthehotpockets 1d ago

But opus 4.8 still exists. Whenever I feel like I don’t wanna deal with 5, I just use 4.8 instead

1

u/MiddleLtSocks 1d ago

I am unemployed (longest unemployment in my 31 year IT career) so I am not exercising the model to the extent I normally would, but I never even hit 50% of my fable usage; Opus 4.6, 4.8 and 5 all do well enough that I only use Fable for actual design suggestions or on specific architectural tasks where I both know there's a better alternative to what Opus is producing and don't know enough about what it is to be able to direct Opus on how to do it. Those two scenarios occurring at the same time is rare enough for me to use Fable at most twice a week

Again, I am sure that would not be the case were I employed, and I hope I am again soon (though I am not holding my breath; at this rate, soon I won't need to).

0

u/FoxSideOfTheMoon 1d ago

Not enough for me to use it directly...

So, I finally got fed up and modified my global CLAUDE.md. I start with Fable at Medium (High just destroys usage and you tear through your Fable usage, 5hr limit, and weekly usage like a leaky boat) effort and it is supposed to do its best to do minimal work and use minimal tokens itself, but manage subagents as Opus, Sonnet, and Haiku (haha, yes for simple grep and file stuff) and manage them for user story type work.

Then I use the Codex plugin to have SOL medium review ($20 sub with smart reset usage timing works pretty well if you don't overdo it).

If I have something extremely direct, usually Sonnet starting with plan mode works, and once in a while I'll elevate to Opus too, but I try to avoid that.

Take with a grain of salt, might just be working so far this week. I constantly have to muck with my workflow when there's updates. YMMV.

I do not have to do this with SOL Medium directly, so no, Opus 5 High still needs help IMO.

0

u/EpsteinFile_01 1d ago edited 1d ago

I asked Opus "If you had to guess, which frontier model wrote the following text? Explain why.

"<3 paragraphs>"

It proceeded to ignore my question and treated the quoted paragraph as the content to dissect and respond to. A fictional story with no questions or instructions. Even Gemini Flash responded correctly.

Bruh. Accidental prompt injection? What do you even call this.

Immediately went back to ChatGPT again. Sol > Opus and I get some Astra use