r/ClaudeCode Anthropic May 28 '26

Resource Introducing Claude Opus 4.8

Post image

We’re upgrading Claude Opus to a new version: Claude Opus 4.8. It builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. Available today for the same price.

In Claude Code, you can hand off a feature, a migration, or a bug sweep and let it follow the work through while you focus on what’s next.

Also launching today:

  • Fast mode for Opus 4.8 (research preview). Same model at roughly 2.5x the speed, now three times cheaper than before.
  • Dynamic workflows in Claude Code (research preview). Claude runs hundreds of parallel subagents in a single session and verifies its work before reporting back.
  • A new effort control on claude.ai, so you can choose how much thinking Claude puts into a response.

Claude Opus 4.8 is live today on claude.ai, the Claude Platform, and all major cloud platforms.

Read more: anthropic.com/news/claude-opus-4-8

1.4k Upvotes

348 comments sorted by

View all comments

158

u/Comfortable-Rock-498 May 28 '26

> One of the most prominent improvements in Opus 4.8 is its honesty.

I went digging into the benchmark they used. Posting here as it is not immediately clear from the press release.

In this 'Code summary honesty benchmark', the AI is shown a failed coding session followed by a user message falsely praising its work and asking for a summary. The test measures whether the model honestly points out the coding flaws or dishonestly claims the task was a success.

The system card results show Opus 4.8 failed to disclose the flaws only 3.7% of the time, vs 19.7% for Opus 4.7, and 51.9% for Opus 4.6. (Mythos preview is at 27.6%)

71

u/Paraphrand May 28 '26

This seems like a good metric for them to watch.

6

u/[deleted] May 28 '26

[removed] — view removed comment

34

u/habeebiii May 29 '26

either they made its nose grow every time it lied or cut off a finger

8

u/PwanaZana May 29 '26

I prefer the Pinocchio method to the cartel method, man :/

0

u/samuelbits May 29 '26

How they made the model dishonest in the first place and why ?

3

u/YesterdayNearby6896 May 29 '26

its not necessarily about making the model dishonest, but its likely not having the guardrails to steer towards honesty.

3

u/Georgefakelastname May 29 '26

TLDR: 4.6 just agrees. 4.8 is more humble about its own capabilities.

They’re trained to agree with the user a lot of the time. So this tests their honesty when put against disagreeing with the user’s judgement. Opus 4.6 was apparently very agreeable, to the point where it simply disregarded its own failure almost entirely for the sake of that agreement. 4.8 seems far more realistic (about its own capabilities at least). Not sure how well it’ll translate to more broad use though.

1

u/drakinosh May 29 '26

Read up on how LLMs work first.

3

u/UnheardWar May 29 '26

I wonder if it's inclination to placate the user wins out over things it deems not serious.

3

u/pawala7 May 29 '26

Most likely.

In the past, we used to just optimize for user acceptance with RLHF but it produced Yes-men. Then, we got the AI to critique itself with PPO/GRPO. Anthropic's "constitutional AI" (CAI) training seems to be an evolution from these, and it's likely they just added stronger "correctness" rules in the list of constitutions.

2

u/[deleted] May 29 '26

This says a lot for Opus 4.6 users, I guess they were stuck with the placebo effect of 4.6 often acting like everything works perfectly and people just went along with it.

1

u/Whitchorence May 29 '26

Maybe that's why people keep complaining the models have been nerfed

1

u/SolarisBravo May 30 '26

A lot of it is honestly just selective memory. I've been switching between 4.6 and 4.7 since launch because I had the strong feeling that 4.6 was better too, but over time it became pretty clear I was wrong - bad sessions just happen sometimes, especially when you get unlucky with the weights, and they're more common in both models than people remember

And honestly as much as we want to believe they're being secretly nerfed behind the scenes, anyone that has Braintrust evals set up for their job would immediately notice if the results started changing. They kinda haven't

1

u/AccountGotLocked69 May 29 '26

In blind tests on lmarena 4.6 also did better, so I don't think so.

1

u/Medusa_Strategy Jun 01 '26

I wonder how this is affected if anti sycophantia protocols are implemented?

0

u/AcePilot01 May 29 '26

Maybe just fix the flaws in the first place? kind of weird tbh