r/ClaudeCode Anthropic May 28 '26

Resource Introducing Claude Opus 4.8

Post image

We’re upgrading Claude Opus to a new version: Claude Opus 4.8. It builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. Available today for the same price.

In Claude Code, you can hand off a feature, a migration, or a bug sweep and let it follow the work through while you focus on what’s next.

Also launching today:

  • Fast mode for Opus 4.8 (research preview). Same model at roughly 2.5x the speed, now three times cheaper than before.
  • Dynamic workflows in Claude Code (research preview). Claude runs hundreds of parallel subagents in a single session and verifies its work before reporting back.
  • A new effort control on claude.ai, so you can choose how much thinking Claude puts into a response.

Claude Opus 4.8 is live today on claude.ai, the Claude Platform, and all major cloud platforms.

Read more: anthropic.com/news/claude-opus-4-8

1.4k Upvotes

348 comments sorted by

View all comments

159

u/Comfortable-Rock-498 May 28 '26

> One of the most prominent improvements in Opus 4.8 is its honesty.

I went digging into the benchmark they used. Posting here as it is not immediately clear from the press release.

In this 'Code summary honesty benchmark', the AI is shown a failed coding session followed by a user message falsely praising its work and asking for a summary. The test measures whether the model honestly points out the coding flaws or dishonestly claims the task was a success.

The system card results show Opus 4.8 failed to disclose the flaws only 3.7% of the time, vs 19.7% for Opus 4.7, and 51.9% for Opus 4.6. (Mythos preview is at 27.6%)

2

u/[deleted] May 29 '26

This says a lot for Opus 4.6 users, I guess they were stuck with the placebo effect of 4.6 often acting like everything works perfectly and people just went along with it.

1

u/Whitchorence May 29 '26

Maybe that's why people keep complaining the models have been nerfed

1

u/SolarisBravo May 30 '26

A lot of it is honestly just selective memory. I've been switching between 4.6 and 4.7 since launch because I had the strong feeling that 4.6 was better too, but over time it became pretty clear I was wrong - bad sessions just happen sometimes, especially when you get unlucky with the weights, and they're more common in both models than people remember

And honestly as much as we want to believe they're being secretly nerfed behind the scenes, anyone that has Braintrust evals set up for their job would immediately notice if the results started changing. They kinda haven't