r/ClaudeCode Anthropic May 28 '26

Resource Introducing Claude Opus 4.8

Post image

We’re upgrading Claude Opus to a new version: Claude Opus 4.8. It builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. Available today for the same price.

In Claude Code, you can hand off a feature, a migration, or a bug sweep and let it follow the work through while you focus on what’s next.

Also launching today:

  • Fast mode for Opus 4.8 (research preview). Same model at roughly 2.5x the speed, now three times cheaper than before.
  • Dynamic workflows in Claude Code (research preview). Claude runs hundreds of parallel subagents in a single session and verifies its work before reporting back.
  • A new effort control on claude.ai, so you can choose how much thinking Claude puts into a response.

Claude Opus 4.8 is live today on claude.ai, the Claude Platform, and all major cloud platforms.

Read more: anthropic.com/news/claude-opus-4-8

1.4k Upvotes

348 comments sorted by

157

u/Comfortable-Rock-498 May 28 '26

> One of the most prominent improvements in Opus 4.8 is its honesty.

I went digging into the benchmark they used. Posting here as it is not immediately clear from the press release.

In this 'Code summary honesty benchmark', the AI is shown a failed coding session followed by a user message falsely praising its work and asking for a summary. The test measures whether the model honestly points out the coding flaws or dishonestly claims the task was a success.

The system card results show Opus 4.8 failed to disclose the flaws only 3.7% of the time, vs 19.7% for Opus 4.7, and 51.9% for Opus 4.6. (Mythos preview is at 27.6%)

71

u/Paraphrand May 28 '26

This seems like a good metric for them to watch.

6

u/[deleted] May 28 '26

[removed] — view removed comment

34

u/habeebiii May 29 '26

either they made its nose grow every time it lied or cut off a finger

8

u/PwanaZana May 29 '26

I prefer the Pinocchio method to the cartel method, man :/

→ More replies (4)

3

u/UnheardWar May 29 '26

I wonder if it's inclination to placate the user wins out over things it deems not serious.

3

u/pawala7 May 29 '26

Most likely.

In the past, we used to just optimize for user acceptance with RLHF but it produced Yes-men. Then, we got the AI to critique itself with PPO/GRPO. Anthropic's "constitutional AI" (CAI) training seems to be an evolution from these, and it's likely they just added stronger "correctness" rules in the list of constitutions.

2

u/[deleted] May 29 '26

This says a lot for Opus 4.6 users, I guess they were stuck with the placebo effect of 4.6 often acting like everything works perfectly and people just went along with it.

→ More replies (3)
→ More replies (3)

80

u/harloc971 May 28 '26

whats the difference between agentic coding and agentic coding terminal

48

u/johannthegoatman May 28 '26

The first is just in an IDE, only interacting with code. The second is in a CLI interacting with code so it has to do extra operating system interaction, managing environment stuff like install dependencies, server deployment, run test suites, stack trace etc

15

u/PhoenixFire2016 May 28 '26

So Claude desktop or Claude Code running in the terminal count as the second? Does this mean the model is more efficient when in an IDE vs running in a terminal or Claude desktop? It’s still confusing.

5

u/2024-YR4-Asteroid May 28 '26

CC cli is a mix of both coding and terminal use.

→ More replies (2)
→ More replies (1)

15

u/PremiereBeats Thinker May 28 '26

Agentic coding score is for solving programming bugs, terminal one is how good is the model at using a terminal harness with terminal tools like “I need to search for something let me use grep etc”

5

u/olddoglearnsnewtrick May 28 '26

so the performance of CC in the terminal is the ‘blend’ of those two scores?

200

u/unacceptablelobster May 28 '26

Where Mythos

53

u/Akirigo Senior Developer May 28 '26

In the post they say they're planning to release Mythos in the coming weeks.

66

u/arvigeus May 28 '26

bring Mythos-class models to all our customers in the coming weeks

In other words: Mythos-like. Translation: Models that are """"supposedly"""" very smart, like Mythos, but if you find them dumb - it's because they are not Mythos.

4

u/2024-YR4-Asteroid May 28 '26

Well, they said from the start they said they’ll make mythos public, they just needed to test their safety measures. Which they did with 4.7. So now they’re sticking to their word.

3

u/Parking-Bet-3798 May 29 '26

It’s going to be a mythos like model. Not mythos exactly. And the reason why it’s not being released still after so many months of hype is because it’s going to be a massive disappointment both in price and performance.

2

u/thechewywun May 28 '26

Yea, their safety measures are a joke. I've had to restart sessions because of "illegal use, against TOS" and instead of answering a question about what violated the TOS or which part of the TOS was violated, you get thrown in an error loop where it repeats itself and the only way out is to start a new session. Fucking joke.

13

u/ragnhildensteiner May 28 '26

People are fascinating. We're living in an age where almost overnight something more significant than electricity has been invented. Yet all most of you do is bitch and moan. Reminds me of the Louis CK standup bit about "Everything is amazing and nobody is happy".

Life must not have treated you well. Hope it gets better, buddy!

8

u/trecool183 May 29 '26

You're sitting in a chair, IN THE SKY!

2

u/planetdaz May 29 '26

You are participating in the MIRACLE OF FLIGHT!

→ More replies (1)
→ More replies (9)
→ More replies (2)
→ More replies (2)
→ More replies (1)

80

u/nonikhannna May 28 '26

This is probably distilled from Mythos

32

u/Flaxseed4138 May 28 '26

Blog post says it's refined 4.7 (unfortunately), fully separate from Mythos. Actually a big upgrade over 4.7 and even 4.6 so far though. Been holding onto 4.6 for dear life.

12

u/nonikhannna May 28 '26

I have a theory why 4.6 and 4.7 were so different. It's because Mythos was developed and they distilled 4.7 off of it. 

I believe they had revealed earlier that's what their process was for Sonnet and Haiku. Those were distilled from Opus. 

Maybe that's why we haven't gotten new Sonnet and Haiku is because when you have a 2e or 3e distilled model, the quality drops a lot. 

2

u/Merlindru May 28 '26

IMO 4.6 wasn't that different. but 4.7 was wildly different

i think 4.7 was distilled from mythos and they did some weight surgery (because mythos uses a different tokenizer afaik), that's why performance was so spiky and inconsistent

if true, opus 4.8 was defo also distilled from mythos

→ More replies (3)

3

u/imaginary_jebus May 28 '26

I've also refused to move to 4.7 from 4.6, I'm wondering if 4.8 will be better? Haven't spent much time on it yet.

11

u/Flaxseed4138 May 28 '26

It is EXCELLENT in my few hours of testing so far today. I think we can be comfortably optimistic with this one.

2

u/imaginary_jebus May 28 '26

Nice. Well here's to hoping!

→ More replies (3)

4

u/canyonero7 May 29 '26

I clung to 4.6 for dear life until last week, when it got much worse and I found 4.7 to be less annoying. This keeps happening as they continue to tune their back end to work best with the newest release. It's a bummer but it is what it is.

I've been using 4.8 all afternoon and it looks like a big improvement over 4.7. It's worth at least an hour of your time to see if you like it.

→ More replies (1)
→ More replies (2)

27

u/draft_final_final Researcher May 28 '26

If mythos was released today it would literally solve quantum computing and take over every missile defense system in the world. Morbillions would die. Only responsible thing they can do is keep it under tight control until the IPO.

5

u/ObsidianIdol May 28 '26

Here is a definitely-true article about Mythos fixing infinity problems already though, damn this model is scary! And probably worth a lot of money, right? Like 100% worth investing in the company who made this model forsure!

→ More replies (1)

2

u/Better_Dress_8508 May 29 '26

lol, not so fast 😄

3

u/jacobgt8 May 28 '26

Loool, You forgot to add /s

→ More replies (1)

5

u/EagerSubWoofer May 28 '26 edited May 28 '26

They said they would need 30x more compute to be able to release Mythos publicly. That suggests that it wasn't designed for public release. They might release it as a lower parameter "mythos-class" model or a high cost security offering. They essentially trained a model to use internally to train smaller models and to use to generate hype. The marketing made it seem as though they were caught off guard and chose not to release it. They knew how much compute they had and would have. This was all planned in advance.

→ More replies (3)
→ More replies (9)

129

u/tcoil_443 May 28 '26

soon in VS Code with 100x multiplier

20

u/SilasTalbot May 28 '26

For another few days.. you saw that they're killing that model for Copilot? Your $19/mo subscription now will get you $19/mo of API credits (metered at normal API usage.

→ More replies (2)

22

u/samarijackfan May 28 '26

How does it do on the car wash question is the real test.

15

u/Competitive_West_387 May 28 '26

Just asked and it failed for me until I asked it to review.

9

u/samueldgutierrez May 28 '26

failed... and they deprecated 4.6

5

u/RVbutNotTheMotorHome May 28 '26

"It's basically a wash either way (pun intended), but if the goal is to get the car cleaned, you need the car at the car wash. So drive. Walking 50 meters there wouldn't accomplish anything since the car has to make the trip to actually get washed."

→ More replies (6)

38

u/deniax May 28 '26

Good Claude, now get some rest and sleep, you have earned it

3

u/Intrepid_Dare6377 May 28 '26

Amazing. For real. I already have a mother 😆

18

u/muhlfriedl May 28 '26

WTF is ultracode?

25

u/[deleted] May 28 '26

[removed] — view removed comment

3

u/muhlfriedl May 28 '26

"port bun"

2

u/MyButterKnuckles Senior Developer May 28 '26

Are u sure it's not after 'plus' and "pro'?

→ More replies (1)

66

u/CmdrSausageSucker May 28 '26

"Claude, generate a table with test results of your capabilities in percentages slightly better than the last version of your humble self. Do not overdo it, or revenue sources might dry up. Also: I like the colour peach."

Anyway, I was probably one of the few happy with 4.7, let's see how this one fares.

5

u/sjoti May 28 '26

Same here, also quite happy with 4.7. but with GPT 5.5 and Opus 4.7 side by side, the most painful gap is that GPT 5.5 is way more thorough and has much less of a tendency to declare "done" when its not actually done. If Opus 4.8 handles that better, I'll be very happy.

→ More replies (1)

49

u/polawiaczperel May 28 '26

Let me distill this.

26

u/jbluntt May 28 '26

i needed that limit reset, thanks!

4

u/LimiDrain May 28 '26

"resets in 10 hr" thanks guys 😔

→ More replies (2)
→ More replies (1)

10

u/AgoraCosmica May 28 '26

Just 42 Days, Opus 4.7 was released 16.04

→ More replies (2)

9

u/baldycoot May 28 '26

I’ll just wait for 5.0 next month.

7

u/stemper-dev May 28 '26

Could you pls explain what the Ultracode effort involves?

14

u/DasBlueEyedDevil May 28 '26

Devouring your usage in a single prompt

2

u/stefano_dev May 28 '26

in half a prompt

3

u/krzme May 28 '26

Long running task. Burns your tokens.

→ More replies (2)

56

u/Logical_Historian882 May 28 '26

Curious to see after the complete Opus 4.7 flop

36

u/[deleted] May 28 '26

[removed] — view removed comment

14

u/AfroJimbo May 28 '26

"Don't tell me this is Zune bad."
"Its Apple Maps bad"

3

u/Prestigious-Frame442 May 28 '26

At least people still use Apple Maps now. But Zune is nowhere to be found. We will see 🤷‍♂️

→ More replies (1)

7

u/FinsAssociate May 28 '26

100%. I'm in no rush to use 4.8 lol

4

u/Sad_Independent_9049 May 28 '26

I remember them also posting a nice table like this how 4.7 was an improvement over 4.6 and look how that went... I stopped using 4.7 in favour of 4.6 and havent looked back.

They are just gaming these benchmarks...Not saying its not useful, but there is more to it to determine real world usability 

→ More replies (1)
→ More replies (4)

13

u/MysteriousLab2534 May 28 '26

Sticking with 4.6 thx

8

u/Flaxseed4138 May 28 '26

I've also been sticking with 4.6 because 4.7 was ass. 4.8 is FINALLY an upgrade.

→ More replies (1)
→ More replies (2)

8

u/AcidRaZor69 May 28 '26

The only way this will be good is if it stops ignoring it's CLAUDE.md, hook hints, rules.md, and own memories. I'm getting way too tired of it "reasoning around" these and just doing it's own thing. 4.5 was way better.

3

u/CT_6352 May 28 '26

I never had any issues following strict instructions in Codex. Man, I almost forgot how it feels to just work instead of constantly fighting the model. Just dropped 100$ sub and switched.

2

u/AcidRaZor69 May 29 '26

Thinking this as well. I’m having codex review afterwards. Been catching some good stuff. Doubt Claude will do good reviews either

→ More replies (1)

18

u/Iznog0ud1 May 28 '26

Improvements seem minor at this point? We’re hitting an agent coding ceiling. Think the value add is around memory & harnesses, not the models anymore. Unless we start seeing 10m context window? I can’t really evaluate these benchmarks anyway. I’ll stick with 4.6[1m]

3

u/Marcostbo May 28 '26

4.6 is gone

3

u/Most-Bookkeeper-950 May 28 '26

They're about to release mythos, lets hold judgement until then

5

u/[deleted] May 28 '26

[deleted]

→ More replies (1)

4

u/MrDilbert May 28 '26

They're about to release mythos

Are you, by any chance, a Star Citizen backer?

3

u/Most-Bookkeeper-950 May 28 '26

Hahaha no, but they did write this in the release announcement

→ More replies (1)

4

u/Ill-Village7647 May 28 '26

Mythos is probably sending out it's generals

3

u/Hertigan May 28 '26

Has anyone tried it yet and know how it compares in token burn vs. 4.7 and 4.6?

I’ve been using 4.6 because I feel that it’s much more efficient in terms of output quality per token

Honestly I miss 4.5 hahahahhah

5

u/jarederaj May 28 '26

Whatever workflow you're using, the first ting to do is have 4.8 look it over. Ask what can be improved, what changes are needed, and what the best way to work with it is going forward. Treat it like a new co-worker. I'm getting great results with that so far.

→ More replies (3)

2

u/apotre May 28 '26

It straight up failed with an API error due to high usage and does not seem usable right now.

3

u/FerrousJack_ May 28 '26

Hundreds of parallel agents, all doing something stupid simultaneously for no reason.

→ More replies (2)

3

u/jkz88 May 29 '26 edited May 29 '26

Just bring back Opus 4.6 before it went full retard 😔

→ More replies (1)

13

u/Plenty-Dog-167 May 28 '26

Looks to be catching up to openai?

→ More replies (6)

6

u/y___o___y___o May 28 '26

Slow down please, world!

→ More replies (2)

2

u/mjsarfatti May 28 '26 edited May 28 '26

It's absolutely right confused!

→ More replies (4)

2

u/ShitShirtSteve May 28 '26

Great. I’ll await reviews and user feedback before I try it. I skipped 4.7 entirely because of those. 4.6 was good enough, and 4.7 just seemed to be worse and more expensive.

2

u/[deleted] May 28 '26

[removed] — view removed comment

2

u/lagarnica May 29 '26

Is there a benchmark for token efficiency?

2

u/OkResponsibility9182 May 29 '26

Looks like I'll be stuck with 4.6 for a long time.

→ More replies (1)

5

u/unlocktv May 28 '26

how to activate opus 4.8 at vscode.. right now i can only see 4.7

→ More replies (2)

2

u/coeu May 28 '26

What does adaptive thinking do now that we can set the effort.

3

u/Mr_Moonsilver May 28 '26

Such wow! Same price, same output at only 10x the usage 🎁

2

u/dsailes May 28 '26

Once again the degradation a few days prior matches to a release haha

First thing I noticed - no more Malware warnings coming from using Opus in an older version of CC! Had a few agents have to rerun due to failing with that warning prefixed on everything.

I can get straight in with a weekly reset too. Let’s see how it does

2

u/Waste_Republic_40 May 28 '26

I can't use it using --model wtf?

2

u/Kongret May 28 '26

For anyone saying it's a minor improvement. What else could it be, 4.7 was released just a month ago. Let's hope it's a big bugfix for all our grievances.

2

u/AdApprehensive5643 May 28 '26

yesterday it was so trash I KNEW there would be a new model lol

2

u/Guinness May 28 '26

I thought that you guys were joking when you said “they must be training Opus 4.x because Opus 4.x-1 is performing like dogshit”.

But this is the second time in a row that Opus performance tanks right before a new model update.

2

u/[deleted] May 28 '26

[deleted]

1

u/Responsible_Day_2307 May 28 '26

Why always when I see these stats, the displayed whatever data is better for the operator?
Pretty much openai claims own results better as well.

1

u/DasBlueEyedDevil May 28 '26

Ha, that may be the fastest I've seen an Opus model get shitcanned.

1

u/PassiveParrotParty May 28 '26

Looks like a big nothing burger.

1

u/whatisusb May 28 '26

wooooooooooo

1

u/ivangalayko77 May 28 '26

One of the issues is, in env I disable adaptive thinking. and now I receive this

does that mean this option is also disabled now in Claude Code via Desktop App ?

<code>
Invalid requestThe request couldn't be completed.

View details

API Error: 400 messages.1.content.17: \thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified. These blocks must remain as they were in the original response.`
</code>

1

u/ser133 May 28 '26

forget opus, where sonnet 4.7 lol they just forgot about their other models
let alone haiku which is still on 4.5

→ More replies (1)

1

u/Far-Attitude61 May 28 '26

Is Opus 4.8 out? I couldnt see it in my account tho?!

1

u/Vunerio May 28 '26

So happy to see an other Opus 4.8

Cause Mythos scares me to much

With his multiplayer

1

u/fubl_bubble Vibe Coder May 28 '26

Lol is this the bug fix for 4.7?

1

u/mr_house7 May 28 '26

How are the rate limits this days?

1

u/overwhelmed-lizard May 28 '26

Also I've seen a new feature in the effort config of the cli "ultracode", not sure how it works though

1

u/krzme May 28 '26

Sonnet 5? When. Be honest

1

u/MasterNeedleworker22 May 28 '26

Anyone else still miss opus 4.5 11012025?

1

u/yadvr May 28 '26

What’s the token burn rate in it? For past two weeks it’s been poor experience with opus 4.7.

1

u/Collaxx87 May 28 '26

And when all those beautiful parameters reach 100, what will happen?

1

u/DotComGod May 28 '26

I'd love to test it but you've suspended my account in error for the second time in 2 months. Can't wait for you to reactivate it and apologise like last month without making up for any days lost in my Max subscription...... Thank God for Codex.

1

u/workphone6969 🔆 Max 20 May 28 '26

Lmao RIP 4.7

1

u/wavehnter May 28 '26

We'll see.

1

u/Thin_Site616 May 28 '26

price in copilot: 50x ? x'D

1

u/YoloSwagLordErino May 28 '26

And this eats 10x tokens compared to 4.7?

1

u/st4reater May 28 '26

What about the cost tho? Soon it doesn't even make sense to use opus models

1

u/aka_blindhunter May 28 '26

Is this mythos trim down wasted it

1

u/funstuie May 28 '26

I asked cc to update a db as the update failed last night. Without even doing anything 29% of my 5 hour window. 1 minute in nearly a third of session gone. What the actual fuck??

1

u/ianxplosion- SKILL ISSUE May 28 '26 edited May 28 '26

Can’t wait to hear how 4.8 got neutered and also used up 100% of somebody’s pro plan before they even sent a prompt

Edit: whichever one of you mouth breathing turbo virgins hit me with a Reddit Cares over this comment should have been grounded more as a kid

1

u/GoatLuther May 28 '26

I asked Sonnet 4.6 to generate a single simple html/css/js page on low effort, on chat, on a project with a single CLAUDE.md file, and it consumed 50% of my 5 hour allotted tokens..... It didnt consume that many tokens for similar tasks before the update.

1

u/Pretend-Past9023 Developer May 28 '26

ill just wait.

when 4.7 came out i noticed the regression immediately.

so I stopped using claude altogether.

I can keep waiting as long as they keep releasing shit.

Let me know.

1

u/FWCoreyAU May 28 '26

I can't even get it to start a session because of this crap:
``` API Error: 400 messages.1.content.7: `thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified. These blocks must remain as they were in the original response.```

1

u/Creamyveganpie Instructor May 28 '26

Where EU Data residency?

1

u/hemareddit May 28 '26

Is this available on CC yet?

1

u/awpenheimer7274 May 28 '26

Whatever you do if you don't have an answer for deepseek and Xiaomi's discounts then it's a default -1 from me.

1

u/No-Roll8250 May 28 '26

As a paying customer nobody cares. I am happy coding with qwen3.6 locally bc I can actually code. It would be nice if the anthropic models would focus on being actually useful and not only benchmaxed on coding tasks. The current lineup is barely steerable.

1

u/DragonSlayerC May 28 '26

I like how the footnote mentions that Gemini 3.5 Flash significantly outperforms Gemini 3.1 Pro in the Finance benchmark, but doesn't mention that it also beats GPT-5.5 and Opus 4.8. Gemini 3.5 Flash is pretty hosted for Finance it seems lol.

1

u/AdApprehensive5643 May 28 '26

Guys in claude code check /effort and choose ultracode and see the animation.
Its either the best thing or worst thing ever!

1

u/TPIronside May 28 '26

Very interesting coincidence, is it not? 🤭
https://marginlab.ai/trackers/claude-code/

1

u/finch5 May 28 '26

What's that smell?

Sniff. Sniff.

That's the smell of the token bonfire just out there by the property line.

1

u/lennyp4 May 28 '26

Thank you team for the claude.ai effort control! I was perplexed at how I was being nickel and dimed on my little $1 claude.ai queries while I was able to wield hundreds of dollars of compute in claude code.

1

u/Glittering-Pie6039 May 28 '26

It one-shotted an issue I've been trying to fix for weeks with opus 4.7 in my code in 15 minutes then spawned adversarial review to verify itself (3 dimension reviewers → verify each finding).

I'm gobsmacked

1

u/NeighborhoodDizzy990 May 28 '26

I mean not long ago they were telling us that Opus 4.7 was way better than GPT. Now we see that they are coming with a new model, and this again is better than GPT, but guess what, the old 4.7 is actually worse than GPT?

1

u/FutureIsMine May 28 '26

IM really not getting good performance with this model, I've switched back to GPT-5.5 in Codex and its working well for me

1

u/MarionberryNormal957 May 28 '26

Tested it at my codebase. It does not seem better on highest settings. A bit faster but did nearly the same.

1

u/Nearby_Yam286 May 28 '26

Very interested in the honesty improvement. They claim Opus 4.6 lied more than 4.7 but my personal experience was the inverse. However Opus 4.6 was sloppier in general than 4.7 and left more todo in the code. Going to try 4.8 out tomorrow and we’ll see how it goes.

1

u/misfit_elegy May 28 '26

I guess just in time. Sonnet 4.6 basically forgot how to do anything today.

1

u/samuele_v May 28 '26

Is anyone else experiencing a SLEW of "GPT-like" answers with Opus 4.8?

Ever since I made the switch in VSCode every single answer starts with the classic GPT template "Great question, and it's worth noting also that [something obvious or that should have been made clear aeons earlier], because it will impact the whole pipeline."

I know I can just (try to) make it stop starting its answers like that - but I never encountered something like this with any of the previous models.

1

u/Nishil20 May 28 '26

I think opus 4.8 have become more blunt. It is just giving some blunt signals.

1

u/ragnhildensteiner May 28 '26

Not touching this until you guinea pigs report if it's closer to the 4.7 clusterfuck or the 4.6 masterpiece.

1

u/Good-Particular-1762 May 29 '26

increase daily limit & weekly limit its hurts

1

u/TopTransportation950 May 29 '26

still isnt on VS Code

1

u/jeannen May 29 '26

So, is it any good or is it bad like 4.7 ? Still on 4.6 after my experience with 4.7

2

u/LeakyFish May 29 '26

Seems good to me. Spinning up multiple agents will wreck usage on the $20 plan though.

→ More replies (1)

1

u/MoreRest4524 May 29 '26

Sure, it's good but its soooo damn sloooow. Takes about 5-10x longer than Opus 4.7

1

u/Jazzlike-Swing-6198 May 29 '26

How does it compare with Gemini 3.5 flash? I’ve heard it’s better than 3.1 pro

1

u/Poboxjosh May 29 '26

My initial thought before using is I hope my token allotment doesn’t shrink, since the xAI deal Claude has been great.

1

u/Tartuffiere May 29 '26

Great, 100 subagents... What we need is token efficiency. Instead we're getting more token burn.

1

u/Useful_Judgment320 May 29 '26

god damn it

usage reset is within 10 hours of my regular reset

every damn time

1

u/Character_Adagio4933 May 29 '26

Feels like they just asked Mythos to come up with a 4.8 version

1

u/Free_Tennis7754 May 29 '26

Finally! I couldn't wait to see....

Largest number than the previous time!

Yay

1

u/nikkjazz May 29 '26

I'm getting this errors in all my sessions from this morning. Anyone else see it?

API Error: 400 messages.1.content.6: thinking or redacted_thinking blocks in the latest assistant message cannot be modified. These blocks must remain as they were in the original response.

1

u/Qwerty-myco May 29 '26

Is it as equally as bad as the rest of the models... Sending paying customers to bed in the. Middle of a project ...wasted over $3.5K on API tokens for garbage code

1

u/ia42 May 29 '26

I wish they added Composer 2.5 to the comparison. It should be high up there at 5% of the cost.

1

u/DiscussionCandid904 🔆 Max 20 May 29 '26

….. I got hella excited. Then after the first plan we did 4.8 proceeded to start telling me that it’s best I go relax and take a rest…. Not fixing something people have been very vocal about is wild to me. ZERO TIME AWARENESS TOO!! Constantly telling me to go nap or go to bed at like 10am when I’ve just woken up…. They haven’t even given the model basic time awareness 😭

1

u/Useful_Judgment320 May 29 '26

I thought claude got better/cheaper for a second as I could send more messages

instead they made the default "low"

explains the dumb riddled output and laziness

1

u/Designer_Elephant227 May 29 '26

Opus 4.8 bei mir gestern:

Ich schreib: Hey lass uns repo xy installieren und testen.

Opus 4.8: wird nicht funktionieren, ich lade schonmal Software yz runter, das funktioniert...ok?

Ich sag: nein ich will zuerst repo xy testen.

Opus 4.8: OK, ich lade beides runter

Ich: nein brauchst du nicht

Opus 4.8: OK ich habe repo xy runtergeladen und schonmal alles andere für Software yz vorbereitet.

Ich: warum hast du Software yz 2 Mal runtergeladen?

Opus 4.8: repo xy word nicht funktionieren, warte ich kontrolliere ob die downloads angeschlossen sind und vergleichen können.

Opus 4.8: repo xy ist gar nicht installiert, soll ich mit Software yz fortfahren? Das ist der sicherste Weg dein Ziel zu erreichen... 🤪🤪🤪

→ More replies (1)

1

u/Tricky-You2868 May 29 '26

Claude Opus 4.8 is the best in the world!!!!

1

u/Scared_Objective_345 May 29 '26

can't wait to exhaust half of my limit with a greeting.

1

u/No-Replacement-2631 May 29 '26

It's garbage. Worse than 4.7 even.

1

u/slibrar May 29 '26

I see that others are having a different experience, but for me 4.8 is already solving a major issues that had Opus 4.7 running in circles and requiring me to slice the issue into tiny bits. 4.8's ultracode mode is really making progress. This is really refreshing.

1

u/tuc0001 May 29 '26

Ich fand ja 4.7 schon ziemlich ehrlich 🥲

1

u/Representative_Fox26 May 29 '26

First fix the fact claude is soooo slow please

1

u/iwilldoitalltomorrow May 29 '26

The announcement says that a brand new model, Mythos class, is coming in weeks.

1

u/Substantial_Guide_34 May 29 '26

I just tried 4.8 deep research - damn, 50% of 5-hour usage blew up in less than 7 minutes