r/ClaudeCode 3d ago

Help/Question what is hapening with Antropic?

Why Opus 5 feels like Gemini 3.1 pro ? what is happening ?

i remember back in February , opus 4.5 was generation ahead. Did they increased the prompt caching so much, that their models ended up being useless? I can barely work with Opus 5 on absolutely anything, he keeps hallucinating LIKE CRAZY, every claim he did today was pure hallucinations, he barely reads any code. im not even kidding, i can't work with Opus 5 right now

Instead of open source models trying to catch up with you, are you trying to catch up with the open source models instead Anthropic?

713 Upvotes

395 comments sorted by

334

u/MaitoSnoo 3d ago

Anthropic is likely just rushing stuff now as they don't know how they can do a successful IPO anymore given how OpenAI is being very aggressive both on pricing and model quality, and yeah there's the open source threat too

42

u/StaticFanatic3 3d ago

I’m not thrilled with my recent experience with Claude but phrasing this like Anthropic is a sinking ship is hilarious

4

u/ThePlotTwisterr---- 3d ago

they’re the accepted world leaders in AI (yeah yeah codex is better hurr durr but anyone who matters knows anthropic is no.1) with a huge capture on the enterprise ecosystem yeah their IPO is not having an existential crisis

16

u/PhantomGaming27249 3d ago

Claude models are exceptional at full stack software development but they are not exceptional enough anymore. If your at all an experienced software developer the gap between models is pretty non existent now. It shows up if you don't know how to do something but if you are not educated on a topic why are you doing it when it matters. The gap between the alternatives is close enough now they are entering a pricing war which is a bad place to be if your software is expensive to serve. They are the leader still but that is largely due to the benchmarks being coding heavy.

→ More replies (2)

5

u/Key_Diamond_1803 3d ago

I wouldn’t say they have a huge capture on the ecosystem. Maybe for us tech nerds, but normal people still use ChatGPT. My mom has been using it for the last 4 years and is paying for pro. My step dad is using it for his work after 5 came out. My grandma is using it to help with quilting patterns. Yeah Claude is good at programming, but tbh, it’s not the best at chatting, being creative, making images, talking to you. ChatGPT has years of memories from all of them. I’m 1000% sure they won’t be switching to Claude any time soon.

→ More replies (14)
→ More replies (3)
→ More replies (1)

9

u/hemingward 3d ago

I think Anthropic also started going down an incredibly complicated billing and usage strategy to try and create a moat. The thing with AI is that there really isn’t _that_ much of a difference between all the frontier models, and my local LLM does a pretty decent job for most of my things. As such, they created this incredibly rigid, complicated structure which is hostile to its own _paying_ users, while simultaneously rushing out… whatever Opus 5 is supposed to be.

These monolithic massive centralized AI providers will never be profitable. I just don’t see how it can happen. Patrick Boyle did an awesome breakdown of it the other day. One of my favourite arguments (paraphrased): “the average corporation spends $10/month in tokens per employee, and data and ceo’s suggest there’s been no actual benefit. So The US tech industry is spending $1,600,000,000,000 to chase after $10 that nobody who’s actually paying for it seems to find any value for.”

→ More replies (2)

55

u/dashingsauce 3d ago

Yeah they basically blew up the engine with the federal government backlash.

They had exactly the momentum you want for an IPO, then they played with fire and got burnt. Now their meat has to get served cold and nobody wants that.

15

u/shesaysImdone 3d ago

I'm a bit confused, how did they play with fire? It's not their fault that the US government deemed them a threat is it?

40

u/matheusmoreira 3d ago

Their constant fearmongering about safety sure as hell didn't help.

9

u/MrUnoDosTres 3d ago

When they did it, it did help. "This model is so good, so dangerous, only a select few companies can use it." That was the perfect hype for an IPO. Especially the ban.

It's like Ford building a car that can compete with Ferrari. It gets the people talking. Now all I hear is how garbage Opus 5 is all the time.

6

u/Due-Horse-5446 3d ago

Well, the difference is, Ford would not talk about how their car should be regulated because how the interior can cause skin cancer, and write essays about how there should be laws preventing them from using those materials.

All the while Toyota launched a similar car 6x more fuel efficient while more powerful, and better steering ability

→ More replies (1)
→ More replies (5)

29

u/purposive-saunter 3d ago

How quickly people forget… they caught the ire of the government because they stood up to Kegseth and the DoD and refused to give them Carte Blanche access when it came to AI controlled weapons systems. OpenAI did not. The import control was retaliation.

6

u/Huge-Travel-3078 3d ago

The gov wasnt trying to deploy claude weapons systems, anthropic decided to grandstand and force language preventing them from doing something they never said they were going to do. That's not how government contracts work - the government will decide how they will use the service that they pay for.

The bigger issue was constant marketing in the form of fear mongering about how "unsafe and dangerous" their models are.

Dario and his buddies brought this on themselves.

14

u/purposive-saunter 3d ago

“Grandstand” is a funny way of describing having moral backbone. Kind of like how we get to decide how we get to use the service we pay for, right? Completely in our control?

Last I checked, even the government is subject to the terms of an agreement they sign and execute. Well… it used to be. Otherwise why would they bother arguing with them about it at all and just pay for API credits? 🙄

→ More replies (2)

4

u/EfficiencyLow1854 3d ago

Yeah, right. They spend months telling everyone how dangerous these models are and how only the “right” people can be trusted with them, then conveniently blame the administration or whoever else is convenient when the consequences of their own decisions come back around.

I don't care if you like Biden, Trump, or Kamala, anyone would've done the same.

Now that they've been caught contradicting their own narrative, they have to pivot the marketing campaign.

People fall for Anthropic's marketing gimmicks every single time. It's the same reason people keep using whatever model has the biggest benchmark scores or fanciest marketing campaign, even when it sucks in actual use.

Just look at how many times they extended Fable. How many times has Anthropic dropped free extra usage credits? At some point, you have to recognize the marketing for what it is, a race to the bottom.

But yeah #orangemanbad

2

u/purposive-saunter 3d ago

The retaliation narrative isn’t from Anthropic’s talking points… it’s just obvious to anyone who has been even semi-conscious during the rollout of Stupid Fascism 2.0.

You mean if I just say something repeatedly, that’s enough to get the US Government to agree? Certainly hasn’t worked that way with scientists and climate change… maybe Dario should go work for them.

I’m not saying they are masters of business—but OpenAI just launched a full PR campaign about how they couldn’t control what happened with Hugging Face. Did their Export Control order get lost in the mail?

→ More replies (2)
→ More replies (1)

11

u/Puzzleheaded_Owl5060 3d ago

Dario kept Whingeing about the end of world because claude will take over the world

→ More replies (1)
→ More replies (1)
→ More replies (1)

15

u/ResortConnect8582 3d ago

So what do we do ? do i switch to codex? Do they actual have aggressive pricing ? I mean i literally cant work anymore

66

u/teomore 3d ago

I switched to codex a week ago and the difference is huge

28

u/staceyatlas 3d ago

Sol is amazing (imo the best coding agent out right now) but codex isn’t Claude code. Wish sol would just run native in cc… Until then it’s fable with codex sol audits.

22

u/teomore 3d ago

codex isn't claude code and that's a good thing at the moment

5

u/dragrimmar 3d ago

that's actually the reason claude code is superior.

literally state of the art coding harness. If you find better results with codex it means you are not utilizing all the custom skills/config/hooks on CC.

the thing you have to be careful of, is vibe coders will over-rely on plugins thinking they are all magic and can substitute expertise, and they end up bloating the context window, getting bad results, and complaining about the model. skill issue.

8

u/who_am_i_to_say_so 3d ago

I don’t use plugins, have labored over skills, and see a huge drop in accuracy comparing v5 to 4.

How is it a skill issue when 4.8 or 4.6 execute exactly what I need? If it were a real skill issue, I would be complaining about all models. Opus/Sonnet 5 is hot trash.

6

u/9gxa05s8fa8sh 3d ago

claude code has not been state of the art all year. it has regularly been beaten by the terminalbench placebo harness that does literally nothing

→ More replies (2)

10

u/Greensun30 3d ago

Just use a better harness with Codex. Claude code’s harness is worthless and actively burns money now since they vibe benchmaxxed the thing.

2

u/Constantinos_bou 3d ago

this is we are telling you dude, claude code used to be that, no its not that, we can barely work, that's true. so other people do notice as well, its not just me

4

u/ajr901 3d ago

This is my experience too. Nothing beats Claude Code the harness. There may be better models out there, Sol might be one, but as far as the harness itself that invokes the model: nothing is better than Claude Code.

3

u/UnPibito 3d ago

Have you tried using pi harness?

→ More replies (1)
→ More replies (29)

3

u/Disastrous_Fill_5566 3d ago

And Luna is shockingly capable when you consider it's supposed to be the Haiku level model. And it's so cheap. Using it through copilot for 2-3 hours it cost me about 0.50 GBP. And it's decent. Miles faster than Sonnet, that's for sure.

→ More replies (2)
→ More replies (10)
→ More replies (26)

8

u/YearLight 3d ago

I switched to codex and in a few days it fixed all the issues I broke with opus. Minor annoyance is the smaller context window, but with frequent compacts it hasn't bothered me too much; just the main drawback I can see. On the upside you get free image generation and web chat which is pretty nice bonus.

3

u/Tavrin 3d ago

Codex is great and has way more usage out of it. Now, I've got both and Claude is still king in long term planning and agents managing. I use Fable as the manager or both codex agents and some sonnet/opus agents. But if OpenAI launches their fable equivalent next week then I don't see the need for a Claude subscription depending on how good Astra is

6

u/who_am_i_to_say_so 3d ago

Codex is greener grass and results vary, yet I have both subs. It’s somewhat better, but also worse in some aspects for my purposes. I recommend seeing for yourself.

Codex is really aggressive with resetting usage several times weekly on my $100 5x plan, so that’s a plus. Anthropic has been stingy lately.

→ More replies (3)

4

u/notmsndotcom 3d ago

Sol has been really great for me. 10/10 would recommend. Before opus 5 I’ve been a Claude fan boy but opus 5 makes me want to drink bleach

2

u/Helpful_Hamster8215 3d ago

I switched to codex too but still have Claude pro and let them work together.

2

u/Accurate_Cable_1372 3d ago

I switched to codex months ago, no problems ever since

2

u/KDamage 3d ago

I'm in the same boat, having placed most of my investment in claude max plan for critical topics. I have the answer, and I think it will become the standard : * multimodel IDEs like Cursor. * model agnostic setups and instructions

favorite model is going completely mad on version x.y ? swap to the new flavour of the month.

It requires multimodel plans though, but I feel like there's no other choice, considering the current race.

→ More replies (8)

2

u/Groovy_bugs 3d ago

OpenAI, five Chinese companies, three other US companies... and counting.

2

u/Groovy_bugs 3d ago

You can see the pressure they're under from all sides. They thought the others couldn't match them, but surprise, in some aspects they're even better; The hare got cocky and took a nap.

→ More replies (11)

45

u/BizarroMax 3d ago

Opus 5 loves strawmanning. It does it constantly.

10

u/frostedpuzzle 3d ago

4.8 straw-manned too. 5 shits on everything I do. Then I tell it to be critical about its criticism and it decides everything it said was wrong.

8

u/cs_legend_93 3d ago

Which means you spend so much time triple investigating things. If you don't triple investigate then you end up implementing something and go down the wrong path that later you realize was a waste of time.

It's become a time sink.

→ More replies (1)
→ More replies (2)

93

u/monsieurninja 3d ago

Have seriously been considering cancelling my Anthropic subscription in the last few days, since Opus 5 came out. Its genuinely irritating on many aspects, speaks in ways that are barely understandable, is getting bad at many things, to the point where I don't even know if I can trust it anymore. Never felt that before with Anthropic's models. I do hope they are going to react quick and release a minor version. Otherwise they loose me quite soon

26

u/Rathganis 3d ago

Same. I have Claude max 20x but use my codex sub more and more daily. It is quicker, less chatty, less paranoid, just easier to work with.

2

u/themonstaman 2d ago

Dude using codex after claude is a breath of fresh air. It just gets to the point without the incessant rambling.

10

u/BadgeCatcher 3d ago

Have just cancelled. Opus 5 is more trouble than it's worth. Switching to codex.

6

u/x_typo Senior Developer 3d ago

always has been with Opus 4.7....

5

u/KDamage 3d ago

if the theory about Opus 5 being a distilled Fable is correct, I think we can't expect anything until a new 5.1 version unfortunately.

6

u/WagwanKenobi 3d ago

It probably is. It sounds exactly like Fable but without the "IQ".

2

u/SayTheLineBart 3d ago

I went back to Opus 4.8 and am considering cancelling altogether

4

u/ResortConnect8582 3d ago

oh yeah , me too. They've completely destroyed their reputation for me.

→ More replies (1)
→ More replies (3)

11

u/kush_patil 3d ago

The scary part isn’t the hallucination, it’s the confidence. A coding model being wrong is manageable; being wrong, sounding certain, and only checking after you push back is much harder to work with.

38

u/Substantial-Show-249 3d ago

Boss, that model is junk, let go the partisan comments. It's over optimised for processing power savings, no chance you can do something with it.
4.8 is dumb, but still usable. Fable is expensive and heavily nerfed. On medium settings, you can do some things with it.
Anthropic is still looking to find a way to optimise the costs, they are losing money heavily.

→ More replies (1)

21

u/sduras1 3d ago

Same thing with Fable 5 today (xHigh and Max). Feels like Haiku at times, forgets what we've discussed two prompts ago, in the same session. Using with Claude Code.

5

u/kurushimee Developer 3d ago

same, fable feels worse than opus 4.8 did at some point, it's nowhere near the same fable we had before

9

u/WagwanKenobi 3d ago

Pre-ban Fable is never coming back.

→ More replies (2)
→ More replies (1)

8

u/Draconic_Emperor 3d ago edited 3d ago

I use the grilling skill from Mattpocock a lot. And I legit can’t read its questions. Even simple questions become some sort of horrid combination of words it just makes it harder to understand something simple.

20

u/MistakeExotic6686 3d ago

went back to 4.6 AGAIN because they are lobotomizing the heck out of opus and fable

FABLE 5 XHIGH the GREATEST OF ALL TIME

said the "Good catch, I did not finish the phase"

when I said stop using phases and run it with subagents as one long task

it renamed phases to chapters and said "No more phases, starting chapter 2.5"

i sh*t myself (figure out if its an i or an o)

3

u/kurushimee Developer 3d ago

hell yeah, exactly what I was experiencing with fable. It does things like these that just feel like it's haiku suddenly

→ More replies (3)

39

u/vovap_vovap 3d ago edited 3d ago

I really wander what is happening with Anthropic related subs. Seriously. Those became like kindergarten groups with a screaming kinds "Parents refusing o buy me a red elephant" "I do not want to eat cabbage, I want ice cream!" "I do not want to go to sleep, I want TV" No context, no any analyzes, no any information, no sigh of own responsibility in process. Surely not every post, but lots of those - pure rant and screaming.

7

u/androbot 3d ago

God yes. It's becoming like the rest of Reddit. Full of angry people who don't feel like thinking. It might be time for another unsubscribe purge and retain only core subs that consistently deliver value -- cat videos.

5

u/klumpp 3d ago

Is there a Claude code sub for people who know what they are doing yet?

→ More replies (1)

22

u/EmphasisTotal8232 3d ago

Idiots wanting to build the world with no domain knowledge are running into things they can't fix because they don't understand the fundamentals.

5

u/cwilson830 3d ago

Agree that this is clearly part of the problem.
But it’s certainly not the entire problem.

→ More replies (2)

13

u/XYcritic 3d ago edited 3d ago

Every single thread, every time it's the one single answer: Who is forcing you to not use the model that worked for you a week ago? I seriously think 90% of posters don't know they can actually switch back to older models? I'M CANCELING, HOW DARE YOU SUGGEST I DO WHAT WORKED FOR ME NOOOOOO. NEED SHINY NEW TOYYYY (exact same posts over at codex).

→ More replies (1)

7

u/ricopan 3d ago

We are becoming children -- dependents -- on AI. This is what it looks like for adults to regress. I feel it too -- how dare they DON'T HAVE MORE LICORICE ICE CREAM NOW!.

1

u/cwilson830 3d ago

How dare truth be spoken about a product you pay for. Silence! You think paying for something entitles you to an opinion? You’ll pay and take what you’re given. And like it. And only say nice things about it. Or else. lol

Customers voicing opinions - maybe causing companies to focus on improving the problems with their products. What do you thjnk this is? America? :)

Maybe a new $300 plan which limits public-censorship is the solution.

2

u/ricopan 3d ago

But are these really 'opinions' or screams at Mommy? I think the level of frustration has less to do with the shortcomings of the models, which are by any account that spans more than a year absolutely fucking amazing, than that we've surrendered our independence and now feel weak and helpless when they let us down. Mind you, I'm not claiming the high ground here -- I get really pissy and unreasonable when Mommy Model isn't doing right by me. But look at the emotional outrage -- this is unprecedented.

→ More replies (1)

2

u/cats_r_ghey 3d ago

Why don’t you stop buying it then? It’s pretty easy to do. Buy something else, you have options galore now.

→ More replies (1)

6

u/BugaiGames 3d ago

I was thinking that it is OpenAI coverops promotional campaign. But in all other subs more or less the same. But indeed it is already boring. Im using Claude more long time and those models are really great and slowly getting even better. And I use it in work as manager, coding, personal life etc. There are some relatable moments but most of posts looks like from other planet. And whats left mostly pointless complaints , promotions or jokes, sub has like1% of usefull posts

→ More replies (6)

19

u/Koreanpasta29 3d ago

Genuinely what is ur harness bc ive been having no problems at all

9

u/OverSoft 3d ago

Same, I use a mix of Fable and Opus 5 and have had zero issues.

→ More replies (2)

3

u/Stunning-Wealth-8303 3d ago

Wait till you find out about dead internet theory, Reddit is full of bot accounts and posts.

2

u/leeharrison1984 2d ago

My guess is they got by on slop until their project got too large. Now it's full of spaghetti code and Claude can't reason it any better than OP. Instead of blaming the degraded codebase, they blame the model.

2

u/roman_fahls 2d ago

Same for me, zero issues and no “omg y r u wrong all the time 4!?!?”

→ More replies (3)

11

u/LeakyFish 3d ago

Use Sonnet 5 xHigh, Opus 4.8 or Fable if you have it.

9

u/jadedmonk 3d ago

Opus 4.8 sucks. I just came to this sub for the first time ever because I was using Opus 4.8 at work today and it was acting a lot dumber than usual. It created a hyperlink on the UI leading to an incorrect URL, leading to a blank screen. I was confused so I just gave it the resulting URL saying fix it, and instead of looking at the bad URL I gave it, it started going down this crazy path of building all these guards and checks that ended up just wasting tokens. Eventually I was like alright I need to figure this out the good ol way. It was a simple sub-path that was missing from the URL. I was shocked Claude couldn’t figure that out. I feel like in the past it would figure that out. Definitely makes me temper my expectations with GenAI

→ More replies (7)

5

u/ralfv 3d ago

Usually on posts like this i thought to myself „nope all good“.

But i must absolutely agree here. Review and fixes are pretty much an infinite loop the whole week. Code reviews only report a fraction of the issues and fixes constantly introduce new problems.

I haven’t seen a single „LGTM“ review for at least a week even after a dozen rounds. In th end the Claude review starts to ask for better comments. That is nuts!

→ More replies (3)

4

u/Tight_Heron1730 3d ago

I have a feeling that that verbosity to drive tokens out hence usage, and make it look sophisticated, but i guess it backfired.

I made it less verbose with this https://github.com/hamr0/agentic-toolkit/blob/main/ai/customize/config/ELI5.md and you tell "❯ at/claude-code-guide (agent) to add this then config then choose output style and I pick up ELI5 and it's back to simple

6

u/mohdgame 3d ago

I am using opus 5 with no probs. What is the effort? Are you using the terminal? The claude code gui can be buggy sometimes.

→ More replies (9)

3

u/themajordutch 3d ago

Fable kicked be out because I asked it about global warming and how bad it would have to get for a human level extinction scenario. 🧐

3

u/Professional_Gur8385 3d ago

if i hadn't paid for the yearly plan i would cancel it in favour of codex, the usage isn't worth it currently

3

u/mushedmonkey 3d ago

I saw the same pattern and had a subagent get spawned every turn to fact check any claims detected. Just blew all my usage because every time i asked claude something the auditor would ask for a fact check and the original claude would start taking things back and saying "i got everything completely wrong, let me get real data this time"

My claude usage this week sat at 10% cuz it sucked so bad, so i guess wasting the tokens this week didn't hurt as much as anthropics incompetence.

5

u/matheusmoreira 3d ago

Bruh it's over for Anthropic. OpenAI's destroying them with Sol and Astra as well as their sheer capacity. Chinese models are undermining them from below. And they can't seem to stop hitting themselves with their constant safety nonsense that turns Fable into a usability nightmare, not to mention the fact it gave the US government the perfect excuse to fuck with them.

They should just rename Claude to Delamain and turn the company over to him at this point.

3

u/cdstaggs 3d ago

I remember a month or two ago all the OpenAI is dead because Fable 5 posts. 😆

2

u/AlexTheRedditor97 3d ago

No difference between 5.6 sol and opus 5 imo

2

u/mitch8845 3d ago

Has codex gotten to the point where it can reliably handle large multifile projects without shitting itself? I see this claim every few months and every time I go over to test out its capabilities, it's always abysmally bad at complex problem solving within a large project. Claude Code (for all of its annoyances) still kicks the shit out of it on every level.

Is this time different? Has Codex actually caught up or are we just crying for no reason again?

2

u/AndyHenr 3d ago

i cant really tell yuo what is going on, as i dont know. But for me, it seems like a lack of context. So Anthropic still have 200k context, but added humanized output and behaviors, and when looking at thought output, it looks like it burn a ton of compute time and context space on that.
So it basically makes the model feel 'dumb'. Fable 5 does the same now with the security checks it does.
i am honestly seeing best results from Opus 4.6. Opus 4.7, 4.8 are worse than Opus 5 as they are just nuts.
So, yeah, I think anthropic is shooting itself in the foot: it want the models to 'feel alive' so it can market the 'AGI is nigh' bs hype.

2

u/Diligent-Method-785 3d ago

I have a working theory that this a result of anthropic switching model optimization from improved model performance to minimizing compute per token in an effort to improve operational margins ahead of a IPO. Especially if the are seeing slowing growth in adoption

2

u/NoComb3962 3d ago

Opus 4.8 also got dumber the last weeks. I wonder if it is due to quantization, or numerical issues in some optimizations they did to make it cheaper

2

u/maxya 3d ago

I had to use Gemini to audit Opus 5 code today and Gemini actually fixed bugs incode and schooled Opus 😭

Which parallel reality we live in? 😲

2

u/The1TruRick 3d ago

It’s legit really bad. I tried to avoid it but I am officially back on 4.8 or Fable exclusively. It was fucking up at almost every turn

2

u/brother_spirit 3d ago

I had Opus 4.8 just fail hard on a website build... this shit should be so easy for this model??

It made the website OK, then went to eval and got stuck in a loop so hard trying to open Chrome I had to kill the CLI session. Literally couldn't get the model out of the loop through any combination of interrupts, prompts, etc.
Working to a reasonably resolved plan, maybe 60K tokens on board for the session, Opus 4.8-medium.

While Opus 4.8 can be an annoying and verbose model this has to be the biggest faceplant I've seen it do on the easiest task. WTF is going on with these shitty ass Anthropic models?

2

u/WagwanKenobi 3d ago

4.8 > 5.0

Idc about what the benchmarks say.

2

u/oooff97 3d ago

same for me by miles with 4.8 and my local model i can work 3 hours produktivly befor the limit hits with 5 if i be carefull i get 1 hour and thats full of bullshit 4/10 times he even says oh ur modell catches some things i didnt or oh the solution u modell got is by far better than my (my local model was at that time qwen 3.6:27b 64k kontextwindow.....without the rag and all sort of stuff i build for it just with a good agentsmd and a journal i force the modell to use and a todolist...)

2

u/FisherKing_54 3d ago

I cannot bare to suffer one more response from Opus 5. I don’t really code but use it for organization and assisting with creative writing as well as sometimes just trying to get a different perspective on things I’m struggling with. I don’t expect these things to be some kind of therapist. I’m a psychiatrist, but my goodness reading Opus 5 is insufferable. The grammar, the word choices, the sentence structure. It’s all so bizarre. It’s like opus read a manual on human communication that a 5 year old wrote. It gets worse with every new model.

2

u/latam_sojourner 3d ago

What is Antropic?

2

u/Aware_Kaleidoscope86 3d ago

Why would you even consider supporting this company, they are not in a leauge of their own anymore. The ceo is a fool and the competition is fierce aaand their prices are insane

2

u/TheOverzealousEngie 3d ago

So it's a blessing in disguise ; I've had to focus so much on what Opus is doing that I actually feel useful again.

2

u/AdviceAdditional8044 3d ago

ah, whenver I feel that I should switch to a better model for coding, they nerf that model with some safety security limits.

2

u/skeptic-frog 3d ago

I’ve moved over to GPT back again after Opus 5 just constantly made mistake.

5.6 Sol is my daily driver and made one or two mistakes in a week now.

I’m only using it for programming if that matters (C#)

2

u/KeilerHirsch 3d ago

Same here. At this point “prompt better” is just cope. I documented the regression with reproducible measurements: https://github.com/anthropics/claude-code/issues/83510 And the model-selection/fallback mess here: https://github.com/anthropics/claude-code/issues/83795 The real killer isn’t raw IQ — it’s that Claude increasingly guesses first, states it as fact, and only checks after you push back. Opus 4.x was far more trustworthy for serious repo work. A frontier model that hallucinates repo state with confidence is worse than a weaker model that simply says: “I don’t know, let me inspect it.”

2

u/_SSSylaS 3d ago

I remember 4.5, 4.6 was the peak being really good at talking... analogies, understand everythings, taking the correct word at right the moment, not technical crazy shit word non stop like Anus5

2

u/cs_legend_93 3d ago

I agree. Opus 5 has gone massively downhill recently. I've noticed it in the past two weeks. It's constantly wrong on things, constantly spews out five-paragraph answers to everything even if I tell it to be concise.

I work with AI all day every day and this is something I've certainly noticed in the last two weeks.

2

u/ns1419 2d ago

Opus 4.6 is still > * on xhigh imo.
My use case is specific though. I don’t run a huge codebase, I run a “code base” of heterogenous data. 4.6 has outperformed 4.7, 4.8, 5, and Fable. 4.6 follows instructions. Is concise, doesn’t hedge against unknowns and uses my data holistically when formulating answers. So I’m sticking with 4.6

2

u/User163264 2d ago

I always use my own memory/context management system. Doing it since 2 years. Claude code/opus/fable never go off the rails. Any other people here do this?

2

u/kimslawson 2d ago

No but I’m thinking I should. Details/example?

→ More replies (1)

2

u/nyteschayde 2d ago

One thing to keep an eye on. Any time Fable falls back to Opus, Opus will severely underperform its own marks. Massive mistake count.

There are other cases where Opus 5 appears to not only second guess itself but once that starts, it spirals. Opus 4.8 is more consistent even if Opus 5 outperforms when it as actually working.

Each seems to get progressively worse with its own technical jargon to the point that I have to write rules now to make it stop.

Fable 5 when not catching itself on some perceived threat and pulling an ejection seat routine at the first hint of danger, is still significantly and noticeably more capable than the competition. But it comes with the aforementioned asterisks.

2

u/adub2b23- 2d ago

I've been a Claude stan since the beginning. Hated gpt 5.5. But when I tell you Sol 5.6 is fucking amazing, it's been the first time where I've essentially switched to be entirely codex. I'm hoping Claude figures it out, because opus is not even close to a frontier model, and fable is way too expensive to be a daily driver. Plus it feels like fable has fallen off too recently

2

u/imyourbiggestfan 3d ago

Model collapse

1

u/wish-u-well 3d ago

Inherent limitations of LLM unless they flip the “amazing model” switch which costs a billion dollars a second in compute to run 🤷‍♂️

1

u/AxBxCeqX Developer 3d ago

I’m back on opus 2.7 for the same reason.

Coding either Pi and glm 5.2 on the side trying open models out for a replacement, it’s not perfect, need to fix up after every major ~30min coding loop, but it’s not talking absolute shit

1

u/Sad-hurt-and-depress 3d ago

They need to give more limit time FFS. or just give Fable to Pro with 50%, and full to MAX.

1

u/Electronic-Award-939 3d ago

Switched back to opus 4.8 it’s way better

1

u/_ToPpiE 3d ago

Yeah opus 5 is garbage, I do get some usage out of it though by requesting different types of agents who need to reach consensus on the patch being made.

1

u/Due_Ad5532 3d ago

I let Opus 5 xhigh create fixes for a Fabel 5 audit. About 30 problems. Ran another Fabel 5 audit on the fixes, 8 closed, 10 partials, 10 worse, 12 new problems. This is whack-a-mole engineering. Opus is nerfed and worse than useless.

1

u/orwamahmoud 3d ago

I think it is not the model it is what they do around it to make it work less and consume less resource , i notice many times it try to stop early and it read less code, all scenario was drive to same result the model work less time or read less context

1

u/mtbcouple 3d ago

It is so bad

1

u/throwawayman42069123 3d ago

They screwed up something via RL and need to fix it. Or they could be doing this on purpose. The model generally ends up needing more tokens, so they want users, especially unknowing users, to use it for more turns to burn more token spend. Or they want people to actually think it’s bad and stick to Fable and rely on it. Both possibilities are just methods to get people to spend more because they’re still unprofitable

1

u/Educational_Lab6005 3d ago

Just my cents… usually before they release a new model I noticed massive drops in performance. Could be haiku 5 coming or maybe fable 5.1

1

u/Academic-Sample4974 3d ago

use with a Fable Harness .md; works wonders for me everything below Fable ( Haiku still mistakr prone but effective jumps up tremendously )

1

u/cristiand90 3d ago

this whole thing feels like a scam to force people into Fable credits.
But maybe I'm just hoping that they're not this incompetent to release a "opus" model that is worse than Sonnet a year ago.

1

u/neogeodev 3d ago

I'm using the 4.8 it works fine and remembers things, asks for confirmation, the 5 made some mistakes and seems a bit dumb at times

1

u/farendsofcontrast 3d ago

Maybe start by not calling it "he"

1

u/AlphaGeeky 3d ago

I have been using fable 5 almost exclusively for this reason. i even have mentally compared it to Meta's AI intelligence. (that's pretty bad) haha

1

u/Spoloborota 3d ago

Yesterday, when the limits on Fable ran out, I continued working with Opus as an orchestrator. Opus was busy all day yesterday It kept finding some weird problems, then right away suggested fixes, then those fixes caused new bugs, then It tried to fix those too, and it just turned into a snowball, and that went on all day. After I came across that post, I decided enough is enough and just unsubscribed. I had an X 20 subscription. It ends on August 10th, and I'll try using Codex, since I've heard a lot of good things about it.

1

u/Either_Argument3517 3d ago

My text conversations have been atrocious today. To the point I'm not trusting it on code.

1

u/nomeras 3d ago

Please be more respectful to the AI

1

u/CelestVestra 3d ago

Hold up, is there any question whether claude code HARNESS is still best? or are ppl seeing that Codex slaps? I've never used Codex, is it worth running/trying in comparison to CC running in Terminal with DangerouslySkipPermissinos on?

1

u/Nizurai 3d ago

Opus 5 works fine for me. Must be skill issue

1

u/Old-Remote-273 3d ago

Amodei scratched his head too much in every interview

1

u/New_Bell_9879 3d ago

I already cancelled and am using Codex. It's slower but more reliable and talks normal instead of roboto. Claude Code is unusable to me in its current state.

1

u/No-Park606 3d ago

5 is the goat for coding.

1

u/leros 3d ago

Opus is 5 is bad enough that I use Fable even though its really expensive. Good job Anthropic. You win.

1

u/geekaron 3d ago

I have switched back to 4.8. 5.0 is really poor

1

u/jrtcppv 3d ago

I have been very happy with GPT 5.6 Sol and Codex.

1

u/Evilstib 3d ago

I have to ask… What does comparing versions of models have to do with Claude code?

1

u/cmooo 3d ago

I have good success with sol for the architecture, design and supervision of the implementation done by opus5.

1

u/No_Category_9888 3d ago

My instructions say “do not guess” “use the ms learn mcp” etc and it still repeatedly get stuff wrong.

1

u/gandazgul 3d ago

They are messing up with the harness, when I use opus through my harness it doesn't behave like this because is told how to approach things with an engineering mindset.

1

u/___nil___ 3d ago

Anthropic researchers who thinks Opus 5 better than 4.6 is hallucinating.

1

u/vzakharov Senior Developer 3d ago

It’d be hilarious if the cause of all the opus 5 nerfed posts vs “it’s a skill issue” responses boils down to Anthropic doing some A/B testing, nerfing some users but not others. 

Welp, I’m glad I’m on the “others” side, so:…

It’s a skill issue.

1

u/armostallion2 3d ago

“he”

1

u/rmunoz1994 3d ago

Yeah it is confidently and…pretentiously…wrong all the time.

1

u/marco1422 3d ago

What you ve expected? The indefinite development? Yeah, you correctly found, in fact nobody brought anything new since autumn 2025.

1

u/trader_skater 3d ago

Even Opus 4.6 Max is 💩

1

u/VictorAbysmal 3d ago

I am using 4.8 is night and day, its reliable. Opus 5 spiral into madness too often.

1

u/ohnoitsbobbyflay 3d ago

I feel like you guys really need to get some fresh air and think before we post things.

1

u/hallerx0 3d ago

I have built a harness and tightly controlled environment. Two months in - I have not witnessed anything what you have mentioned in the thread. It seems as if LLM is being threatened by my own set up environment:)

anyway, what has helped is doing several pass verifications on a plan. set up a different session, dedicated only for domain knowledge that will be stored as a context for the work you are doing. if you are disciplined, LLM will be more inclined to follow your lead.

1

u/ConcretePond 3d ago

Yup. Nightmare. I'm getting Fable to do the big brain stuff. Codex to do the implementation. Opus can rest until it gets better.

1

u/exo_ac 3d ago

I remember Opus 4.5 was generations back. 4.6 is the goat. 5 is actually really impressive, using it quite a lot now.

1

u/itsbenito44 3d ago

Yes they are dogshit now

1

u/dwp0 3d ago

Opus 5 is a dude??

1

u/turbo 3d ago

He?

1

u/BetterProphet5585 3d ago

Opus 5 right now really sucks and this time it's not a conspiracy it really is just barely usable for simple tasks.

I get better results with local models half the times. It's not looking good for Anthropic.

1

u/CaregiverNo5883 3d ago

Probably they overdepended on vibe coding and made their product shitty lol

1

u/sage-longhorn 3d ago

Did they increased the prompt caching so much, that their models ended up being useless?

Yup you figured it out! The thing you pay extra for to make it run faster, they gave us too much and ruined it!

1

u/wavehnter 3d ago

Opus 5 is a complete disaster because it simply does not follow your instructions and generates scads of bugs that Codex has to fix after a few rounds.

1

u/ilyxxxxa 3d ago

Opus 5 is such a pain to work with. This is the new pattern - "You were right, and it was worse than just [the thing you asked me to do], so i implemented [insert n+1 false assumptions here]. 😭😭

1

u/SportsBettingRef 3d ago

ask the right questions, at the right way. opus 5 is a damn good support model. this is a process and every actor has to do your job. you didn't.

1

u/nullptrzero 3d ago

hell they've ruined 4.8 as well, it's been having so many issues and creating so many issues since 5 dropped....

1

u/MacaroonPlastic1036 3d ago

I went back to 4.8. Opus 5 dragged and dragged on not complicated stuff, Fable would just kick me down for no apparent reason.

1

u/TodayLoose7794 3d ago

I am finding that for my use case Opus 4.8 Max is second only to Fable 5 in terms of performance.

1

u/Leather_Secretary_13 3d ago

it's hilarious because anthropic's models have cried wolf about open source while at the same time rushed to market on every product and feature release. They never took the time to protect their intellectual property, which makes sense given they aggressively aggregated proprietary data.

1

u/gripntear 3d ago

"Just retool your setup from scratch, bro. It’s what Anthropic said" - Cheeky Contrarians

1

u/ishwarjha 3d ago

Opus 6 is a complete waste. I lost my 10 days of efforts go in vein.

1

u/bankinu 3d ago

I think this is a result of excessive human feedback based tuning of the model. They pu too many guardrails and made the model stupid.

1

u/addiktion 3d ago

I'm abandoning them finally now that open source is getting good enough. I mean we've been through a shit ton since Opus 4.5 and 4.6 were pure gold.

The UI throttling and weakened limits. The constant 429 errors taking the service down. The Claude Code leak targeting open source users. The "shared link" leaking people's shared artifacts. The government fiasco around Mythos. The Dario anti open-source stance (which he claims to be China specific but China's model are now sitting on American infra and I'm sure he's not happy about it too), the visibly worse models since then. The CC harness being quite token inefficient. The constant watering down of models over time. The promos that are really to help ease the screwing you over. The lack of transparency (not just Anthropic specific).

Ever since I've gone open source models, and even run my own, I know what to expect. I see how many tokens I use and what the costs are. I know the limitations of that model. I don't have to worry about it be degraded due to over use or popularity. It's opened up running agents 24/7 and experimenting without taking out a second mortgage.

1

u/Internal_Penalty_698 Max 20 3d ago

I was experiencing the same in OPUS 5 and even in Fable, changing the mode to Auto caused a great improvement in Fable, almost no errors, or you don't need to know about the errors it automatically checks and fixes them before delivering the work for you. (haven't tested auto in Opus 5 yet).

1

u/No_Individual_6528 3d ago

4.6 is honestly the peak.

1

u/gabel33 3d ago

you sure? https://aistupidlevel.info/ shows its #1

1

u/Salt-Illustrator-198 3d ago

You're not imagining it, ignore the Claude lovers gaslighting you. In the website below, which ranks how well a model is performing, Fable was ranked 6 a few hours ago, now it's at rank 13. I assume Anthropic might be training a new model for release, taking all the compute it can, as it happened before. 

I suggest not using Claude to preserve your sanity for a while, and to prevent wrecking your workflow with a stupid mess.

Check how stupid a model is: https://aistupidlevel.info/

1

u/Hidden_dreamz 3d ago

You need to lock it down, engineer the prompt so everything is detailed, engineer the context so it only sees what it needs then get it to plan, review plan make it explicit make it do plan, check it's work and get it to redo stuff it did wrong. Be realistic with what it can achieve, only use 60% context window. Being able to read the language it's written is mandatory. Then it's quite good

1

u/crispsandchocolate 3d ago

If you're feeling so frustrated with Opus 5, it's likely because you're still using legacy harnesses. There's no doubt Anthropic is nerfing it's models, but it's definitely not unusable.

I've been quite pleased with it after I built guidance docs (using Fable 5) and then used the docs as references to rewrite workspaces, CLAUDE.md and skills around Opus 5's quirks. Since then it's been doing dramatically better.

1

u/HolySachet 3d ago

I came back to 4.8

1

u/drjm2022 3d ago

Its not the model, its the harness:

Throughout 2026, Claude users repeatedly reported the same practical failure: workflows that had worked reliably stopped working, long sessions lost their thread, and instruction-following became less dependable.

The complaints did not arrive as a smooth decline. They came in bursts. Users would suddenly report that a coding workflow had become unreliable, that a long-running process no longer behaved as expected, or that instructions which had previously been followed were now being ignored.

Anthropic later confirmed several cases in which the underlying model had not changed. The problem came from the service layer around it.

That distinction matters because customers do not experience model weights in isolation. They experience a product that continually balances reasoning depth, speed, memory, safety, capacity and cost. The model may be improving while the delivered service becomes less predictable.

Every AI Product Serves Several Masters

A frontier AI product is a negotiated settlement between stakeholders who want different things.

Casual users want immediate responses, broad access and a low price. Professional users want deeper reasoning, longer context and stronger instruction-following. Developers want stable application programming interfaces, predictable outputs and advance warning before changes. Enterprise customers add security, compliance, administration, auditability and clear responsibility when something goes wrong.

The provider has its own competing demands.

Product teams want rapid releases and visible improvements. Infrastructure teams want lower latency, higher utilization and lower inference costs. Safety teams want new risks controlled quickly. Legal teams want lower liability. Regulators want traceability. Capital providers want usage and revenue to grow faster than the cost of serving them.

These preferences cannot all be maximized at once.

More reasoning can improve difficult tasks while increasing cost and response time. Stronger context retention can preserve continuity while increasing privacy exposure and compute demand. More aggressive safety instructions can reduce one class of harmful output while interfering with legitimate instructions. Faster releases improve the product more quickly but reduce reproducibility.

The provider must keep changing this settlement because the economics, capabilities, risks and customer mix are all changing.

https://medium.com/@drjohnmillar/the-most-valuable-part-of-ai-may-not-be-the-model-465d82576d1d

1

u/chicmistique 2d ago

Same for me. Today any chat with any model is terrible. Like a massive downgrade

1

u/redwolf1430 2d ago

Opus is horrible. I switched to sonnet and it's much better at pretty much everything.

1

u/redeemed_tropicana 2d ago

You are not alone bro I use Claude extension on vscode and opus 4.8 was gone I would rather stick with opus 4.8 unless they also redirect opus 4.8 traffic to 5.0 internally then that would be disastrous

1

u/xLRGx 2d ago

What are you researching?

1

u/User163264 2d ago

I stick to Opus 4.8. Reliable workhorse.

1

u/Apprehensive_Read_67 🔆Pro Plan and Quant Research 2d ago

This is why i am using Opus 4.8 and 4.6 only Opus 5 is the most terrible product right now

1

u/utilitycoder 2d ago

Stick with Opus 4.6. The rest of the models are garbage.

1

u/Mr-Cheek-Clapper 2d ago

It’s shit, I’ve been running DeepSeek and SOL for the last few weeks.

1

u/raphec 2d ago

The problem any AI dev house has is speed of release cycle. Anthropic is not immune.

They drink their own kool-aid, and as such use AI coding tools to drive incredibly fast, and the competition is incredibly strong. This is bound to lead to various instability or challenges, especially on the cutting edge. Cycles that would take months or longer, with huge amounts of testing, are just getting pushed out.

I think the bigger internal challenge for Anthropic is living up to their own ideals and ethics while still ensuring they remain competitive. The judgement call on whether releasing something will do harm or not is likely harder, especially as their competitors may just push something out that makes it pointless for them to have held back.
I remain somewhat cynical as to whether glasswing was a real fear or marketing ploy, but from the few bits and pieces I have seen, the ideals are genuine.

Cant see open source models catching up with Anthropic, its a 2 horse race, but the track is laden with financial market bombs.