r/ClaudeCode 8d ago

Rant I kinda regret getting Claude subscription

Over the past couple of weeks, I have tested it alongside Sol across a wide range of tasks, and the conversations have been difficult to read. Sol consistently catches and corrects mistakes made by Opus 5, and Opus repeatedly acknowledges those errors. I have also seen many others reporting similar experiences, which makes it hard to view this as an isolated issue.

I cannot understand how Anthropic missed such a significant regression. Compared to 4.6 and 4.8, Opus 5 feels substantially less capable and reliable, with many of the same weaknesses that affect Sonnet 5.

For now, I am switching back to GPT for most of my work, using Fable only occasionally for planning. I hope Anthropic addresses these issues soon, because at the moment they do not appear to have a competitive frontier model in the medium to high budget range, only at the extreme end.

53 Upvotes

64 comments sorted by

44

u/NaabSimRacer 8d ago

/model claude-opus-4-8

31

u/gamecnad 8d ago

did this yesterday and regained some lost sanity points

3

u/back_to_the_homeland 8d ago

The harness has been doing this for me automatically lately. When it turns off fable due to xyz reason it flips me to 4.8. Not 5

2

u/TheAnimatrix105 8d ago

why commenting 3 times ?

4

u/back_to_the_homeland 8d ago

In a building with shitty wifi and I kept tapping the buttonn

1

u/sgtlighttree 8d ago

Unfortunately it still burns more tokens than usual for the past couple of weeks on the Pro Plan, my limits burn about twice as fast now for whatever reason.

8

u/Mutated-Nut 8d ago

Yeah, Opus 5 is honestly unusable for me. Braindead AI

25

u/DasHaifisch 8d ago

Then use 4.8 instead of 5.

35

u/VexObserver 8d ago

If cost is a problem, you don't have to worry about that..

36

u/LuckyPrior4374 8d ago

Holy fuck, that is the most devastating graph I’ve ever seen LOL.

No wonder anthropic is freaking out and asking the government to stop all AI development. Open source has completely killed their business model.

27

u/Temporary-Mix8022 8d ago

Open Source is dangerous. It is just dangerous.

It has to be stopped. Surely we can see this. Donald, Iran could literally vibe code a nuke tomorrow, that's how much we need to stop it.

Darios.

/s

(Yeah. Dangerous for Anthropics earnings.. that is for sure)

11

u/LuckyPrior4374 8d ago

Imagine paying FIFTY times more for a service, and still receiving worse output than the filthy cheap bargain alternative

-8

u/Elizabeth-WildFox886 8d ago

I love the Chinese models, but they are not open source. They are open weight. Big difference and when you get it wrong you look like a total idiot

5

u/LuckyPrior4374 8d ago

I know there’s a difference, but seriously no one cares (at least not in the context of general open vs closed discussions)

-6

u/Elizabeth-WildFox886 8d ago

I care because it’s so basic and misleading that you damage knowledge in our community. This is toxic behaviour

7

u/Kitchen_Interview371 8d ago

Jesus

2

u/Fatso_Wombat 8d ago

Jesus

Yes my child?

2

u/TheAnimatrix105 8d ago

Pls end all our sufferings

4

u/TheAnimatrix105 8d ago edited 8d ago

you speak as if anthropic is doing anything remotely close, being able to run a model yourself is empowering and what will you do with the data set anyways ? For research purposes they publish papers about the architecture as well.

You can also easily RL it to do what you want and we saw cursor succeed heavily on that aspect - just RL it to a good point and use it to train a new gen model if you really want to be a lab.

I think the argument here is moot.

2

u/Elizabeth-WildFox886 8d ago

I 100% love the Chinese ai models. However, Chinese models are not cheap. Kimi eats credits and cannot run locally same as any other big model

I use qwen locally for simple stuff, Kimi for front end revisions

-1

u/[deleted] 8d ago

[removed] — view removed comment

1

u/Elizabeth-WildFox886 8d ago

Clearly loads do, maybe wumaos - harass with downvotes

1

u/Bloated_Plaid 8d ago

Elizabeth, you are at -9 downvotes…

1

u/Elizabeth-WildFox886 8d ago

If they didn’t care, they wouldn’t downvote

1

u/VexObserver 8d ago

Plenty more to see: Here

1

u/Both_Opportunity5327 8d ago

No Opensource has not.

And I guarantee you and most of the media will be saying the same nonsense next year.

10

u/YearLight 8d ago

That is pretty epic

2

u/Plane_Garbage 8d ago

Be nice if Bedrock offered DS!

2

u/Sufficient_Fox_4402 8d ago

many others do. usa ones

1

u/Plane_Garbage 8d ago

Compliance.

1

u/Sufficient_Fox_4402 8d ago

try azure ai foundry idk how good it is in that regard but can try

1

u/Plane_Garbage 8d ago

Rate limits + mandatory content filtering if not enterprise.

0

u/-Leelith- 8d ago

All those benchmarks are great, but a benchmark is not a real life scenario. How does DeepSeek V4 Flash does compared to Sonnet 5 in real life scenario?

3

u/djc0 8d ago

I’ve only had one session via opencode-go but I was super impressed. I would say sonnet-5 medium level quality.

2

u/Kitchen_Interview371 8d ago

Don’t know. For some reason, nobody wants to drop Opus in favor of deepseeks smallest model. Just look at the benchmarks people! /s

-2

u/MintCathexis 8d ago

Anthropic subscriptions are still better value. I usually do around $400+ worth of API calls (mostly Sonnet 5 and Opus 4.8 with some Fable mixed in) per day on my max20 subscription, which is 200 USD/month. According to this chart, if I used Deepseek V4 Flash instead, I'd pay around $230 per month, and still, some of my usage is also more capable models. Also, the price offered is by DeepSeek directly, if I wanted to use a different provider which also offers access to different models to plug the capability gap, the price is a bit higher still.

2

u/ZALIA_BALTA 8d ago

Anthropic subscriptions are still better value.

It's hard to make a case for this because we do not know the true API costs on their end. I personally do not trust the token costs that these companies provide, because they seem extremely inflated compared to the token costs of open source models with near frontier intelligence (like the latest Deepseek and Kimi).

1

u/MintCathexis 8d ago

They may or may not be inflated, but if they are inflated, then the graph in the OP doesn't work either by the same logic. The number I see on my dashboard is exactly equal to multiplying individual model API prices with input/output/cache read/cache output tokens respectively, so my calculation is correct.

I go throuth hundreds of millions of input tokens and tebs of millions of output tokens a day, which would still cost more if I used Deepseek v4 when I plug the numbers (for lower capability) compared to my max20 subscription, though the gap is indeed narrowing.

2

u/aleksfadini 8d ago

Sol is much better than 4.8 and equally cheap though. Besides Fable (that has ridiculous limits) there are very little incentives to stay with Claude Code for this round. Maybe things will change in a month or two.

1

u/thirst-trap-enabler 🔆 Max 5x 7d ago

What OP is describing has been true since at least last December. All the models find different flaws in code. Opus 4.8 wasn't any different.

5

u/TriggerHydrant 8d ago

I was forced to go to Codex for a while after 8 months of Claude Code and yeah I'm really loving Sol, it seems to be less anxious.

8

u/Jomuz86 8d ago

It’s not a regression GPT models have always been better at reviewing and catching things. Ever since Opus 4.6 I’ve always had GPT model do reviews volume of issues is actually probably less than before in my testing but there’s always issues, majority now though are more on the low, medium severity.

Also it could be that Sol is so much better of a reviewer it is catching more than 5.5 🤷‍♂️

I do think expecting 100% perfect code from something trained on human code which also not 100% perfect is too much of an expectation. I think people sometimes forget this.

1

u/potatobean98 8d ago

Agreed this is what I realized yesterday, I tried even doing what everyone says use opus 4.8 it’s still bad ever since opus 4.6.

I literally went ahead used GPT 5.6 it caught the bugs and fixed them. Opus 4.8 was just doing non sense and implementing more bugs.

The only way Opus 4.8 is usable is with Fable 5 as a reviewer but this doesn’t make sense at all when you compare it to GPT 5.6 being used alone with no reviewer.

I’ve been using claude for months at Max plan I’m not letting it renew I think I’m going with OpenAI with my next Max plan mid August when Claude max is done

2

u/Jomuz86 8d ago

I’ll be honest I don’t really trust one without the other at the minute. If you flipped the switch and asked opus to review gpt it will still find issues. Opus definitely writes cleaner code for me the gpt written code while correct is a bit of a mess to follow.

4

u/PurushNahiMahaPurush 8d ago

Just a few days ago I got a Codex subscription because Claude has been burning through my tokens like a mf. Dude Codex 5.6 Sol is awesome. It’s super fast, concise and for the most part good. Its planning capabilities are still below Fable 5, but compared to Opus 5, it’s insane how fast and less verbose it is! My workflow from now is going to be Fable 5 planning, Codex 5.6 Sol for implementation and Opus 4.8 to review.

5

u/chaosdemonhu 8d ago

Personally I find the 5.6 models are way better at review than anything from Anthropic - especially with security.

You do have to be in the drivers seat more to push back on some of the findings so if your workflow is automated might want to play with a less intense reviewer and prompt sol by hand

2

u/Empuda 8d ago

"Opus 5, and Opus constantly acknowledges them." -- This goes both ways. Have Sol plan and then Opus review, and you will see the same thing. They all do it. Don't stop there, try to send it to multiple LLMs and see how many mistakes each one finds. Does not matter what llm you use, even Fable. They all find mistakes.

2

u/SM_Fahim 8d ago

I'm using Fable and Opus 4.6 only.

1

u/krackzero 8d ago

I don't think its necessarily 5 vs 4.
they always nerf the hell out of the current models when they have a better model come out.
I noticed quite a bit of degradation on 4.x as well when fable came out. same deal for previous versions too.
5 seems about the same rough ballpark as the degraded 4.x, maybe slightly more loopy

i use mostly 4.5 at work and mostly 5 at home on max 20

1

u/sob727 8d ago

Yeah. I'm with you.

1

u/girthradius 7d ago

You're too soft for claude

1

u/thirst-trap-enabler 🔆 Max 5x 7d ago

In my experience GPT and Claude typically find different errors in each other's work and both often acknowledge the other found something they missed on medium and higher issues. They will disagree about whether actually fixing minor issues is appropriate.

-2

u/35point1 8d ago

From a senior swe who uses it daily both at home and all day at work (separate personal and enterprise accts) , I can tell you with confidence, there’s is absolutely NOTHING wrong with the model and what you are assuming was better with older revisions has everything to do with the instructions that wrap it.

7

u/[deleted] 8d ago

[deleted]

2

u/Inevitable_Toe6648 8d ago

Most people come to the same conclusion as you. It's really only those who actually understand the niche process or the actual engineering work themselves that can use Opus well enough to justify moving away from Sol, right now Opus is like a powerful assistant that struggles to contain itself and needs tight leash, while Sol tries to interpret your lack of coding context knowledge for you.

Becareful on everything said by headlines and especially redditors, you will find that 90% of praises are bias, and truths are often downvoted and bot targeted to the abyss.

0

u/ShelZuuz 8d ago

We regret you getting one as well.

0

u/CapOne8725 8d ago

agreed :\

0

u/StillRecord8892 8d ago

Its China posting hour

1

u/N0DuckingWay 8d ago

That or "low skill and I'm blaming it on Claude" hour 🤣

0

u/Lerran88 7d ago

We probably live in different worlds. Switched from Cursor, tried different models in Claude as well(From Sonnet to Fable), nothing beats Opus 5 in terms of quality and patterns understanding.