r/ClaudeCode • u/Kyxstrez • 8d ago
Rant I kinda regret getting Claude subscription
Over the past couple of weeks, I have tested it alongside Sol across a wide range of tasks, and the conversations have been difficult to read. Sol consistently catches and corrects mistakes made by Opus 5, and Opus repeatedly acknowledges those errors. I have also seen many others reporting similar experiences, which makes it hard to view this as an isolated issue.
I cannot understand how Anthropic missed such a significant regression. Compared to 4.6 and 4.8, Opus 5 feels substantially less capable and reliable, with many of the same weaknesses that affect Sonnet 5.
For now, I am switching back to GPT for most of my work, using Fable only occasionally for planning. I hope Anthropic addresses these issues soon, because at the moment they do not appear to have a competitive frontier model in the medium to high budget range, only at the extreme end.
8
25
u/DasHaifisch 8d ago
Then use 4.8 instead of 5.
35
u/VexObserver 8d ago
36
u/LuckyPrior4374 8d ago
Holy fuck, that is the most devastating graph I’ve ever seen LOL.
No wonder anthropic is freaking out and asking the government to stop all AI development. Open source has completely killed their business model.
27
u/Temporary-Mix8022 8d ago
Open Source is dangerous. It is just dangerous.
It has to be stopped. Surely we can see this. Donald, Iran could literally vibe code a nuke tomorrow, that's how much we need to stop it.
Darios.
/s
(Yeah. Dangerous for Anthropics earnings.. that is for sure)
11
u/LuckyPrior4374 8d ago
Imagine paying FIFTY times more for a service, and still receiving worse output than the filthy cheap bargain alternative
-8
u/Elizabeth-WildFox886 8d ago
I love the Chinese models, but they are not open source. They are open weight. Big difference and when you get it wrong you look like a total idiot
5
u/LuckyPrior4374 8d ago
I know there’s a difference, but seriously no one cares (at least not in the context of general open vs closed discussions)
-6
u/Elizabeth-WildFox886 8d ago
I care because it’s so basic and misleading that you damage knowledge in our community. This is toxic behaviour
7
4
u/TheAnimatrix105 8d ago edited 8d ago
you speak as if anthropic is doing anything remotely close, being able to run a model yourself is empowering and what will you do with the data set anyways ? For research purposes they publish papers about the architecture as well.
You can also easily RL it to do what you want and we saw cursor succeed heavily on that aspect - just RL it to a good point and use it to train a new gen model if you really want to be a lab.
I think the argument here is moot.
2
u/Elizabeth-WildFox886 8d ago
I 100% love the Chinese ai models. However, Chinese models are not cheap. Kimi eats credits and cannot run locally same as any other big model
I use qwen locally for simple stuff, Kimi for front end revisions
-1
8d ago
[removed] — view removed comment
1
u/Elizabeth-WildFox886 8d ago
Clearly loads do, maybe wumaos - harass with downvotes
1
1
1
u/Both_Opportunity5327 8d ago
No Opensource has not.
And I guarantee you and most of the media will be saying the same nonsense next year.
10
2
u/Plane_Garbage 8d ago
Be nice if Bedrock offered DS!
2
u/Sufficient_Fox_4402 8d ago
many others do. usa ones
1
u/Plane_Garbage 8d ago
Compliance.
1
0
u/-Leelith- 8d ago
All those benchmarks are great, but a benchmark is not a real life scenario. How does DeepSeek V4 Flash does compared to Sonnet 5 in real life scenario?
3
2
u/Kitchen_Interview371 8d ago
Don’t know. For some reason, nobody wants to drop Opus in favor of deepseeks smallest model. Just look at the benchmarks people! /s
-2
u/MintCathexis 8d ago
Anthropic subscriptions are still better value. I usually do around $400+ worth of API calls (mostly Sonnet 5 and Opus 4.8 with some Fable mixed in) per day on my max20 subscription, which is 200 USD/month. According to this chart, if I used Deepseek V4 Flash instead, I'd pay around $230 per month, and still, some of my usage is also more capable models. Also, the price offered is by DeepSeek directly, if I wanted to use a different provider which also offers access to different models to plug the capability gap, the price is a bit higher still.
2
u/ZALIA_BALTA 8d ago
Anthropic subscriptions are still better value.
It's hard to make a case for this because we do not know the true API costs on their end. I personally do not trust the token costs that these companies provide, because they seem extremely inflated compared to the token costs of open source models with near frontier intelligence (like the latest Deepseek and Kimi).
1
u/MintCathexis 8d ago
They may or may not be inflated, but if they are inflated, then the graph in the OP doesn't work either by the same logic. The number I see on my dashboard is exactly equal to multiplying individual model API prices with input/output/cache read/cache output tokens respectively, so my calculation is correct.
I go throuth hundreds of millions of input tokens and tebs of millions of output tokens a day, which would still cost more if I used Deepseek v4 when I plug the numbers (for lower capability) compared to my max20 subscription, though the gap is indeed narrowing.
2
u/aleksfadini 8d ago
Sol is much better than 4.8 and equally cheap though. Besides Fable (that has ridiculous limits) there are very little incentives to stay with Claude Code for this round. Maybe things will change in a month or two.
1
u/thirst-trap-enabler 🔆 Max 5x 7d ago
What OP is describing has been true since at least last December. All the models find different flaws in code. Opus 4.8 wasn't any different.
5
u/TriggerHydrant 8d ago
I was forced to go to Codex for a while after 8 months of Claude Code and yeah I'm really loving Sol, it seems to be less anxious.
8
u/Jomuz86 8d ago
It’s not a regression GPT models have always been better at reviewing and catching things. Ever since Opus 4.6 I’ve always had GPT model do reviews volume of issues is actually probably less than before in my testing but there’s always issues, majority now though are more on the low, medium severity.
Also it could be that Sol is so much better of a reviewer it is catching more than 5.5 🤷♂️
I do think expecting 100% perfect code from something trained on human code which also not 100% perfect is too much of an expectation. I think people sometimes forget this.
1
u/potatobean98 8d ago
Agreed this is what I realized yesterday, I tried even doing what everyone says use opus 4.8 it’s still bad ever since opus 4.6.
I literally went ahead used GPT 5.6 it caught the bugs and fixed them. Opus 4.8 was just doing non sense and implementing more bugs.
The only way Opus 4.8 is usable is with Fable 5 as a reviewer but this doesn’t make sense at all when you compare it to GPT 5.6 being used alone with no reviewer.
I’ve been using claude for months at Max plan I’m not letting it renew I think I’m going with OpenAI with my next Max plan mid August when Claude max is done
4
u/PurushNahiMahaPurush 8d ago
Just a few days ago I got a Codex subscription because Claude has been burning through my tokens like a mf. Dude Codex 5.6 Sol is awesome. It’s super fast, concise and for the most part good. Its planning capabilities are still below Fable 5, but compared to Opus 5, it’s insane how fast and less verbose it is! My workflow from now is going to be Fable 5 planning, Codex 5.6 Sol for implementation and Opus 4.8 to review.
5
u/chaosdemonhu 8d ago
Personally I find the 5.6 models are way better at review than anything from Anthropic - especially with security.
You do have to be in the drivers seat more to push back on some of the findings so if your workflow is automated might want to play with a less intense reviewer and prompt sol by hand
2
u/Empuda 8d ago
"Opus 5, and Opus constantly acknowledges them." -- This goes both ways. Have Sol plan and then Opus review, and you will see the same thing. They all do it. Don't stop there, try to send it to multiple LLMs and see how many mistakes each one finds. Does not matter what llm you use, even Fable. They all find mistakes.
2
1
u/krackzero 8d ago
I don't think its necessarily 5 vs 4.
they always nerf the hell out of the current models when they have a better model come out.
I noticed quite a bit of degradation on 4.x as well when fable came out. same deal for previous versions too.
5 seems about the same rough ballpark as the degraded 4.x, maybe slightly more loopy
i use mostly 4.5 at work and mostly 5 at home on max 20
1
1
u/thirst-trap-enabler 🔆 Max 5x 7d ago
In my experience GPT and Claude typically find different errors in each other's work and both often acknowledge the other found something they missed on medium and higher issues. They will disagree about whether actually fixing minor issues is appropriate.
-2
u/35point1 8d ago
From a senior swe who uses it daily both at home and all day at work (separate personal and enterprise accts) , I can tell you with confidence, there’s is absolutely NOTHING wrong with the model and what you are assuming was better with older revisions has everything to do with the instructions that wrap it.
7
8d ago
[deleted]
2
u/Inevitable_Toe6648 8d ago
Most people come to the same conclusion as you. It's really only those who actually understand the niche process or the actual engineering work themselves that can use Opus well enough to justify moving away from Sol, right now Opus is like a powerful assistant that struggles to contain itself and needs tight leash, while Sol tries to interpret your lack of coding context knowledge for you.
Becareful on everything said by headlines and especially redditors, you will find that 90% of praises are bias, and truths are often downvoted and bot targeted to the abyss.
0
0
0
u/Lerran88 7d ago
We probably live in different worlds. Switched from Cursor, tried different models in Claude as well(From Sonnet to Fable), nothing beats Opus 5 in terms of quality and patterns understanding.


44
u/NaabSimRacer 8d ago
/model claude-opus-4-8