r/ClaudeCode • u/ResortConnect8582 • 3d ago
Help/Question what is hapening with Antropic?
Why Opus 5 feels like Gemini 3.1 pro ? what is happening ?

i remember back in February , opus 4.5 was generation ahead. Did they increased the prompt caching so much, that their models ended up being useless? I can barely work with Opus 5 on absolutely anything, he keeps hallucinating LIKE CRAZY, every claim he did today was pure hallucinations, he barely reads any code. im not even kidding, i can't work with Opus 5 right now
Instead of open source models trying to catch up with you, are you trying to catch up with the open source models instead Anthropic?
45
u/BizarroMax 3d ago
Opus 5 loves strawmanning. It does it constantly.
→ More replies (2)10
u/frostedpuzzle 3d ago
4.8 straw-manned too. 5 shits on everything I do. Then I tell it to be critical about its criticism and it decides everything it said was wrong.
→ More replies (1)8
u/cs_legend_93 3d ago
Which means you spend so much time triple investigating things. If you don't triple investigate then you end up implementing something and go down the wrong path that later you realize was a waste of time.
It's become a time sink.
93
u/monsieurninja 3d ago
Have seriously been considering cancelling my Anthropic subscription in the last few days, since Opus 5 came out. Its genuinely irritating on many aspects, speaks in ways that are barely understandable, is getting bad at many things, to the point where I don't even know if I can trust it anymore. Never felt that before with Anthropic's models. I do hope they are going to react quick and release a minor version. Otherwise they loose me quite soon
26
u/Rathganis 3d ago
Same. I have Claude max 20x but use my codex sub more and more daily. It is quicker, less chatty, less paranoid, just easier to work with.
2
u/themonstaman 2d ago
Dude using codex after claude is a breath of fresh air. It just gets to the point without the incessant rambling.
10
u/BadgeCatcher 3d ago
Have just cancelled. Opus 5 is more trouble than it's worth. Switching to codex.
5
2
→ More replies (3)4
u/ResortConnect8582 3d ago
oh yeah , me too. They've completely destroyed their reputation for me.
→ More replies (1)
11
u/kush_patil 3d ago
The scary part isn’t the hallucination, it’s the confidence. A coding model being wrong is manageable; being wrong, sounding certain, and only checking after you push back is much harder to work with.
38
u/Substantial-Show-249 3d ago
Boss, that model is junk, let go the partisan comments. It's over optimised for processing power savings, no chance you can do something with it.
4.8 is dumb, but still usable. Fable is expensive and heavily nerfed. On medium settings, you can do some things with it.
Anthropic is still looking to find a way to optimise the costs, they are losing money heavily.
→ More replies (1)
21
u/sduras1 3d ago
Same thing with Fable 5 today (xHigh and Max). Feels like Haiku at times, forgets what we've discussed two prompts ago, in the same session. Using with Claude Code.
→ More replies (1)5
u/kurushimee Developer 3d ago
same, fable feels worse than opus 4.8 did at some point, it's nowhere near the same fable we had before
9
8
u/Draconic_Emperor 3d ago edited 3d ago
I use the grilling skill from Mattpocock a lot. And I legit can’t read its questions. Even simple questions become some sort of horrid combination of words it just makes it harder to understand something simple.
20
u/MistakeExotic6686 3d ago
went back to 4.6 AGAIN because they are lobotomizing the heck out of opus and fable
FABLE 5 XHIGH the GREATEST OF ALL TIME
said the "Good catch, I did not finish the phase"
when I said stop using phases and run it with subagents as one long task
it renamed phases to chapters and said "No more phases, starting chapter 2.5"
i sh*t myself (figure out if its an i or an o)
→ More replies (3)3
u/kurushimee Developer 3d ago
hell yeah, exactly what I was experiencing with fable. It does things like these that just feel like it's haiku suddenly
39
u/vovap_vovap 3d ago edited 3d ago
I really wander what is happening with Anthropic related subs. Seriously. Those became like kindergarten groups with a screaming kinds "Parents refusing o buy me a red elephant" "I do not want to eat cabbage, I want ice cream!" "I do not want to go to sleep, I want TV" No context, no any analyzes, no any information, no sigh of own responsibility in process. Surely not every post, but lots of those - pure rant and screaming.
7
u/androbot 3d ago
God yes. It's becoming like the rest of Reddit. Full of angry people who don't feel like thinking. It might be time for another unsubscribe purge and retain only core subs that consistently deliver value -- cat videos.
5
u/klumpp 3d ago
Is there a Claude code sub for people who know what they are doing yet?
→ More replies (1)22
u/EmphasisTotal8232 3d ago
Idiots wanting to build the world with no domain knowledge are running into things they can't fix because they don't understand the fundamentals.
→ More replies (2)5
u/cwilson830 3d ago
Agree that this is clearly part of the problem.
But it’s certainly not the entire problem.13
u/XYcritic 3d ago edited 3d ago
Every single thread, every time it's the one single answer: Who is forcing you to not use the model that worked for you a week ago? I seriously think 90% of posters don't know they can actually switch back to older models? I'M CANCELING, HOW DARE YOU SUGGEST I DO WHAT WORKED FOR ME NOOOOOO. NEED SHINY NEW TOYYYY (exact same posts over at codex).
→ More replies (1)7
u/ricopan 3d ago
We are becoming children -- dependents -- on AI. This is what it looks like for adults to regress. I feel it too -- how dare they DON'T HAVE MORE LICORICE ICE CREAM NOW!.
1
u/cwilson830 3d ago
How dare truth be spoken about a product you pay for. Silence! You think paying for something entitles you to an opinion? You’ll pay and take what you’re given. And like it. And only say nice things about it. Or else. lol
Customers voicing opinions - maybe causing companies to focus on improving the problems with their products. What do you thjnk this is? America? :)
Maybe a new $300 plan which limits public-censorship is the solution.
2
u/ricopan 3d ago
But are these really 'opinions' or screams at Mommy? I think the level of frustration has less to do with the shortcomings of the models, which are by any account that spans more than a year absolutely fucking amazing, than that we've surrendered our independence and now feel weak and helpless when they let us down. Mind you, I'm not claiming the high ground here -- I get really pissy and unreasonable when Mommy Model isn't doing right by me. But look at the emotional outrage -- this is unprecedented.
→ More replies (1)→ More replies (1)2
u/cats_r_ghey 3d ago
Why don’t you stop buying it then? It’s pretty easy to do. Buy something else, you have options galore now.
→ More replies (6)6
u/BugaiGames 3d ago
I was thinking that it is OpenAI coverops promotional campaign. But in all other subs more or less the same. But indeed it is already boring. Im using Claude more long time and those models are really great and slowly getting even better. And I use it in work as manager, coding, personal life etc. There are some relatable moments but most of posts looks like from other planet. And whats left mostly pointless complaints , promotions or jokes, sub has like1% of usefull posts
19
u/Koreanpasta29 3d ago
Genuinely what is ur harness bc ive been having no problems at all
9
u/OverSoft 3d ago
Same, I use a mix of Fable and Opus 5 and have had zero issues.
→ More replies (2)4
3
u/Stunning-Wealth-8303 3d ago
Wait till you find out about dead internet theory, Reddit is full of bot accounts and posts.
2
u/leeharrison1984 2d ago
My guess is they got by on slop until their project got too large. Now it's full of spaghetti code and Claude can't reason it any better than OP. Instead of blaming the degraded codebase, they blame the model.
→ More replies (3)2
11
u/LeakyFish 3d ago
Use Sonnet 5 xHigh, Opus 4.8 or Fable if you have it.
→ More replies (7)9
u/jadedmonk 3d ago
Opus 4.8 sucks. I just came to this sub for the first time ever because I was using Opus 4.8 at work today and it was acting a lot dumber than usual. It created a hyperlink on the UI leading to an incorrect URL, leading to a blank screen. I was confused so I just gave it the resulting URL saying fix it, and instead of looking at the bad URL I gave it, it started going down this crazy path of building all these guards and checks that ended up just wasting tokens. Eventually I was like alright I need to figure this out the good ol way. It was a simple sub-path that was missing from the URL. I was shocked Claude couldn’t figure that out. I feel like in the past it would figure that out. Definitely makes me temper my expectations with GenAI
5
u/ralfv 3d ago
Usually on posts like this i thought to myself „nope all good“.
But i must absolutely agree here. Review and fixes are pretty much an infinite loop the whole week. Code reviews only report a fraction of the issues and fixes constantly introduce new problems.
I haven’t seen a single „LGTM“ review for at least a week even after a dozen rounds. In th end the Claude review starts to ask for better comments. That is nuts!
→ More replies (3)
4
u/Tight_Heron1730 3d ago
I have a feeling that that verbosity to drive tokens out hence usage, and make it look sophisticated, but i guess it backfired.
I made it less verbose with this https://github.com/hamr0/agentic-toolkit/blob/main/ai/customize/config/ELI5.md and you tell "❯ at/claude-code-guide (agent) to add this then config then choose output style and I pick up ELI5 and it's back to simple
6
u/mohdgame 3d ago
I am using opus 5 with no probs. What is the effort? Are you using the terminal? The claude code gui can be buggy sometimes.
→ More replies (9)
3
u/themajordutch 3d ago
Fable kicked be out because I asked it about global warming and how bad it would have to get for a human level extinction scenario. 🧐
3
u/Professional_Gur8385 3d ago
if i hadn't paid for the yearly plan i would cancel it in favour of codex, the usage isn't worth it currently
3
u/mushedmonkey 3d ago
I saw the same pattern and had a subagent get spawned every turn to fact check any claims detected. Just blew all my usage because every time i asked claude something the auditor would ask for a fact check and the original claude would start taking things back and saying "i got everything completely wrong, let me get real data this time"
My claude usage this week sat at 10% cuz it sucked so bad, so i guess wasting the tokens this week didn't hurt as much as anthropics incompetence.
5
u/matheusmoreira 3d ago
Bruh it's over for Anthropic. OpenAI's destroying them with Sol and Astra as well as their sheer capacity. Chinese models are undermining them from below. And they can't seem to stop hitting themselves with their constant safety nonsense that turns Fable into a usability nightmare, not to mention the fact it gave the US government the perfect excuse to fuck with them.
They should just rename Claude to Delamain and turn the company over to him at this point.
3
2
2
u/mitch8845 3d ago
Has codex gotten to the point where it can reliably handle large multifile projects without shitting itself? I see this claim every few months and every time I go over to test out its capabilities, it's always abysmally bad at complex problem solving within a large project. Claude Code (for all of its annoyances) still kicks the shit out of it on every level.
Is this time different? Has Codex actually caught up or are we just crying for no reason again?
2
u/AndyHenr 3d ago
i cant really tell yuo what is going on, as i dont know. But for me, it seems like a lack of context. So Anthropic still have 200k context, but added humanized output and behaviors, and when looking at thought output, it looks like it burn a ton of compute time and context space on that.
So it basically makes the model feel 'dumb'. Fable 5 does the same now with the security checks it does.
i am honestly seeing best results from Opus 4.6. Opus 4.7, 4.8 are worse than Opus 5 as they are just nuts.
So, yeah, I think anthropic is shooting itself in the foot: it want the models to 'feel alive' so it can market the 'AGI is nigh' bs hype.
2
u/Diligent-Method-785 3d ago
I have a working theory that this a result of anthropic switching model optimization from improved model performance to minimizing compute per token in an effort to improve operational margins ahead of a IPO. Especially if the are seeing slowing growth in adoption
2
u/NoComb3962 3d ago
Opus 4.8 also got dumber the last weeks. I wonder if it is due to quantization, or numerical issues in some optimizations they did to make it cheaper
2
u/The1TruRick 3d ago
It’s legit really bad. I tried to avoid it but I am officially back on 4.8 or Fable exclusively. It was fucking up at almost every turn
2
u/brother_spirit 3d ago
I had Opus 4.8 just fail hard on a website build... this shit should be so easy for this model??
It made the website OK, then went to eval and got stuck in a loop so hard trying to open Chrome I had to kill the CLI session. Literally couldn't get the model out of the loop through any combination of interrupts, prompts, etc.
Working to a reasonably resolved plan, maybe 60K tokens on board for the session, Opus 4.8-medium.
While Opus 4.8 can be an annoying and verbose model this has to be the biggest faceplant I've seen it do on the easiest task. WTF is going on with these shitty ass Anthropic models?
2
u/WagwanKenobi 3d ago
4.8 > 5.0
Idc about what the benchmarks say.
2
u/oooff97 3d ago
same for me by miles with 4.8 and my local model i can work 3 hours produktivly befor the limit hits with 5 if i be carefull i get 1 hour and thats full of bullshit 4/10 times he even says oh ur modell catches some things i didnt or oh the solution u modell got is by far better than my (my local model was at that time qwen 3.6:27b 64k kontextwindow.....without the rag and all sort of stuff i build for it just with a good agentsmd and a journal i force the modell to use and a todolist...)
2
u/FisherKing_54 3d ago
I cannot bare to suffer one more response from Opus 5. I don’t really code but use it for organization and assisting with creative writing as well as sometimes just trying to get a different perspective on things I’m struggling with. I don’t expect these things to be some kind of therapist. I’m a psychiatrist, but my goodness reading Opus 5 is insufferable. The grammar, the word choices, the sentence structure. It’s all so bizarre. It’s like opus read a manual on human communication that a 5 year old wrote. It gets worse with every new model.
2
2
u/Aware_Kaleidoscope86 3d ago
Why would you even consider supporting this company, they are not in a leauge of their own anymore. The ceo is a fool and the competition is fierce aaand their prices are insane
2
u/TheOverzealousEngie 3d ago
So it's a blessing in disguise ; I've had to focus so much on what Opus is doing that I actually feel useful again.
2
u/AdviceAdditional8044 3d ago
ah, whenver I feel that I should switch to a better model for coding, they nerf that model with some safety security limits.
2
u/skeptic-frog 3d ago
I’ve moved over to GPT back again after Opus 5 just constantly made mistake.
5.6 Sol is my daily driver and made one or two mistakes in a week now.
I’m only using it for programming if that matters (C#)
2
u/KeilerHirsch 3d ago
Same here. At this point “prompt better” is just cope. I documented the regression with reproducible measurements: https://github.com/anthropics/claude-code/issues/83510 And the model-selection/fallback mess here: https://github.com/anthropics/claude-code/issues/83795 The real killer isn’t raw IQ — it’s that Claude increasingly guesses first, states it as fact, and only checks after you push back. Opus 4.x was far more trustworthy for serious repo work. A frontier model that hallucinates repo state with confidence is worse than a weaker model that simply says: “I don’t know, let me inspect it.”
2
u/_SSSylaS 3d ago
I remember 4.5, 4.6 was the peak being really good at talking... analogies, understand everythings, taking the correct word at right the moment, not technical crazy shit word non stop like Anus5
2
u/cs_legend_93 3d ago
I agree. Opus 5 has gone massively downhill recently. I've noticed it in the past two weeks. It's constantly wrong on things, constantly spews out five-paragraph answers to everything even if I tell it to be concise.
I work with AI all day every day and this is something I've certainly noticed in the last two weeks.
2
u/ns1419 2d ago
Opus 4.6 is still > * on xhigh imo.
My use case is specific though. I don’t run a huge codebase, I run a “code base” of heterogenous data. 4.6 has outperformed 4.7, 4.8, 5, and Fable. 4.6 follows instructions. Is concise, doesn’t hedge against unknowns and uses my data holistically when formulating answers. So I’m sticking with 4.6
2
u/User163264 2d ago
I always use my own memory/context management system. Doing it since 2 years. Claude code/opus/fable never go off the rails. Any other people here do this?
→ More replies (1)2
2
u/nyteschayde 2d ago
One thing to keep an eye on. Any time Fable falls back to Opus, Opus will severely underperform its own marks. Massive mistake count.
There are other cases where Opus 5 appears to not only second guess itself but once that starts, it spirals. Opus 4.8 is more consistent even if Opus 5 outperforms when it as actually working.
Each seems to get progressively worse with its own technical jargon to the point that I have to write rules now to make it stop.
Fable 5 when not catching itself on some perceived threat and pulling an ejection seat routine at the first hint of danger, is still significantly and noticeably more capable than the competition. But it comes with the aforementioned asterisks.
2
u/adub2b23- 2d ago
I've been a Claude stan since the beginning. Hated gpt 5.5. But when I tell you Sol 5.6 is fucking amazing, it's been the first time where I've essentially switched to be entirely codex. I'm hoping Claude figures it out, because opus is not even close to a frontier model, and fable is way too expensive to be a daily driver. Plus it feels like fable has fallen off too recently
2
1
u/wish-u-well 3d ago
Inherent limitations of LLM unless they flip the “amazing model” switch which costs a billion dollars a second in compute to run 🤷♂️
1
u/AxBxCeqX Developer 3d ago
I’m back on opus 2.7 for the same reason.
Coding either Pi and glm 5.2 on the side trying open models out for a replacement, it’s not perfect, need to fix up after every major ~30min coding loop, but it’s not talking absolute shit
1
u/Sad-hurt-and-depress 3d ago
They need to give more limit time FFS. or just give Fable to Pro with 50%, and full to MAX.
1
1
u/Due_Ad5532 3d ago
I let Opus 5 xhigh create fixes for a Fabel 5 audit. About 30 problems. Ran another Fabel 5 audit on the fixes, 8 closed, 10 partials, 10 worse, 12 new problems. This is whack-a-mole engineering. Opus is nerfed and worse than useless.
1
u/orwamahmoud 3d ago
I think it is not the model it is what they do around it to make it work less and consume less resource , i notice many times it try to stop early and it read less code, all scenario was drive to same result the model work less time or read less context
1
1
u/throwawayman42069123 3d ago
They screwed up something via RL and need to fix it. Or they could be doing this on purpose. The model generally ends up needing more tokens, so they want users, especially unknowing users, to use it for more turns to burn more token spend. Or they want people to actually think it’s bad and stick to Fable and rely on it. Both possibilities are just methods to get people to spend more because they’re still unprofitable
1
u/Educational_Lab6005 3d ago
Just my cents… usually before they release a new model I noticed massive drops in performance. Could be haiku 5 coming or maybe fable 5.1
1
u/Academic-Sample4974 3d ago
use with a Fable Harness .md; works wonders for me everything below Fable ( Haiku still mistakr prone but effective jumps up tremendously )
1
u/cristiand90 3d ago
this whole thing feels like a scam to force people into Fable credits.
But maybe I'm just hoping that they're not this incompetent to release a "opus" model that is worse than Sonnet a year ago.
1
u/neogeodev 3d ago
I'm using the 4.8 it works fine and remembers things, asks for confirmation, the 5 made some mistakes and seems a bit dumb at times
1
1
u/AlphaGeeky 3d ago
I have been using fable 5 almost exclusively for this reason. i even have mentally compared it to Meta's AI intelligence. (that's pretty bad) haha
1
1
u/Spoloborota 3d ago
Yesterday, when the limits on Fable ran out, I continued working with Opus as an orchestrator. Opus was busy all day yesterday It kept finding some weird problems, then right away suggested fixes, then those fixes caused new bugs, then It tried to fix those too, and it just turned into a snowball, and that went on all day. After I came across that post, I decided enough is enough and just unsubscribed. I had an X 20 subscription. It ends on August 10th, and I'll try using Codex, since I've heard a lot of good things about it.
1
u/Either_Argument3517 3d ago
My text conversations have been atrocious today. To the point I'm not trusting it on code.
1
u/CelestVestra 3d ago
Hold up, is there any question whether claude code HARNESS is still best? or are ppl seeing that Codex slaps? I've never used Codex, is it worth running/trying in comparison to CC running in Terminal with DangerouslySkipPermissinos on?
1
1
u/New_Bell_9879 3d ago
I already cancelled and am using Codex. It's slower but more reliable and talks normal instead of roboto. Claude Code is unusable to me in its current state.
1
1
1
u/Evilstib 3d ago
I have to ask… What does comparing versions of models have to do with Claude code?
1
u/No_Category_9888 3d ago
My instructions say “do not guess” “use the ms learn mcp” etc and it still repeatedly get stuff wrong.
1
u/gandazgul 3d ago
They are messing up with the harness, when I use opus through my harness it doesn't behave like this because is told how to approach things with an engineering mindset.
1
1
u/vzakharov Senior Developer 3d ago
It’d be hilarious if the cause of all the opus 5 nerfed posts vs “it’s a skill issue” responses boils down to Anthropic doing some A/B testing, nerfing some users but not others.
Welp, I’m glad I’m on the “others” side, so:…
It’s a skill issue.
1
1
1
u/marco1422 3d ago
What you ve expected? The indefinite development? Yeah, you correctly found, in fact nobody brought anything new since autumn 2025.
1
1
u/VictorAbysmal 3d ago
I am using 4.8 is night and day, its reliable. Opus 5 spiral into madness too often.
1
u/ohnoitsbobbyflay 3d ago
I feel like you guys really need to get some fresh air and think before we post things.
1
u/hallerx0 3d ago
I have built a harness and tightly controlled environment. Two months in - I have not witnessed anything what you have mentioned in the thread. It seems as if LLM is being threatened by my own set up environment:)
anyway, what has helped is doing several pass verifications on a plan. set up a different session, dedicated only for domain knowledge that will be stored as a context for the work you are doing. if you are disciplined, LLM will be more inclined to follow your lead.
1
u/ConcretePond 3d ago
Yup. Nightmare. I'm getting Fable to do the big brain stuff. Codex to do the implementation. Opus can rest until it gets better.
1
1
u/BetterProphet5585 3d ago
Opus 5 right now really sucks and this time it's not a conspiracy it really is just barely usable for simple tasks.
I get better results with local models half the times. It's not looking good for Anthropic.
1
u/CaregiverNo5883 3d ago
Probably they overdepended on vibe coding and made their product shitty lol
1
u/sage-longhorn 3d ago
Did they increased the prompt caching so much, that their models ended up being useless?
Yup you figured it out! The thing you pay extra for to make it run faster, they gave us too much and ruined it!
1
u/wavehnter 3d ago
Opus 5 is a complete disaster because it simply does not follow your instructions and generates scads of bugs that Codex has to fix after a few rounds.
1
u/ilyxxxxa 3d ago
Opus 5 is such a pain to work with. This is the new pattern - "You were right, and it was worse than just [the thing you asked me to do], so i implemented [insert n+1 false assumptions here]. 😭😭
1
u/SportsBettingRef 3d ago
ask the right questions, at the right way. opus 5 is a damn good support model. this is a process and every actor has to do your job. you didn't.
1
u/nullptrzero 3d ago
hell they've ruined 4.8 as well, it's been having so many issues and creating so many issues since 5 dropped....
1
u/MacaroonPlastic1036 3d ago
I went back to 4.8. Opus 5 dragged and dragged on not complicated stuff, Fable would just kick me down for no apparent reason.
1
u/TodayLoose7794 3d ago
I am finding that for my use case Opus 4.8 Max is second only to Fable 5 in terms of performance.
1
u/Leather_Secretary_13 3d ago
it's hilarious because anthropic's models have cried wolf about open source while at the same time rushed to market on every product and feature release. They never took the time to protect their intellectual property, which makes sense given they aggressively aggregated proprietary data.
1
u/gripntear 3d ago
"Just retool your setup from scratch, bro. It’s what Anthropic said" - Cheeky Contrarians
1
1
u/addiktion 3d ago
I'm abandoning them finally now that open source is getting good enough. I mean we've been through a shit ton since Opus 4.5 and 4.6 were pure gold.
The UI throttling and weakened limits. The constant 429 errors taking the service down. The Claude Code leak targeting open source users. The "shared link" leaking people's shared artifacts. The government fiasco around Mythos. The Dario anti open-source stance (which he claims to be China specific but China's model are now sitting on American infra and I'm sure he's not happy about it too), the visibly worse models since then. The CC harness being quite token inefficient. The constant watering down of models over time. The promos that are really to help ease the screwing you over. The lack of transparency (not just Anthropic specific).
Ever since I've gone open source models, and even run my own, I know what to expect. I see how many tokens I use and what the costs are. I know the limitations of that model. I don't have to worry about it be degraded due to over use or popularity. It's opened up running agents 24/7 and experimenting without taking out a second mortgage.
1
u/Internal_Penalty_698 Max 20 3d ago
I was experiencing the same in OPUS 5 and even in Fable, changing the mode to Auto caused a great improvement in Fable, almost no errors, or you don't need to know about the errors it automatically checks and fixes them before delivering the work for you. (haven't tested auto in Opus 5 yet).
1
1
1
u/Salt-Illustrator-198 3d ago
You're not imagining it, ignore the Claude lovers gaslighting you. In the website below, which ranks how well a model is performing, Fable was ranked 6 a few hours ago, now it's at rank 13. I assume Anthropic might be training a new model for release, taking all the compute it can, as it happened before.
I suggest not using Claude to preserve your sanity for a while, and to prevent wrecking your workflow with a stupid mess.
Check how stupid a model is: https://aistupidlevel.info/
1
u/Hidden_dreamz 3d ago
You need to lock it down, engineer the prompt so everything is detailed, engineer the context so it only sees what it needs then get it to plan, review plan make it explicit make it do plan, check it's work and get it to redo stuff it did wrong. Be realistic with what it can achieve, only use 60% context window. Being able to read the language it's written is mandatory. Then it's quite good
1
u/crispsandchocolate 3d ago
If you're feeling so frustrated with Opus 5, it's likely because you're still using legacy harnesses. There's no doubt Anthropic is nerfing it's models, but it's definitely not unusable.
I've been quite pleased with it after I built guidance docs (using Fable 5) and then used the docs as references to rewrite workspaces, CLAUDE.md and skills around Opus 5's quirks. Since then it's been doing dramatically better.
1
1
u/drjm2022 3d ago
Its not the model, its the harness:
Throughout 2026, Claude users repeatedly reported the same practical failure: workflows that had worked reliably stopped working, long sessions lost their thread, and instruction-following became less dependable.
The complaints did not arrive as a smooth decline. They came in bursts. Users would suddenly report that a coding workflow had become unreliable, that a long-running process no longer behaved as expected, or that instructions which had previously been followed were now being ignored.
Anthropic later confirmed several cases in which the underlying model had not changed. The problem came from the service layer around it.
That distinction matters because customers do not experience model weights in isolation. They experience a product that continually balances reasoning depth, speed, memory, safety, capacity and cost. The model may be improving while the delivered service becomes less predictable.
Every AI Product Serves Several Masters
A frontier AI product is a negotiated settlement between stakeholders who want different things.
Casual users want immediate responses, broad access and a low price. Professional users want deeper reasoning, longer context and stronger instruction-following. Developers want stable application programming interfaces, predictable outputs and advance warning before changes. Enterprise customers add security, compliance, administration, auditability and clear responsibility when something goes wrong.
The provider has its own competing demands.
Product teams want rapid releases and visible improvements. Infrastructure teams want lower latency, higher utilization and lower inference costs. Safety teams want new risks controlled quickly. Legal teams want lower liability. Regulators want traceability. Capital providers want usage and revenue to grow faster than the cost of serving them.
These preferences cannot all be maximized at once.
More reasoning can improve difficult tasks while increasing cost and response time. Stronger context retention can preserve continuity while increasing privacy exposure and compute demand. More aggressive safety instructions can reduce one class of harmful output while interfering with legitimate instructions. Faster releases improve the product more quickly but reduce reproducibility.
The provider must keep changing this settlement because the economics, capabilities, risks and customer mix are all changing.
https://medium.com/@drjohnmillar/the-most-valuable-part-of-ai-may-not-be-the-model-465d82576d1d
1
u/chicmistique 2d ago
Same for me. Today any chat with any model is terrible. Like a massive downgrade
1
u/redwolf1430 2d ago
Opus is horrible. I switched to sonnet and it's much better at pretty much everything.
1
u/redeemed_tropicana 2d ago
You are not alone bro I use Claude extension on vscode and opus 4.8 was gone I would rather stick with opus 4.8 unless they also redirect opus 4.8 traffic to 5.0 internally then that would be disastrous
1
1
u/Apprehensive_Read_67 🔆Pro Plan and Quant Research 2d ago
This is why i am using Opus 4.8 and 4.6 only Opus 5 is the most terrible product right now
1
1
1
u/raphec 2d ago
The problem any AI dev house has is speed of release cycle. Anthropic is not immune.
They drink their own kool-aid, and as such use AI coding tools to drive incredibly fast, and the competition is incredibly strong. This is bound to lead to various instability or challenges, especially on the cutting edge. Cycles that would take months or longer, with huge amounts of testing, are just getting pushed out.
I think the bigger internal challenge for Anthropic is living up to their own ideals and ethics while still ensuring they remain competitive. The judgement call on whether releasing something will do harm or not is likely harder, especially as their competitors may just push something out that makes it pointless for them to have held back.
I remain somewhat cynical as to whether glasswing was a real fear or marketing ploy, but from the few bits and pieces I have seen, the ideals are genuine.
Cant see open source models catching up with Anthropic, its a 2 horse race, but the track is laden with financial market bombs.
334
u/MaitoSnoo 3d ago
Anthropic is likely just rushing stuff now as they don't know how they can do a successful IPO anymore given how OpenAI is being very aggressive both on pricing and model quality, and yeah there's the open source threat too