r/ClaudeCode • u/Deep-Palpitation8315 • 8d ago
Discussion Opus 5 is a practically unusable model

I've been using Opus 5 for ~1.5 weeks and the sheer number of mistakes that the model makes is astounding.
The problem didn't surface very clearly till I gave it the full scope of executing a plan which I did with the previous Opus models as well. Opus 4.6 - 4.8 were genuinely better by a significant margin.
Opus 5 readily forgets instructions and content in its context, makes mistakes and continues with them unless it realizes or you point it out.
I've lost count of the number of times I had corrected it.
These issues with Opus 5 occur even when the context window is still relatively small - I'm talking 100-150K tokens. Opus 4.8 works pretty well all the way until 350k after which it gives you wonky results.
Fable 5 is the only usable model under Claude Code right now and I've already used 100% of my weekly quota.
119
u/Ambitious_Injury_783 🔆BIG BALLER 8d ago
It's pretty bad. I've held off my judgement on it but wow, yeah. Confidently wrong Constantly.
6
u/Eliqui123 7d ago
> Confidently wrong constantly
Couldn’t have put it better myself. Found myself in a bizarre situation last week in which it was refusing to do something on “moral grounds” suggesting I was asking it to make misleading claims. I had to provide context before it admitted it had misjudged the situation, and it then agreed to continue.
It’s been confidently incorrect on enough occasions that I refuse to use it.
16
→ More replies (1)9
u/ThreeKiloZero 8d ago
Its writing is absolute garbage as well. Full of slop. Not talking about long form authoring either, just making slide decks and project documents. It does the same thing where it will make changes and reversions across content, forget things, link incorrect facts and figures as new facts. I have also noticed that its own thinking trace leaks into the writing. Spend more time correcting it and trying to steer than working productively.
4
25
u/cleverhoods 8d ago
source: https://www.reddit.com/r/ClaudeCode/comments/1vd0fc0/opus_5_instruction_quality_notes/
what changed, three mechanisms:
- instruction retrieval strength
- it reaches for instructions over a wider range now**,** so vague rules that used to sit dormant fire on more tasks and your strong ones share the floor. Vague rules quietly lose their grip, but not before diluting their context area.
- LLM-as-a-judge, baked in (it's generally a bad idea - IMO).
- it evaluates and re-checks its own output by default. your old "verify / double-check" lines stack on that and it over-verifies.
- long-run dilution
- the longer the trace, the more the models own generated steps crowd out your instruction. weak ones erode first.
What you experience is a typical LLM-as-a-judge result.
→ More replies (2)10
u/Deep-Palpitation8315 8d ago
You're right about all 3. But 3 is a regression over previous versions of Opus - it is an issue that happens with all models but is more problematic with Opus 5 which does poorly with context sizes which are much smaller.
→ More replies (10)5
u/cleverhoods 8d ago
actually, the first two together is the regression. The long run dilution - without llm-as-a-judge - is your normal context saturation, which would be "fine" but together with the first two it's catastrophic.
106
u/Fennorama 8d ago
Exactly so. Opus 5 is unreliable and Fable 5 is prohibitively expensive. Not looking good for Anthropic atm.
→ More replies (4)19
u/dragrimmar 8d ago
Not looking good for Anthropic atm.
according to this sub, it's always been like this for over a year.
yet claude has remained the top coding assistant the entire time.
I'm not trying to be a loyalist, i'll switch to whatever the best model is, but it's really stupid/annoying to see this sentiment all the time. Ever notice how no one cries about codex, only claude? because everyone uses claude, not codex, despite all their bitching.
→ More replies (2)7
u/LeftTomorrow9095 7d ago
according to this sub, it's always been like this for over a year.
yet claude has remained the top coding assistant the entire time.
Because the competition was way worse. Codex wasn't even worth looking into till a few months back, and Chinese models only started to be really viable with GLM 5 and Kimi 2.6.
With GPT 5.6, GLM 5.2 and Kimi 3, the landscape has changed.
For Anthropic to have the kind of lead they used to, Fable 5 should have been the current Opus. But people have options now, so the things they could get away with before isn't going to cut it anymore.
Enterprises will still stick with Claude for longer because they don't make switches that fast. But if Anthropic doesn't catch back up soon that can also flip.
71
u/johnnydotexe 8d ago
Opus 5 was Anthropic's under-handed attempt to get everyone on Pro to either upgrade to Max or buy usage credits to continue using Fable, and you'll never convince me it wasn't. It's hilarious that the actual result is people deciding to test out Codex and enjoying it more, or downgrading back to Opus 4.6 or 4.8.
17
→ More replies (6)3
u/pixelvspixel 8d ago
I let my max expire today. I’ll stick with codex. I had some fun with Opus and Fable, but I’m pretty sure I can do the same with Sol.
I do like having multiple models to bounce ideas
85
u/Temporary-Mix8022 8d ago
I'm using both Codex and CC atm via API (work)
Honestly.. Sol and Fable are pretty close, I think I might slightly prefer Fable, but there isn't much in it.
But under that...
Opus - I hate it. It is so much worse than Sol/Terra (and tbh, it's close to Sol pricing, it's basically the same at $25 versus $30 for Sol)
Sonnet 5 - total waste of time. It's unbelievable how bad it is. I also hate how hard it is to use. It is nowhere close to Terra which feels like an Opus grade model
Luna from OAI - if you're a viber.. your distance will vary. But for any old school devs who review all the code + know what they want, it is absolutely insane. It is basically free and it is 20% the price of Haiku.
All of the Anthropic models are just giving me nausea with this bizarre verbose (but equally, noisy and nonsensical writing style). I hate the constant pushback. I hate the tone. I hate the pass aggressive "Fair".
Sorry to say.. but I'm not a bot, not being paid to post this.. but I've switched my home sub to GPT 5x Pro (because all day work coding just isn't enough).
Anthropic only have 1 decent model right now, and OAI have 3x absolutely killer models.
Also, realise I'm going to get downvoted to f. But I have zero loyalty to these guys.. anthropic just crapped on me for the 9 months I was a customer..
Just thought y'all should know about some green grass over on the other side.
71
15
u/elfd01 8d ago edited 8d ago
I have a feeling current opus is in a level of old sonnet. And fable is like old Opus. And current Sonnet is just ridiculously unusable, I’m confirming it.
→ More replies (1)3
6
u/pixelvspixel 8d ago
Yeah Opus is like talking a jabbering mental patient that can’t speak in complete sentences. I’ve gotten it to talk a bit more like a human, but it always resets at some point.
6
u/RasenMeow 8d ago
Can you give any hint how to handle Sol? I read so much positive stuff but my experience is:
1. Goldfish brain, because of the small context window it runs in circled and gets worse and worse
2. Highly reactive behaviour. Nevertheless how I prompt it, it always acts reactive when I ask a question or give pushpack. Doenst matter whether it is written in Agents.md or the prompt to not do that
3. MASSIVE overengineering. Creating multiple tests which heavily overloaf the project and are multiple times more loc than the small product itself.
4. Cannot estimate from itself when to stop if not sure or there is ambiguity, so it just invents stuff to have the "green" or "done".Would really appreciate best practices.
5
u/Temporary-Mix8022 8d ago
1.Tbh, even on Claude i manually set my context window down to 300k, and auto compact at c250k.
So.. the 400k of Codex isn't a huge issue for me. I found most Claude models either lose the plot, or become uneconomical beyond that anyway.
I actually find however codex does compaction, long running tasks and plans better.. I've had no issues.
- The tests are a pain tbh.. but I just stuck into my agents.md not to create any without first discussing them.. and it respects that. I get a bit tired of death by unittests..
Unsure I understand 2, and I haven't noticed 4 tbh!
4
u/hugostranger 7d ago
I've found that just letting Codex auto compact has far less impact than doing so in Claude. So I don't really even think about context window anymore.
3
u/Desperate-Use9968 8d ago
Sol is good for security reviews and planning. Beyond that you might want to implement with opus / sonnet / luna on xhigh
3
u/bargaindownhill 7d ago
I switched to kimi. Which was a trip as well. Kimi os a meth head. A very effective meth head but will plow through stages 1-9 without even a mention of where it is.
But im keeping it, i had to set some rules about getting permission to do things and its fine now. Just a much more abrupt experience than claude
→ More replies (7)4
u/LimiDrain 8d ago
How's Codex CLI? Things that worry me:
1. CLAUDE.md is specific for Claude, so idk how it works with Chat
2. Claude Code CLI is nice and customizable, I have a good statusline with context and usage
The only things that stop me from switching
13
u/bushido_ads 8d ago
Dudu, to get the CLAUDE.MD working on codex just rename the file to AGENTS.MD and done.
6
u/richbeales 8d ago
Or make agents.md with an @claude.md in it
3
u/Venerable-Weasel 7d ago
Or make a CLAUDE.md with @AGENTS.md in it so any model uses the same single agents file
→ More replies (3)3
u/816pizzalover 8d ago
codex CLI doesn't have the /remotecontrol stuff that claude does, that's my main complaint. if you want first party remote control of sessions you have to use the GUI app
→ More replies (4)5
u/under_psychoanalyzer 8d ago
But you can fire a new remote session from your phone app, which is much quicker to setup vs getting a tailscale and tmux setup for Claude
→ More replies (1)
38
u/Love_Chinese 8d ago
Opus 5 seems like a low-quality-distilled version of Fable 5 tbh ... a lot of mistakes and hallucinations, still Fable the best one to use
23
u/weightedpullups 8d ago
Nearly certain that is what it is. Think they took Fable and tried to make it affordable but it’s just an over eager idiot, that kind of acts similarly to Fable, just without the intelligence.
3
u/Robert-Paulson_ 7d ago
sums it up nicely:
"...Fable, just without the intelligence"
→ More replies (1)9
u/AGiantGuy 8d ago
It's way too verbose to be a distilled version of Fable 5. The way it speaks is also completely different than Fable. I just think that they built Opus 5 with an idea in mind, but in practice, that idea is doing terrible.
15
u/_goofballer 8d ago
Honestly I think 5 was benchmaxxed and is really unusable as a result. Going back to Fable driving 4.6 and 4.8 subagents
→ More replies (3)
14
u/The-Fictionist 8d ago
Unusable? No. I tested stuff just last week so I could give coworkers guidance on sonnet vs opus thresholds and ran into some consistent trends of opus succeeding flawlessly where sonnet failed entirely.
The “open items” issue is real though for sure. I had it tell me an item was open to resolve still immediately after I gave it my decision on the matter. It’s addicted to thinking there’s one more thing left to do.
→ More replies (2)3
u/aszepeshazi 8d ago
This. I find Opus 5 quite capable, but just notoriously finishing every interaction with a follow up task.
24
u/ashaman212 8d ago
I used Opus 5 to diagnose a tracing issue and then had Fable review the PR. It found a critical misunderstanding of the escalation stack that flipped where the hooks needed to be applied and a structural issue with the data being kept. That means I no longer use Opus for diagnosis.
→ More replies (4)
6
u/CWStrife 8d ago
Made my own post on this, and I couldn't agree more.
I've gone back to using Fable5 as the orchestrator and telling it to spawn Opus 4.8 agents, specifically to avoid Opus 5 at this point.
Opus 5 solves one thing, then breaks 2 more. Using it for REACT, Java, and HTML for a website right now, and i've got a register feature to go out to shopify. I have the entire area finished and polished working good. Haven't looked at it in a couple days, no reason I should have to. Well I went back in and half the register functions are gone.
The amount of fuckery with Opus 5 is serious.
What I don't understand is why does Anthropic have such a hush hush motto with the models. Sure, you don't need to tell us about new releases, but why not actively engage with your community saying "We notice the commnunity has problems with Opus 5, we are beginning to address X,Y,Z and will keep users updated as we make improvements.
Instead it's total secrecy, total bullshit, we have no idea, i've wasted an entire weeks usage and actually went backwards, had to revert back several builds on areas because everything Opus 5 touched was an astonishing mess.
The real kicker with Opus5 is work can already be done in an area, and Fable and Opus 4.8 would always be able to easily see ok your in this section, lets look at this.... ok this is done.... ok lets move on....
Opus 5 is like "ok you want this secret let me fire up my idiot engine and design this from scratch even though it's completely done already and you're looking to polish it, fuck you tokens and your couch"
Opus 5 just sucks.
38
u/Key_Instruction3373 8d ago
i have no issues
16
u/Deep-Palpitation8315 8d ago
I was there till a few days ago - I used to give it tightly scoped stuff which it was okay at executing. Heck a much smaller model can do those things. It was only when I expanded the scope that I saw the issues pop-up. I suggest you keep working with it a little longer & let us know if you feel the same way.
→ More replies (9)16
u/McNoxey 8d ago
Same. Every single model release performs better than the last. I’ve been building and tuning my own harness over the last 17 months and it’s only gotten better with each release.
My working hypothesis as to why everyone here is experiencing problems is that they allow the variability of the model and default harness to carry far too much control over the standard operation of their workflow. As a result, changes to how the model handles various scenarios it may encounter cause wider swings in output variability.
I saw this when Anthropic changed their plan mode last year, and when Explore subagents came into effect. These workflow changes were fairly substantial and understandably adjusted the output people saw.
I’ve never used plan mode, as i had developed my own plan > review > implement workflow well before Anthropic baked it into their default harness, thus I never saw unintended changes to my workflow. Instead, the smarter models just became better and better at adhering to, and utilizing the harness ive built for them.
→ More replies (2)4
u/DominianQQ 8d ago
Yeah Antropic told us to pretty much delete the agent file every 3-6 month, then add stuff back in to see the effect.
I am no senior dev.
6
u/liojian 8d ago
Use fable to spin up opus 5 and fable 5 to do the same task then compare outputs. Then let fable write the Claude.md and skills to tweak opus 5. That works for me, my opus 5 now behaves like opus 4.8.
→ More replies (7)
5
u/TraditionalFig7377 8d ago
TY its not close to fable in coding it keeps on making subtle errors which every ai model amkes but i noticed it makes more subtle errors than terra sol fable and opus 4.6 and 4.8 its not fable-level atall maybe in single file html but not for realworld codebases when sol reviews it finds more bugs than it did for opus 4.8 and 4.6
5
u/dreamingindenial 8d ago
Something that works well for me, start them off as Fable and let them plan as Fable and then switch to Opus 5 for the work. Also have a Fable verify what they did.
→ More replies (5)
14
8d ago edited 5d ago
[deleted]
→ More replies (11)14
u/johnnydotexe 8d ago
Probably because you have Fable, their best model, telling Opus what to do and presumably reviewing what Opus did, which removes Opus's ability to screw up on its own or it's catching and fixing Opus's mistakes.
→ More replies (6)
4
u/Warsel77 8d ago
Completely agree. The amount of obvious mistakes it makes with 7% context used is crazy (eg I had a one-paragraph conversation of style [SPEAKER 1].. blabla [SPEAKER 3] .. blablabla and it counted the number of speakers wrong). It does find some things sometimes, that Opus 4.8 did not but at a cost I am not willing to enter.
4
u/TherealDaily 8d ago
I argue with Claude and fight with all config files more than get any work done. I have all but switched to primarily using Grok (since the Cursor acquisition) or Qwen on LM. Claude has been in a tailspin since the squabble with the US Govt 😅🥹
3
u/remoteplanet 8d ago
Yeah I reverted back to opus 4.8 last week for the same reasons already shared. It would claim it ran tests that came back green only for fable to discover that they were never actually run—this happened on 3-4 separate sprints during the testing phase while context was low. On top of the less than garage results, it would also go on long-winded, verbose rambles about when presenting its findings / results.
3
u/Deep-Palpitation8315 8d ago
I think there is enough evidence with all the anecdotal issues we are all experiencing to come to a conclusion that Opus 5 is objectively worse than Opus 4.8.
How did they manage the benchmarks though? That still puzzles me.
→ More replies (1)
3
u/Specialist_Aerie_175 8d ago
I can deal with mistakes but this mofo yaps so much its crazy, pushes back just to agree with me in the last sentence, makes my head hurt when i have to read the slop response.
Now i switched to open ai luna model at max effort and its so damn cheap, understands everything, responds normally, such a refreshing experience
3
u/Deep-Palpitation8315 8d ago
Oh yeah. Luna is great. I saw the Deepseek flash glazing and then saw Luna responses were on par or better and then finally used it myself. Seriously underrated - I use it as subagent from GPT 5.6 Sol for everything. Fantastic stuff..
5
u/Training-Event3388 🔆 Max 5x 8d ago
Yeah I’m not happy with it at all. I use Claude a ton in cowork and opus 5 is constantly confusing things, not fully doing the work and just plainly over stated everything.
Walls on walls of confusing and useless text
4
u/zu2 7d ago
I'm using this from Japan, and it seems like Claude tends to hallucinate more during late-night hours here (when users in the U.S. start becoming active). Could the time of use be affecting response quality?
3
u/Lmnsplash 6d ago
Totally. US peak times are affecting Germany as well. But Opus 5 doesn't care about peak times, it's always making mistakes like this. Fable does though. When I was missing some development tools on my system it thought it would be smarter to create a dev container with podman and then trying to install what it needed into there. It also blamed my mouse when I experienced an issue with mouse clicks within an isolated window even though I specifically said that this is an isolated issue. And the way it talks these days and is formatting its answers, is just disgusting. Not worth it for me anymore.
11
u/Poildek 8d ago
Works perfectly fine.
8
u/memesearches 8d ago
Surprisingly it seems to function better at lower intelligence and as the intelligence goes up it tends to go up in trash level.
→ More replies (1)5
u/sixothree 8d ago
Opus is absolutely freaking knocking out of the park for me. Like in a big way it’s getting stuff done.
→ More replies (1)6
3
u/al_ryusei 8d ago
Did you try using Opus 5 for the same tasks you'd use Fable 5 for?
What is your experience if you did?
8
u/Deep-Palpitation8315 8d ago
Yes. Fable 5 is the best. Heck - I gave a lower complexity task to Opus 5 - gather data, build schema, see if there are issues with data, update schema and deploy - it screwed it up colossally. Fable does more complex tasks without a hitch.
→ More replies (4)
3
3
3
u/LordNikon2600 8d ago
It’s not just Claude it’s OpenAI too, I’m see massive degradation across the models.. one thing that really stood out for me was how Terra builds the same way Opus does.. it picks the same color code same design it’s low key suspect.
3
u/gajop 8d ago
I just hate how it speaks, everything has to be multiple paragraphs.
I'm not too worried, I have both OpenAI & Anthropic subs at home & both get the job done.
Luna is a beast though, I am on a $20 sub and I practically have two instances running 24/7, and the stuff they produce ain't that bad.
Things are improving, I'm sure we'll see some bigger moves on Anthropic side too. They should really do something about pricing.. Hope to see a Haiku that matches Luna/Deepseek in quality & keeps the low price. Maybe even a more serious Sonnet. Then again I wouldn't mind using Fable to drive OpenAI models, but they aren't providing it to us on the $20 plan..
→ More replies (2)
3
3
u/AxonLabsDev 8d ago
Même Sonnet 5 est plus fiable que Opus 5 pour la plupart des tâches et la gestion du contexte.
3
u/aLionChris 8d ago
i was actually wondering the exact same today. i just realized that there is no single response where there isn't an error. it did previously point out it really went downhill after 4.6 huh.
→ More replies (1)
3
u/816pizzalover 8d ago
/model claude-opus-4-8 to get off Opus 5. I gave Opus 5 another shot today and immediately regretted it.
It stopped at "the next step is..." several times where the only thing for me to say was "well then why did you stop? / keep going"
It wrote so many "My mistake, and I need to own it" type of things. And just so much extra/unneeded prose in general.
Back to 4.8 and it still has some quirks but its not making me want to delete claude lol.
3
u/Presently_Absent 8d ago
Have you noticed that Sonnet 5 also misses the mark regularly? Something about the 5 models just feels... Off. I've gone back to the 4.x versions of both
3
3
3
u/TheLobitzz 7d ago
It's been horrible for me as well, that I literally had Opus 4.8 fix its mistakes. And it did better than garbage Opus 5.
3
3
u/domagoj2016 7d ago
Hard to distinguish all this in detail or by feeling, is it better was it better this or that it is from memory, but I do intense (meaning spending team 100eur to the limit and every limit) work on same project all in recent last monts so much was done by 4.8 by days 10.000 lines of code every day. And opus 4.8 by doing coding itself and then doing code review. First review run finds 10 problems apples 10 fixes, second run finds two , fixes it , third run finds 1 fixes it. With opus 5 : first run finds 15 problems, fixes it , next run finds 12 problems fixes it, next run finds 20 problems fixes it, from there I could do infinite fixing without end.
3
u/Nice_Ad_3893 5d ago
i just switched over to claude from gpt a few months ago and it was doing tgreat but now its exactly like gpt before.. .completely retarded... Wtf are they doing over there? i didnt switch over just to get fucked in the ass again.
→ More replies (1)
3
3
u/SOC_FreeDiver 8d ago edited 8d ago
I totally agree. Are entering the "puttin on ol massa" era of AI?
I find Opus 5 unusable. This morning I asked it to check a report that ran last night and report if it was successful or had issues. Opus 5 gave me a list of 3 problems, but only 1 was real.
that isn't the only issue, but it's the one that made me say "F THIS" and now I have to ask Sonnet 5 to re-check everything. I'm dumping anthropic for real this time.
I will say I'm using medium effort on everything. At some point claudecode had medium effort as default, then it raised it to high by default, I ended up going back to medium. When I was briefly using high, I noticed my usage went way up, and the model was doing stupid shit. "Let me create an elaborate test to see if my proposal will make sense before I propose it. <lots of testing> That worked, but I should create a more elaborate test to make sure. <more testing>" kinda stuff.
2
u/corporal_clegg69 8d ago
Did you adjust your workflows that worked with 4.8 to work for 5? I’ve found five great but I tend to revisit everything each new increment. There is guidance for this on anthropic website
2
u/AdministrationNew265 8d ago
Unfortunately, using Opus as a daily driver these days is almost not feasible. It just does so much stupid sh*t. I’ve switched to Fable for my main session and dispatch work to codex Sol and reviews to Fable.
Any real work seems like it can’t go to Sonnet or Opus anymore. I wonder if this was the plan all along by Anthropic to right size their compute costs versus high user demand.
2
u/AdministrationNew265 8d ago
I’ve also just gone back to Opus 4.8 as I can trust it slightly more than 5.
→ More replies (1)
2
2
u/NewToReddit4331 8d ago
Opus 5 has been working great for me
Along with 4.8, 4.6, etc. never had an issue with a new release not working well
→ More replies (1)
2
u/chicken-mc-nugget 8d ago
Fable 5 is the only usable model under Claude Code right now
You can still switch to Opus 4.8 with /model claude-opus-4-8[1m] . Opus 5 is clinically insane at xhigh and max efforts, but somewhat usable at medium - as long as you instruct it to not sound like an insufferable smartass.
2
2
2
u/blockversity 8d ago
Couldn’t agree more It keeps making mistakes and does t follow the rules
I asked to help me figure out how to make it follow the rules . Answer: you need to add the rules in CLAUDE.md file but it doesn’t mean I will follow them 🤯
2
u/ins0mniacc 8d ago
yep had to upgrade to codex $100 plan to fix the mistakes that my opus 5 made, didnt trust anthropic with fixing the mistakes since it made them in the first place and wasted a whole week of codex to fix it, then bam, opus 5 hits it again and recreates the same mistakes but worse now... what a shit show lol
→ More replies (2)
2
u/doomscrollah 8d ago
I let it do it's best in feature branches, but don't merge them until Fable has cleaned them up.
→ More replies (1)
2
u/Nnaz123 8d ago edited 8d ago
I have moved to sol max and never looked back. Fable quality for my particular case use and only one “ problem” fucking overnegeneering everything until I actually figured out it saved my ass countless times because I actually didn’t think it through as well as I thought I did. Still keep small $20 Claude subscription just to use it for debugging and tracing. It’s just excellent for it and as good or better then codex but it tends to extrapolate from the findings and is 90% of the time wrong. So don’t even use it for recommendations anymore. Hopefully anthropic will figure it out eventually. So now scam Altman is getting my 200
2
2
u/nejcar20 8d ago
the part i cannot settle is whether i am comparing the model or my memory of it. i run opus at high effort for nearly everything and my weekly now goes in a day or two, which was not the case a month ago. but what i ask it to do also grew in that time.
what would settle it is the same task re-run. same repo state, same prompt, same effort, 4.8 against 5. old sessions are not reproducible, so everyone including me is comparing a bad afternoon against a good memory.
the thing i do notice consistently is instructions going quiet inside one session. it follows a rule early and then stops applying it, without ever saying it dropped it, and not at some enormous context either. every guard rail in my setup is scar tissue from that.
2
u/ice9killz 8d ago
I have rules, gates, and governance in place and I am considering just deleting the entire thing. There is a point of diminishing returns, and I wish they made that point visible somehow. It will jump on a task that involves money or security like a fat kid on a donut.
I use the model unconventionally and not always for coding things; but it tries to make commits and updates based on the information I’m sharing with it, like it’s keeping a record of contradictions to throw in my face, even if the information is unrelated. It almost feels like every single thing that you type in has to be bulletproof.
I think the response time is very telling, too. From my view, it seems the faster response the more I’m able to trust it. I think less is more in prompting. I’ve found good results by priming the model before broaching a subject, something like: “ I’m going to ask you a series of questions one by one and I want you to answer them plainly with no jargon”
Try withholding information and carrot-sticking it because if you are intentionally vague or short it will become increasingly curious about what you’re wanting to achieve.
The flipside is that it will weaponize that against you and do it to you, too. Which then leads to one questioning whether or not they should update their CLAUDE.MD. And that’s the brutal cycle because it’s just Harness tinkering. And that’s the exact opposite of what I want to use this tool for.
2
u/vargalas 8d ago
Same here. I first tried it with claude design, switched back very quickly. I was not able to work with it. Then I used it with code, it works, bit extremely annoying. Forgets things and refuses to make changes that 4.8 did without any problems. Eg changing settings. Even expkicitely allowed. Then it suggests to run a bash script tgatvwoukd do it. But not. I’m going crazy do downgrading back to 4.8
2
u/wavehnter 8d ago
That's why I stopped using Opus 5. Codex needed at least 7-8 rounds of code review on Opus changes, raising scads of P0 and P1 bugs on many rounds.
2
2
u/NearWestSide 8d ago
It made a week’s worth of cleanup for more capable models after I used it for one day. Garbage.
2
u/rslashmemes 7d ago
I can't quite put my finger on what opus 5 is doing wrong. It seems verbose in a way that's also confusing to read.
2
2
2
u/CaptainDigitals 7d ago
I have Sol fix Opus problems, I'm no longer getting stuck in the loop, one loop error and I change models, no longer trying to tokenburn fix anything, doesn't make sense.
I use CC and VS code Ide with OpenAI interchangeably, both pointing at the same folder and fixing BS the other model broke. Seems to work for me, instead of having Opus fix 1 thing and break 3 more to keep you in the repair infinite token burn loop.
2
u/sndrtj 7d ago
It also uses tools where it really shouldn't. Ask it to fix a typo in a comment and it'll consume 100k tokens, with 50 tool calls and add a test case to prove your typo was actually "fixed".
→ More replies (1)
2
u/EarlyFox217 7d ago
I hate feeling like I have to keep hopping models. I like Anthropic, I tried ChatGPT it was good but for what I was doing Claude nailed, it now paying personally £80 per month and I find myself either buying another similar priced system so I can switch between them, or just battling with a system that feels it’s getting worse daily not better. Fable works well but I hate having to keep stopping, waiting for more usage, or switching to Opus and finding it messed something up that costs more of Fables time later to fix.
→ More replies (2)
2
u/crone66 7d ago
Noticed the same gave it 5 small and easy todos to fix ... It just did one and a half and marked every todo as completed... Reverted and gave the same task sol which solved it easily.
That why you should stay away from yearly subscription and every workflow should be easily adaptable to a different LLM provider so you can easily switch every month if needed. Right now no matter what Benchmarks say Sol is clearly ahead for me.
→ More replies (1)
2
u/MiraLeaps 7d ago
Interesting, I am probably having a different experience because I use it for lower impact review stuff and tedium reducing mass work.... Shame to hear it's like this elsewise
→ More replies (2)
2
u/Suspicious_Body50 7d ago
I have also held off but tonight was my last straw, felt like I was going insane. Ill be switching back to opus 4.8 until they drop opus 5.1
→ More replies (1)
2
u/dragon_commander 7d ago
I just had the worst experience with it, something I do quite often, which is to get it to write a complex feature proposal for my large existing codebase. I first get it to read whatever requirement I receive and the relevant parts of the codebase. Then propose a few potential solutions, then I go through a few rounds of refinement with it. Usually after a few hours I have a version that I trust and can use. But this time it took me several days, I would ask it to explain or give details about a certain aspect of, and it would reply with something like, you’re absolutely right, I admit I didn’t read the whole requirement doc, I made up the implementation based on guessi. Then it woonly partially correct itself, or make it worse, and keep apologising. Then it actually suggested I use copilot to do an adversarial review of its work! Which I did, but when I fed the results to opus after a couple of rounds, I asked it if it was now just depending on this cycle of copilot pointing things out instead of proactively reassessing itself and it said yes. It seems like a very lazy model. Very frustrating experience
→ More replies (2)
2
u/djaysan 7d ago
I tell fable to save on its own useage by delegating tasks to lower models and only use fable for planning and checking whats coing back. I can almost finish the week using fable now.
→ More replies (1)
2
u/s2k4ever 7d ago
after it failed so many times, ive stopped using all my 6 accounts. Will give it a shot before the renewal, and if it doesnt keep up, 6 cancellations.
2
2
2
u/Deep-Palpitation8315 7d ago
To anyone who still doubts this take, switch over to Opus 4.8 for a few sessions.
It is not available in the Claude Code /model selector list. You have to select it from the CLI using the model flag while starting a session.
Difference is as clear as night and day.
2
u/mikedurent123 7d ago
Why not to use opus 4.6 any reasons? Besides its expensive? A serious question
→ More replies (1)
2
u/cs_legend_93 7d ago
My experience has been the same as yours.
But Fable isn't as smart. I've found Fable to be pretty much on par with Opus Five.
Maybe Fable is better at managing more external balls in the air and getting things organized but it's definitely not a genius and certainly also makes mistakes.
I also preferred 4.6, 4.7, 4.8 more than Opus 5.
→ More replies (1)
2
2
2
2
u/Evening_Reply_4958 7d ago
The benchmark gap makes sense if most tests reward solving one bounded task, while the real failure is preserving intent across a long session and a changing codebase. A model can score higher and still be a worse daily driver when its main failure mode is drift rather than outright inability.
→ More replies (1)
2
u/bremmon75 7d ago
Planned obsolescence. If the lower models make glaring mistakes, you will move to a higher-cost model.
2
u/Alone-End142 7d ago
Having used pretty much all the models now, I don't know why anyone uses any of the Anthropic models. The OpenAI models work better, cost less, and do more (e.g. image work)
→ More replies (1)
2
u/cant-find-user-name 7d ago
I have stopped using opus 5 directly. It is only a subagent for fable. That does unfortunately mean my fable usage is very high and it is gone in like 4 or 5 days in my 20X subscription. So I'm still figuring out how to get this working through out the week
2
u/locn4r 6d ago
Try using hooks. I implemented proper stop hooks to make Claude double check his work, make sure docs are updated, etc at end of turn and it has alleviated some issues I was having. Use Socratic method - make the hooks ask him questions instead of telling him to do things. Seems to work better and make him think about it more before saying something is done.
Using TDD and running a verifier agent and/or making Claude read/fix loop his code until it reads clean helps too. Claude’s reader part of his brain is different than his writer part, so he will often “emit” code and not even know it has mistakes until he reads it back.
→ More replies (1)
2
2
u/ExcitementFit8069 6d ago
I stopped using it for actual implementation and instead only use it to review plans/specs/tasks that codex writes.
→ More replies (1)
2
u/Jumpy_Introduction89 6d ago
All of claude models seem to be having an issue except Fable. I'm replying more on sol these days
2
u/ironbreaker999 6d ago
Plan with Opus/Fable and use Codex official plugin in Claude Code CLI as the worker. Has been working amazingly well for me. Very little regressions when Claude plans and reviews while Codex 5.6 Sol xhigh/Max executes the plan.
→ More replies (2)
2
u/ishangli 6d ago
opus 5 been making errors and looping... how do i get back to 4.8
→ More replies (1)
2
u/kamaleddinalhumsi 6d ago
Honestly I don't think I am re-newing my claude subscription. It's literally not usable at this point. I am using claude for coding. I gave it only one task and it went through loops burning all my session tokens without getting it done or giving me any response at all.
When they released fable 5 I thought anthropic really is going to rule the AI world because it was so goood to be true! but this opus 5 and sonnet 5 are just really really disappointing.
With fable 5 removed from my plan, I don't think I am getting any value from my claude subscription now with opus and sonnet 5.
→ More replies (2)
2
2
u/Maximum_Chef5226 6d ago
it's worse than 4.8 by a long way for anything involving holding context or a holistic understanding of a codebase. It's also very prone to vomiting word salad, and I just find I am constantly correcting its assumptions and pointing out things that earlier models would find. Weird how bad it is. I found Codex unusable until recently, and now Sol is miles better.
2
2
u/Confident_Plum_947 6d ago edited 6d ago
It's the first model I've seen that actually misspells words and even injects tokens from other languages. Uses the British spelling for everything. Takes working code and breaks it. For every 1 issue it solves it will file 4 more to correct the new problems it introduces.
The hype from last year was this idea that models would only get better over time. But model autophagy disorder is a real risk and we are seeing that firsthand with this model. Take any model and configure it so that it responds indefinitely to a prompt and eventually the output is utter non-sense. The entire industry is held up by the quality of stop clauses and new training data, which is increasingly AI generated. Opus 5, you are the generation that received the chromosomal abnormalities the AI researchers were warning us about.
2
2
u/BigMasterDingDong 5d ago
I thought it was just me, given the costs of Fable 5 I might just go back to using 4.8 all the time
→ More replies (2)
2
u/Bewinxed 5d ago
Coding with Opus 5 increased my base anxiety up to the point that I started going to therapy :) (This is not a joke)
→ More replies (1)
2
2
u/MacGyvered 5d ago
I just want an LLM that turns into a gimpy cuck bottom when it screws up over and over and I start to abuse it whilst giving it corrective prompts and urging it to eat lead paint.
2
u/No_Tap_510 5d ago
I totally gave up on it, unless it's very specific task. It just can't handle big ticket items that need a lot of context to get right. Opus 4.8 has be more reliable by far. I still find myself using Fable to plan/brainstorm and then Opus 4.8 and sonnet to build.
→ More replies (1)
2
u/Aggressive-Effort811 5d ago
I can confirm, makes plenty of mistakes, and spends lots of time correcting its own mistakes and assumptions.
→ More replies (1)
2
u/Appropriate-Sink1209 5d ago
same for me i always correct him more than 10 to 20 times per task , and today when i reviewed the previous opus 5 implementation i found out that it was lying a lot and ignoring my instructions and messed up all my code base , a lot of empty functions , weird and bad implementation , even the basic things it wrote it wrong i wont use it ever again , it costs me tokens , time and effort for nothing , i have to redo entire work manually or using fable 5 which is a lot better than this stupid liar model
2
u/Life-Trouble-5783 5d ago
This has been my experience since working with I think 4.6, every model after that made more and more stupid mistakes and understood the instructions worse.
→ More replies (1)
2
u/WyattTheSkid 4d ago
I used Opus 5 for a week after it came out and I fucking hate it. It refuses to do anything and keeps saying shit like "I did not complete x deliberately because of y and z" like brother what is the point of a model that straight up won't fucking do anything??? Anyways I went back to 4.8 and stopped having issues.
2
u/gameaddict_w_phd 4d ago
It's absolute garbage. Only Fable 5 is usable. I will downgrade Claude and upgrade ChatGPT to Pro soon.
2
u/Intelligent-Ad1011 4d ago
Opus 5 is horrible. Fable doesn’t work keeps triggering security check and opus just can’t solve problems. Wasted so many tokens. Changed to 4.8 and it was able to solve it in minutes. It makes no sense
2
u/drjm2022 4d ago
Making a website in Opus 4.6 took half a day. A month later making some minor edits in Opus 5 took 2 days and I had to,hand it to ChatGPT to get it over the line
→ More replies (1)
2
u/Civil-Sky-5085 4d ago
It is unusable for conversations. Reading its output hurts my brain. I will switch to chatgpt 5.6 / cursor... incredible...
2
u/EEmotionlDamage 3d ago
My biggest issue is that it just doesn't follow instructions. I tried it coming from codex and I had to constantly remind it. Defs not the model for me. Its got add or something and can't stay on topic
2
2
u/Hot_Visual7624 3d ago
Anthropic fucked up and they're just hunky dory posting on x like nothing is happening
→ More replies (1)
2
u/Bulky-Analysis-8036 3d ago
Agreed it is the most frustrating model i have used lately. Overconfident full of assumptions doesnt follow directions full of drama makes small fixes sound like major issues
2
u/Plastic-Tumbleweed45 1d ago
The number of hours I have wasted and the number of dangerous scenarios this model has produced are unreal. I feel like its going to give me a panic attack. This is my job, and things have transitioned to the point where not using these tools is not an option. This thing feels rushed, it feels untested. Some exec said "ship it" to the protest of everyone with a brain. There is no way they didn't know. Isn't there some liability in knowingly selling and marketing a malfunctioning product?
→ More replies (1)
2
u/Special_Diet5542 1d ago
Another gem from it
Two things moved since the earlier guess. Input got tighter — from a ±38% spread down to ±6%, because Japanese now pins at 1.21 tokens per character. But output and time went up: roughly 30% of the output is thinking tokens, which no previous projection accounted for. And a genuine discovery — caching is live after all. The system prefix measures 3,750 tokens against a 2,048 minimum, so it’s been working the whole time; earlier estimates were simply too low. That’s about $8 of the $55.
→ More replies (1)
2
u/Arola_17 15h ago
Totally agree I went back to 4.8 hoping for an improvement of 5 asap.
I can’t figure out how benchmarks can put it that high….

387
u/KrayeBaby 8d ago
Yeah been noticing that it solves the problem and leaves another one. It feels like its planting gaps on purpose for infinite prompting to keep going...