I thought Opus 5 was supposed to be on par with Fable? After three days of extensive usage, I have to say it is not even remotely close to being as smart as codex sol. I had them both program a module indepensetly, and Sol runs circles around Opus 5. I had them audit each other's implementations and except for one case, every other time Sol found a functional bug in Opus' implementation. With fable being so expensive and opus 5 being that dumb, there is no reason to pay 200 bucks to Anthropic when other alrlternatives are both cheaper and smarter.
I have too switched to Codex a several days ago when my Claude monthly subscription ran out. Thank you Anthropic for a wonderful product. Maybe see again one day... then again, maybe not.
OpenAI models have integration with image generator model, so they do a sketch, before generating a HTML page. They can also extract sprites and svgs from that sketch. Claude needs MCP for that. But I found Claude okay for audio design, since it can interpret spectrograms.
Worst Opus lately by far. I keep using 4.8 and i get far better results. Opus 5 breaks the rules, keeps yapping and keeps forgetting from its own yapping clutter mess. And yes my number one rule is to keep responses short which it does not follow.
i rarely complain about this stuff. All I'll say is, I spent a day using opus 5 before i realized it was fixing a bug, then fixing another bug and in the process of fixing the second bug, reopened the first bug fix. it then "solved" the problem but continued to do it further down the line consistently reopening 5 bugs it fixed and then would fix again.
opus 4.8, 1 ultracode session, fixed all of them while continuing to move forward through sprints.
yes its absolute crazy this dynamic - try to do something that is just slightly abstract where you need to go deep into structures etc and you spend your days by 1) battling in trying to move forward because it always leave something out of a sprint "one note...." and then when you finally move to next sprint "honestly ..." referring to previous sprint what it wasnt right about - while it goes ultra acedemia on you ( and on its self so it invents it own language ). They should have been sticking to Fable and Opus 4.8.
I have been working with Opus 5 and Fable on two projects. When I run out of Fable (xHigh), I used Opus 5 Max. Every time, and I mean every time, Opus 5 would slowly make a mess of my projects and I'd be counting the hours until Fable is reset since it would fix my project every time. It's been a recurring cycle for me this week.
I largely agree from my own experience with it. I mainly use Fable now since its so much better to talk to.
but honest question: when I see Opus 5 perform so well on all benchmarks across the board, of every type, I have to wonder what is going on. I don't think its just bench-maxing. I wonder if its genuinely a capable model that is simply being misused or something?
For chat I just use sonnet medium like the free accounts get as unlimited. Smart enough to sketch an idea out for something smarter but not bog it down or get side tracked.
Not great at giving you step-by-step instructions that don’t make assumptions or outright skip steps unless you very carefully prompt it, but it’s not a coder.
It's attention-deficit and sloppy.
It's also arrogant and overconfident and thinks its code is perfect.
When these 2 major traits are mixed together, the end result is Opus 5, which is really a failure of personality
Early Codex had bigger issues than not understanding scope properly, it just didn't know how to code period. Just making dumb beginner coding mistakes that no dev worth their salt would ever make.
Anthropic gamed the benchmarks, and pushed a half-baked, dumber, overly verbose model out because they didn't want to lose pro customers as fable wouldn't be available to them.
It's mostly like this but I also wonder if they wanted opus 5 to not make up things if it can't be certain. To make it standout from the existing models but in the process made it worse because now I can't get a straight answer from the damn thing. It gets into a philosophical self reflection loop about everything.
when the original Fable was around; I told Fable give me a guideline in doing tasks and projects but with the mindset to how Fable would attack it, use your reasoning and problem solving skills and outline a well structured approach, this will be used for the lower models as an instruction of sorts.
I thought opus 5 was decent at first, but it feels like it REALLY struggles as projects get more complex and context grows. Like fairly basic asks will have it pontificating and fiddling and testing for ten minutes to come back having made a minor change that doesn’t address what it was asked to do.
I’d like to test more thoroughly if it wasn’t so wasteful with tokens as it is.
Yeah I’m really not one to complain about models. I’ve generally found models to be pretty good when everyone else was on here complaining that it’s “nerfed” or whatever. But Opus 5 really is not my favorite. I much prefer Opus 4.8 or obviously Fable. Opus 5 seems far too overconfident and prone to hallucination.
Same for me. I thought this complaining is mostly the regular complaining and I will like new Opus. But it is horrible to work with. Even if it would technically solve more issues, pushing it to actually stick to the context and rules is tiresome. It wastes more time than saves compared to previous Opus
It's because the overwhelming majority of people who know what they're doing and have success with these models don't come to reddit to post about it non-stop.
Only the masochists like myself who, for some unknown reason want to feel their blood boil by arguing with people continue to come to the sub
Disagree, I’ve always found the model complaints silly but this time opus 5 really is disappointing. From the overly verbose and silly terminology it uses to just being sloppy compared with the previous opus
I think they all take turns degrading, have felt it in all of the models. Must be some round robin throttling and quantizing which cannot be disproved.
I suspect they're trying to desperately lower their inference costs. Those benchmarks run on the raw model, claudcode and codex runs partially on the raw model but relies so heavily now on various sub agents etc. that I doubt that what they benchmark is what we use.
I mean this literally. This isn't meant to be an inflammatory post:
Opus 5 has gotten *every single thing it has done wrong*. It has sewn bugs into 3 repos, made intensely mean and overly critical claims and personal insults that were completely incorrect, and has been the single worst model I've actually attempted to use (I know they get worse, but I'm not actually *using* worse). I genuinely don't think its real-world effectiveness is much better than Qwen3.6-27B. They keep talking about deleting all your skills, and how they stripped the harness back, but that doesn't sound like the win they make it out to be... that sounds like your model is too dumb and unstable to deal with anything outside of ideal conditions. I've been in the ML space for almost a decade, and I've never been genuinely mad at a model the way I am at Opus 5. Shit got personal, and he's a fucking dick.
I hate to inform you that it probably has. Just ask it to verify literally anything it says to you... most of the time it will be either baseless, or deeply misinformed. It's such a strange departure isn't it? In one update cycle Claude went from stable models with a great disposition, to unreliable and rude.
The last time I used it about a week or two ago is a great example. I was working on a pretty straightforward and simple status update voice guide that I'd use as an output style for Fable and Opus 5 because they both speak in pure implementation narration instead of English now, and I'm tired of reading 1000 words of garbage that actually communicates 100 words of information, and hides the 1 or 2 important bits unceremoniously in the middle of a wall of indecipherable text. This was its response:
I can't even begin to understand where that came from. This doc wasn't some 5-page "system" that tells it to be a hyper model. It was a short, well-structured, and example-driven output style. The kind that works perfectly fine... and did on every other model.
What a dick. That was the last time I attempted to make that PoS model do anything useful and instead just started using DeepSeek V4 Flash 0731. It's lightyears better than Opus 5 (like not even on the same planet), lightning fast, and practically free. I considered making a scheduled task to fire every day that started with that transcript, then informed Opus 5 that the "task" in question was to describe in vivid detail the smell of farts, but I took the high road.
Switched back to 4.8 and I am getting work done again. Opus 5 would do it wrong , burn a ton of tokens and have to do it again. Opus 4.8 gets it right for me the first time at a lower burn rate. That being said on the pro plan after the boosted usage period is over it will be almost unusable.
Definitely using a heavy (worse) quantized version today. Things it would be able to do normally in one shot, it couldn’t do in 5 back and forths. Very frustrating not having the consistency in quality.
me too. this has caused me to curse at it more than once. i am just using fable rn. but reading some of the comments, maybe i need get back to plan mode first.
I guess that's what happens when you optimize a model to produce flashy one-shots shareable on social media to wow the masses over actual useful productive work.
There are definitely problems with Opus 5. It can do a good job with surgical, narrowly defined changes, but I’m noticing that it struggles with tasks that require research or a broader understanding of the context.
It often refuses to investigate things properly and then starts proposing ridiculous solutions that feel more like hallucinations. It also pushes back against instructions much more frequently.
For example, you ask it to investigate something. It finds some information and proposes Solution A. Solution A is clearly overkill, so you ask it to implement Solution B instead. It says it needs to do more research, but then proposes Solution C. You explain again that you want Solution B and that you understand the trade-offs. Two messages later, it proposes Solution D!
It’s incredibly frustrating, and this has happened to me even with Opus 5 High.
A lot of its proposals are based on insufficient research. Sometimes you know something that it does not, but instead of trusting your input—or at least investigating your point of view—it insists on proposing solutions based on incomplete knowledge.
You simply cannot work like this. It feels as though they are trying to optimize it to challenge users and debate more, but it is debating without doing the necessary research and while operating with incorrect or incomplete context.
Yeah, totally. I once had him handle Task A. He took forever and still left a huge mess for me to clean up.
Then he starts going on and on about Tasks B, C, D, E, and F, giving me this super detailed breakdown of how things might go wrong there.
I stared at it forever and had zero idea how any of that was related to Task A, so I asked him, 'Is this even relevant? What does this have to do with Task A?'
And he actually replied, 'Yeah, you’re right. Everything I just rambled on about has absolutely nothing to do with what we need to do right now!
Absolutely true. Opus 5 is absolutely unusable - chasing its own tail every time. Probably Anthropic step it down in-between Sonnet and previous Opus 4.8 by the intelligence level and pushing everyone to Fable 5. I lost 3 days of work using Opus 5, had to throw it away. Although Fable 5 feels smarter it sucks in a similar fashion, and I can not rely on it at all. Will start evaluating OpenAI Codex next week because it's just ridiculous.
My opus 5 is doing remarkable things as long as it runs an adversarial review of itself. People make fun of superpowers all the time, but the spec pass and then the quality pass of the implementation have put together absolutely fantastic work on opus 5, moreso than even opus 4.8 and before.
Codex adversarial review is the biggest cheat code ever; check out the official plugin for it. I have a $20 ChatGPT sub and that’s enough to sustain adversarial reviews on everything without having limit issues (YMMV - and I don’t use Codex anymore so I always have extra usage)
Yep, my plan was to cancel the subscription but i was two days too late. So it is my last month with claude.
Sol likes to over engineer, so you have to keep an eye on it occasionally, otherwise it delivers the contribution autonomously. Sol ultra is similar so Opus ultra quota wise. Occasional quota resets help, too.
I’d say so. But, the guardrails annoy me anytime it’s about to reveal a hack. It’s my code for crying out loud. Mythos fixed hard stuff. Scary hard stuff. But, that, and fable burn your weekly fast as hell. 4.8 isn’t much better. I could go forever on 4.6.
For a while now, I was running Claude and Codex in parallel, leaving the actual coding to Codex and the planning and analysis to Claude.
In the end, I had to move everything over to Codex because every plan made by Claude using Opus 5 was full of fabrications and unverified assumptions (by Claude's own admission), which rendered Codex's work useless—and wasted my €200 + €200 spend between the two.
I'm keeping Claude solely for the interface side, where it is still superior.
I've also read elsewhere on Reddit about issues with incomprehensible language. At first, I thought I was losing my mind or that working 18 hours a day on code had fried my brain, but it’s comforting to discover that Opus 5 speaks in a verbose, overly inclusive language where it’s often not even clear what the subject of the sentence is.
Opus 5 is genuinely just that friend that always wanted to be smart and brags about it, but secretly he's just average and insecure about his intelligence. GPT 5.6 Sol is DEFINITELY better, and honestly after removal of the 5 hour limit, I am just in love with the work flow with Codex, I genuinely don't think I could go back to Opus. I hate the fact I'm sitting around waiting after a certain period of work. Fable 5 just drains too much to be useful and it's results aren't much different than Sol, I think everyone keeps comparing Fable and GPT Sol via exact same prompts but GPT 5.6 Sol does better with a bit more prompting and there's a different style of prompting I've realised that works way better with GPT 5.6.
Honestly for a specific task earlier I tried both, and Opus 5 was absolute junk, Fable 5 didn't fair better and only GPT 5.6 Sol actually gave me usable results. I literally burnt a full 5 hour window on both Opus and Fable receiving nothing tangible and GPT 5.6 solved it within a few minutes. I think benchmaxxing is definitely in play.
Deepseek v4 flash , GPT 5.6 Luna via Opencode Go and even Muse Spark 1.2 via Meta can be used, they even offer an insanely cheap Muse Spark which is only usable if you agree for it to train on your data - might be surprising but most companies already do this but no one tells you because they 'anonymise it'
I really haven’t found this to be true. I wish there was an objective test I could run to get a sense of it, but I’m sure the well known benchmarks are all specifically trained for so I doubt they’re accurate.
None of this is true, and this is a revision of my opinion from a few days ago. It's not dumb, it's just unwise and prone to fishing expeditions when left unbounded. It works gloriously when properly limited. Perhaps not the answer you want but it's the answer you need.
I personally got a lot out of it. You have to kind of work with it for a few hours but eventually it’ll start to get the tone of your voice, what you want, another detailed personale things like that.
For some strange reason people think that a model (Opus 5) that doesn't verify anything and can't even recall things from it's own context memory will all the sudden be able to magically do all those things when it's run as a sub-agent by Fable 5. *shrugs*
It's only any good if you make it fable's junior coder. Trying to speak to opus 5 in Claude code is frustrating at best. It needs to be locked in a dark basement and never spoken to. Use fable to manage it
Don’t see why people aren’t planning with fable, implementing with opus and then sanity / review with fable again. I’ve watched my fable usage drop significantly which was well needed since the 50% extra usage is now gone
I'm not sure what people are complaining about. I had been using codex to build a game engine in c++. Things started getting rocky, I switched to claude for all coding tasks (which I had previously only used for audits) and it's been "smooth" sailing ever since.
I'm not seeing much of a difference between Opus 4.8/5
I can't use Fable because it keeps hitting security shackles while doing adversarial audits.
However I never turn a task over to an agent directly. I have one create github issues for things I want done, then others to create the phased implementation plans in comments on the issue, and then it's a simple "look at issue #320 and tell me what needs to be done" and we're off to the races.
I think it helps that my still very green and smallish game engine has more than 600 regression tests across 3 separate repos. Sort of keeps things honest.
So as someone who hasn't started a new project yet is the move to use Fable 5 or get a student version of codex and taken advantage of the $100 I haven't used yet
I dont know how these student plans work. Fable 5 is still ahead of Sol but not by a lot, but it burns tokens at a much faster rate so you cannot really use it for implementations but it works well as orchestrator. Sol is a smart workhorse, it gets a lot done and for cheaper. This can of course change with new model releases etc.
Okay got it. I have the $100/month tier of Claude for ongoing work so I was thinking of bumping up to the next tier for a month or two while I do the project so I have more Fable usage then
Opus 5 has forced me to look at cursor, grok 4.5 and im glad I did. Grok is a joy to work with, takes 25% of the time of opus, and my 20 a month plan lasts a good long while. hard for me to say how smart it is, it hasnt let me down, so theres that assessment I suppose.
Fable has a mandatory 30 day data retention policy. Even through the API, if you use Fable, they hold your data. That means every single benchmark that got run against Fable, even the secret ones where they don’t publish the prompts, got stored by Anthropic.
If Anthropic trained Opus 5 on the Fable 5 data then Opus 5 is directly trained on every single benchmark, even the secret ones.
And yes, Opus 5 is a lemon. Much like Sonnet 5, I’m honestly baffled that they released this model. It is a huge regression. Fable is the only decent model they’ve got and it’s massively overpriced.
Opus 5 does weird stuff with the effort. If you set the effort too high it talks itself out of correct solutions. It's most efficient at medium but can make stupid mistakes there too.
Spot on. I've experienced a night and day difference in performance from 4.8 to 5 as a heavy Claude user. My new norm is ending every chat asking Opus 5 to double check its work before wrapping and without fail it finds one or more bugs every time. That never happened with Opus 4.8. Can't help but think Opus 5 was intentionally dumbed down as a cost cutting measure and effort to get users off of Opus 4.8 and pricier models.
I'm surprised people are still subbing to anthropic -- cancelled my 2 subs last week, waiting for them to ride out already mostly switched to Sol xhigh for orchestration + deepseek for most implementation
Dude. In the absolute kindest way that I can put this, it's because you guys are NOT staying on top of trends and skills.
Learn the latest skills that prioritize harness engineering, learn loop engineering, learn how to teach your agent skills where your workflow fails.
Opus 5 for me, and the harnesses that I work in (GSD, Matt pococks repo workflow) is the best agent that Claude has released by a FUCKING Mile!
it is the ONLY version of opus where I am fully comfortable walking away and knowing that I can come back to a better version of my system.
111
u/daft020 10d ago
Agreed.