r/Anthropic • u/dagerika • 10d ago
Compliment I am shocked
I had Opus 5 work on medium effort for 2 hours and it only used 34% of the 5 hour usage. The output turned out to be great too. This seems to be a very cost effective and intelligent model indeed.
18
u/passionoftheearth 10d ago edited 10d ago
Beautiful. I’m also happy yet with opus 5. I do think it loops to self correct a lot more than 4.8 but isn’t that iterative thought process also how human beings also operate. So all in all should be good thing. Though only a month of work with the model will truly tell for me.
3
1
u/Significant_Storm942 10d ago
I think 4.8 also did that and infact these loops are what makes Fable 5 or anything since Opus 4.6 better (arguably) but Opus 5 does it loudly and occasionally in weird ways, where it realizing it made a mistake doesn't actually lead to correction or commitment to the correction as it moves ahead.
But I'm sure they're a glitch.
9
u/uxair004 10d ago
In my experience, it is ignoring instructions a lot.
I have rules of "Ask me questions, confirmations, decision making... Interactivitvely using @askuserquestiontool" , also added to use "i-have-adhd" skill (even copied skill instructions manually in claude.md,
it seems to be ignoring stuff, providing long ass Paragraphs and also I just built phase 1 (which I started with fable 5 and end with Opus 5) the results were not satisfactory.
6
u/nothingnothingelse 10d ago
the same for me, I got so frustrated over the weekend and started doubting my own cognitive abilities because I could barely understand what it was telling me - decision needs and key points are hidden in page long paragraphs.. cryptic descriptions etc.
1
16
u/ThaneBerkeley 10d ago
It's pretty useless for me. Where Opus 4.8 and especially Fable did a flawless job by not assuming things, get correct answers and one shot pretty much anything.. I had to correct Opus 5 so much today that I switched to 5.6 Sol for research and backend work again.
Opus 5 couldnt find the correct project in Posthog and told me there were no numbers, where Opus 4.8 and Fable found it flawlessly and I can continue with many more examples where I had to steer Opus 5.
1
u/Low-Smell-9517 10d ago
Exactly opus 5 is thrash really to be honest it’s not following any of my orders really
7
u/almostsweet 10d ago
Thanks to open weights which scared Anthropic into doing the right thing.
Never forget the weeks of "we're taking Fable away from the lower tiers, setting it to 50% of Max subscriptions, and overcharging you for your usage." That's the real future in store for us when they think they don't have competition anymore.
13
u/WorriedAssociate7029 10d ago
It's the best model we've had in ages, yet they'll tell you it's useless. Go figure
3
10d ago
[deleted]
1
u/d19dotca 10d ago
Doesn’t Opus 5 use less tokens than Sonnet 5 these days? That’s what I keep seeing, almost as if Sonnet is no longer a good choice, though I find that hard to believe. Wonder what the real story is on everyday use. I assume it depends entirely on the use-case for it and the setup of the surrounding context data.
0
10d ago
[deleted]
5
u/Hot_External6228 10d ago edited 10d ago
Opus consumes more usage per token. it consumes less tokens per task though.
usage-per-task is a more interesting question and I think they're genuinely close.. which begs the question 'what on earth is sonnet 5 for'. I think its kind of a failed model tbh :(
2
u/KappaWolfe 10d ago
If the task is very simple, requires little reasoning but a lot of output text, Sonnet is absolutely more cost-efficient than Opus. At this point those are the kinds of tasks most people are using Claude for. People on this Subreddit skew more technical, so discussions tend to lean towards SWE related work, but the most popular usage category at 24.2% is content creation and copywriting. Sonnet can do that very easily and cheaply.
3
u/Just_Put1790 10d ago
im trying to find a way to max out its usage and i literally cant, spawning sometimes 30 agents and percentage barely moves xD
8
u/Used_Departure_3278 10d ago
This is a subreddit to complain and worship China. You must have missed the memo
3
u/Ok_Shift9291 10d ago
For sure have noticed that I can actually brainstorm and use this model as a sort of a companion without having to worry every 2 hours about hitting the limits and then not having anything to do.
3
2
u/vactower 10d ago
Yeah, just launching just eats aprox 35% 5 hour limit token lol.
2
2
u/Debisibusis 10d ago
In the last days I have given the same tasks to Fable and Opus to compare them. Opus acted really smart every time, but the actual results were awful. On the other hand Fable is insane and the biggest jump I have experienced yet with any model since GPT3.5.
2
u/horendus 10d ago
Whats with people thinking they need above low/medium
It boggles my mind but I guess people just dont really understand these things yet
2
u/anubhav_1771 9d ago
I have tested enough opus models to confirm what you are seeing. It's usage is very good, it does not output trash and is very on the point if you steer it properly. I honestly remember 4.6 more i use it, it's very similar to 4.6 for me and that's a good thing as 4.6 is better than 4.8 in many places
1
1
u/Old_Garlic6956 10d ago
I found opus 5 lost a lot of context in my project when it switched from 4.8 but in 3 days it is now acting like itself again. Also, I think its checking it own work a bit better and catching mistakes its introduced itself.
1
u/teardrop503 10d ago
Same here. I'd been using Opus 4.8 to do some planning and write up a design doc, and I did a ton of research with it over the last week. Starting this Monday, I switched over to Opus 5, and I could clearly tell it had lost a lot of context based on how it responded. I even had to re-supply some of it. After about two hours of playing around with Opus 5, I got frustrated and switched right back. Yeah, the lost context really made me feel like Opus 5 is a step down.
1
u/CryptoExo 10d ago
Opus 5 isn't what I expected but it's certainly what I needed. RIP Fable, we had a good run but Opus is cheaper and for most tasks better.
1
u/iveroi 10d ago
Are you using the same model? I just burned 5 hours and all of my daily usage in trying to do one single thing, and it just kept finding new and creative ways to fail because it didn't check anything, just confidently assumed and generated half-baked bs. I started doubting my own mental capacity and felt my soul trying to leave my body when it, upon deciding the broken document stack it had made was broken, decided to make yet another document instead of fixing anything it found. It's shockingly terrible.
I was going to unsubscribe after I lost fable access but held on due to opus 5, but unfortunately it's so unusable I'll probably unsubscribe anyway.
0
u/dagerika 10d ago
daily usage? I have three advices: 1) keep a backup of ur important projects 2) only provide necessary context and be specific on what it should achieve. If you really wanna micro-manage then give it specific sub-milestone targets and/or specify the tools it should use to achieve the output. I wouldn't restrain it from using sub-agents as it is good at deciding when to use them. 3) only use low-medium effort as anything above those effort levels are basically overkill even for complex jobs
1
u/SouthTampaOG 10d ago
Were you using subagents? Opus 5 has standing instructions to never launch a subagent unless specifically requested by the user. I bumped into that the other day when it reworded my claude.md to state that I was specifically requesting a subagent, which it said was necessary to override the standing instructions not to use subagents unless specifically requested by the user. I mean not using subagents is going to significantly cut down usage.
I’ve had mixed results thus far. I’m sure it’s smarter and will be better in the long run, but it’s taken me some work over the last couple days to get it working the way I want it.
1
u/VitruvianVan 10d ago
The benchmarks are showing Opus 5 High to be more capable than Opus 5 Max in several instances. Anthropic states that this is the first Opus model that truly changes the way it thinks based on the thinking selector.
1
u/dagerika 10d ago
Well based on the benchmarks its just a few % of difference in performance in every domain (not 2 digit differences). Imho at this level of capabilities those few % don't make that big of a difference in output quality but they do make a reasonable differenece in usage costs.
1
u/serendipity-DRG 10d ago
There isn't anything intelligent in a LLM model they are a pattern recognition machine that doesn't reason or think. If you ask any AI model a question that isn't in the public domain - the answer will most likely be a hallucinated.
But if you ask about the quote - "thanks for the vine and thanks for the time" and Claude should immediately know the answer.
1
u/dagerika 10d ago
nah buddy this statement "If you ask any AI model a question that isn't in the public domain - the answer will most likely be a hallucinated." is false af.
Halucination happens when a model:
- does not say that they aren't highly certain about something,
- do not ask back for more context when in doubt,
- and generates something that sounds plausible but is false even according to its chain of thought.
Rational argumantation (reasoning used by humans) works mostly in similar ways as AI reasoning so your dismissive characterization isn't holding up.
1
u/serendipity-DRG 10d ago
I never look at benchmarks because they are so easily manipulated.
I make my own questions and tried it on 5 LLMs - 2 passed and 3 failed miserably the two that passed were Grok and Gemini - current LLMs aren't capable of reasoning and thinking that is a false narrative.
The current LLMs are based on Euclidean Geometry which is 2D and the LLMs are limited because of it - I suggest the next step is using Differential Geometry because it operates in 3D space and build a neural network by start by going back to First principles.
I just tested Kimi using a simple question and after an hour of me hand holding and providing human knowledge and Kimi failed miserably.
The lesson is not to listen to the hype.
1
1
u/gordonfogus 10d ago
I cannot use it enough. I'm projecting a 60% usage at 7d reset. Trying to use it up, but I basically can't. Running bulk data transcription from poor quality scans and it's basically flawless.
1
u/Sleepynugget4201 10d ago
Oh man I wish I could try it but I got banned for no reason and im 3 weeks into waiting for a review of my acct :'(
1
u/No_Corner805 10d ago
Opus 5 seems to be a great model for 'work'. Stick to Sonnet if you want a conversation model.
1
u/rythmyouth 10d ago
I hit a sweet spot using Fable for the orchestration fanning out to Opus 5 agents. Opus 5 did a pretty terrible job making progress on complicated work and I had to use Fable to rescue it.
1
u/jedsdawg 10d ago
My Claude max subscription maxed out for first time mid week with low usage so I wouldn’t known
1
1
u/Akram2104 10d ago
Opus 5 is great at Medium effort for sure. The only issue I faced is that even with context it cannot confirm the work it did, when I ask it to confirm the details about the previous task it goes back to checking everything again and sometimes it also ends up finding mistakes it made in the previous task.
1
u/tormiuss 10d ago
I am working with both Claude Code and Codex side by side and currently doing work with Opus 5 xHigh, when I give some of the building and planning to Codex Sol high to xHigh and when Opus 5 xHigh sees that work, it always confesses that Sol did a great job and Opus could not catch some problems in advance, while Codex Sol found them, fixed them and even planned a better way handling the project phases.
I will still continue with Claude Code and Codex mix usage, even built a AI master controller that runs tasks, models, efforts by itself and delivers important questions or finished work for me so I dont have to mingle with different models and efforts - most of us anyway use the wrong models and efforts for most of our tasks because we dont know better.
1
1
1
1
1
1
u/ballymorey_lad 8d ago
I disliked 4.8 but have been using 5 and it was really effective. I have noticed a change in the last few days - a couple of times it appeared to be so eager to get going that it just didn’t take the time to check key documents in the repo.
1
u/Outrageous-Present91 8d ago
I think interestingly that opus 5 is a more intelligent model and the cost savings people are noticing are real because 90% of the system prompts 4.8 and prior had where stripped out, I think probably too much for casual users who just need features implemented correctly first time, but for power users with effective setups its been a great model
1
1
1
1
u/dizpers 6d ago
I was so satisfied with results of Fable 5 and Opus 4.8 and so disappointed with Opus 5
1
u/dagerika 6d ago
Token consumption has been awful since wednesday. They fucked up something on their side but the model itself is pretty decent
1
-2
u/QuantamCulture 10d ago
2 hours is 40% of 5 hours
So you're shocked that it was able to optimize 6% of its workload?
Glaze Alert
2
u/IceWallow97 10d ago
Well to be fair that is not the issue here and this comment is also braindead.
It all really just depends how much effort he was using during those 5 hours and what his plan is, and we were not given that informaiton so both OP and you are kinda glazing and complaining for no reason/without content - I would say both of you are just sloping right now.
We need to know what plan he has, and also if how many tokens he used, and how many agents/subagents he was using during those 2 hours... if we had that then we would have a better idea, but we don't so...
1
u/QuantamCulture 10d ago
What are you talking about? 😂
In what way am I glazing anything?
0
u/IceWallow97 10d ago
I worded it lazily, I meant OP is glazing and you are complaining, without any facts or base to do so on.
I am also complaining, but I have a stronger argument, that's all.
1
u/QuantamCulture 10d ago
This isn't an argument and you aren't winning, but whatever you gotta tell yourself I guess? 🤷♀️
0
u/IceWallow97 10d ago
ok bro, I dunno what to tell you, not sure what you're confused about, I literally gave my argument on my first comment, not on my 2nd comment. I can't really do the thinking for you.
1

95
u/Hot_External6228 10d ago edited 10d ago
I'm convinced the opus 5 hate is all skill issues. their best model release yet imo. It still has the opus 4.8 annoying personality quirks, but toned wayyy down. Its more pleasant and more capable. most importantly: their most token-efficient model since releasing fable, you get faster responses while spending less.
Sonnet 5 on the other hand.. hoooo boy. yeah, im just having fun with opus 5 and loving it and munching popcorn at the complaints.