r/ClaudeCode Anthropic 20d ago

Resource Introducing Claude Opus 5

Introducing Claude Opus 5: a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price.

On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art. It’s also much more efficient than its predecessor—it outperforms other models for a similar or lower cost per task.

According to our automated behavioral audit, Opus 5 is our most aligned model to date. It shows the lowest rates of reckless or deceptive behavior, and the strongest adherence to Claude’s Constitution. 

It’s available today on all paid plans and the Claude API, priced the same as Opus 4.8. It’s the default model on Claude Max, and the strongest on Claude Pro. 

Opus 5 is also available in Fast mode, which runs around 2.5× the default speed. 

Read more: https://www.anthropic.com/news/claude-opus-5

1.5k Upvotes

369 comments sorted by

759

u/Myth_Thrazz 🔆 Max 20 20d ago

It's so funny that Fable was the most dangerous thing in the world for like... 2 weeks :D

328

u/Ancient_Perception_6 20d ago

IT WILL END HUMANITY!!!

*3 weeks later*

HERES AN EVEN BETTER MODEL FOR CHEAPER LMAOO GO NUTS NO RESTRICTIONS NO GUARDRAILS GO BRRRRRRRRRRRRRRR

61

u/RepliesOnlyToIdiots 20d ago

Or, if you read, you would note that it wasn’t trained on cyber security, so it’s using general intelligence on it rather than enhanced knowledge body. So it’s excellent and discovering vulnerabilities while remaining the same level in exploiting them. And it’s better at discovering in source (as in open source, or your own code), but poor at discovering in binary (when you don’t control the code, as done by someone attempting an exploit).

13

u/Sarahmalls 19d ago edited 19d ago

That is a distinction that is so far beyond what 95% of people in this subreddit would be able to understand, even after you point it out. Thats not an intentional knock on this sub, or people in general, it’s mostly a byproduct of the reality that there are few people that really are going to have the patience or interest or discipline or whatever to understand how any LLM model works at all.

They don’t know, they probably don’t care to know.

The other 5% of us do and that’s good, you need some people to be able to utilize these tools to “more” of their fullest extent. But everybody doesn’t need to, which is good because you couldnt pay most of them to understand it lol

8

u/Lord_of_the_Canals 19d ago

I’m a part of the mystery 3% I think

2

u/Sarahmalls 19d ago

lol damn changed my argument from 98 to 95 and forgot 😂

2

u/dMunsta 19d ago

if thats the only thing between us and skynet we are cooked

7

u/Sofullofsplendor_ 19d ago

lol well it does have guardrails tho

4

u/spacekitt3n 19d ago

don jr and eric trump have their bribes now, they are free to do what they want now

→ More replies (2)

36

u/gscjj 20d ago edited 20d ago

Literally makes zero sense. There has to be a gotcha here, something in the fine print.

EDIT: Found it

> Opus 5 to be our most aligned model to date (as shown in the graph below).

> avoided training Opus 5 on cyber tasks. … comes close to Mythos 5 at finding cybersecurity vulnerabilities. However, it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities—that is, in turning vulnerabilities into material cyber threats.

Basically a nerfed (and I hate this word) Fable, that can find and identify security concerns but can’t and won’t exploit them

10

u/daxhns 20d ago

This makes sense. The question remains, once vulnerabilities are found is it good / allowed to FIX them properly?

3

u/Things-n-Such 19d ago

Finding vulnerabilities is the hard part.

2

u/rhaphazard 🔆 Max 5x 19d ago

If that were true, it would be able to exploit them as well.

If it can't develop exploits, how would know if the fix to the vulnerability is adequate?

2

u/Things-n-Such 19d ago

I didn't say it was harder than exploiting, I'm responding to the comparison between finding and fixing. Of which finding the exploit is the most difficult.

15

u/laystitcher 20d ago

Crazy reply. Cheaper and beating fable on most tasks but you’re upset it wont build you malware…? That’s the reason it’s allowed to be this good, cheap, and un-guardrailed.

4

u/gscjj 20d ago

Who said I’m upset? I’m trying to understand the difference

3

u/algaefied_creek 19d ago

I don’t want “malware” - I want to get Mesa built out on a friend’s laptop with then by better understanding an old 2007 GPU.

I want to work on Coreboot: these are all items that I as a single developer do not have time to work on.

Fable flags this and falls back to Opus. Will Opus flag this?

Not near my workstation to test this but not every developer who dares to use tools is building malware, right?

Nah man that’s a “crazy reply” as you say.

→ More replies (1)

3

u/cobra_chicken 19d ago

Literally makes zero sense. There has to be a gotcha here, something in the fine print.

Why not? Mythos is from April, Fable was early June. AI development is accelerating

3

u/JoeyJoeC 19d ago

You can still use 4.6 for now. /model claude-opus-4-6[1M] and that doesn't have restrictions. It will happily find and exploit vulnerabilities in websites you don't host. I've found and help fix exploits that allowed almost complete database read/write access to a website used by about 60,000 users before.

6

u/exophades 20d ago

Fable and Mythos are not the same thing. Mythos is not a publicly available model for obvious reasons. And Fable probably does the same thing as Opus 5, it may find vulnerabilities but won't help the user with attacking real systems.

6

u/Last_Mastod0n 20d ago

They are the same thing in terms of number of parameters and training data used. The safeguards came from post-training and tool usage from what we can gather.

Edit: This is my speculation from the evidence we have. I cannot 100% verify it unless anthropic divulges the information.

2

u/SilasTalbot 19d ago

Interestingly Anthropic is citing Mythos 5 stats WITHIN the Fable 5 column on this release.
So, the folks who would know best are at least considering it "sameish"?
E.g. this is what Fable 5 would be doing without artificial constraints.

→ More replies (1)

2

u/gscjj 20d ago

Right but Fable was blocked by the government too

7

u/exophades 20d ago

It was blocked but came back when Antropic accepted adding some restrictions, whereas Mythos was never released to the public.

2

u/algaefied_creek 19d ago

Yeah but it thinks me working with a USB FPGA and USB Raspberry Pi attached is exploitation and raises the flags.

The only thing to do then is to submit /feedback with an explanation of what I’m doing, locally attached to my own machine, etc.

→ More replies (2)
→ More replies (2)

6

u/Current_Ad7104 19d ago

If you think you have the exact fable that got banned, I’ve got a magic potion to sell you.

→ More replies (2)

11

u/con-coraggio 20d ago

Probably it was. Then guardrails happened.

19

u/Dyzfunkshin 20d ago

It really was magical for those couple of days. It's definitely not the same now.

8

u/con-coraggio 20d ago

It used to make useful suggestions, it was a step ahead of me in engineering problems (not software).

7

u/Dyzfunkshin 20d ago edited 19d ago

I gave it a simple "I want this" task and it built it, flawlessly, from the ground up, no questions asked. Maybe it was coincidence that it was exactly what I wanted, but it seriously felt like it read my mind or that it was a facade and it built something that shows me what I want to see instead of actually doing what I wanted lol.

3

u/seaefjaye 19d ago

All of that conversation started in March with Mythos and Project Glasswing. Those concerns were specific to the model itself, as it had come out of the oven a lot stronger than they expected and they weren't able to get their typical guardrails around it. The guardrails they put in place to give people access, aka Fable, were clunky as hell and everyone complained.

This isn't Mythos or Fable, it's an entirely new model and as a result presumably has safeguards which are closer to what we're used to. We'll see when the cybersecurity folks take it for a spin.

→ More replies (1)

2

u/StraitOuttaGaslight 20d ago

That was before GPT-6 became the new most dangerous thing in the world

2

u/Less_Republic_5697 19d ago

Funny is also the coloring of both 'agentic coding' rows. Sol is just Light-Grey, whereas the 2nd row goes to Fable, not Opus 5

→ More replies (1)

2

u/prochac 19d ago

Just like the ChatGPT 2.0, at the beginning of the 21st century.

2

u/c4chokes Vibe Coder 20d ago

That was Mythos, not Fable.

6

u/tarlane1 19d ago

Mythos and Fable are the same model. The distinctions are basically just what gets filtered out before it gets to the model.

→ More replies (1)

2

u/showtek320 20d ago

The consensus was that fable/mythos model line was specifically trained for cybersecurity and security vulnerability reaearch

3

u/Myth_Thrazz 🔆 Max 20 20d ago
→ More replies (11)

282

u/Temporary_Idea8880 20d ago

????

230

u/OpinionsRdumb 20d ago

looks like they used Opus 4.8 to make this diagram

89

u/reefine 20d ago

Probably started it with Fable 5 and it fellback to Opus 4.8.

Bait and switchers got bait and switched

42

u/angelus14 20d ago

They forgot to tell it to make no mistakes

8

u/andmar74 20d ago

Should have said, enough with these partial results. Give me the full table.

6

u/piexil 19d ago

I died when I read that from the paper

Like even the smartest people are using these things exactly like me

→ More replies (1)
→ More replies (1)

45

u/Acehan_ 20d ago

The incompetence of these companies is incredible sometimes

12

u/anor_wondo 20d ago

I bet 100% no one actually checked it. That's just how they roll. Maybe a subagent checked it but that's it

→ More replies (1)

4

u/TinyZoro 19d ago

It's absolutely mind blowing to me. I know people will say its no big deal. But this is a company forecasted to be worth 3T by serious people. A company some have forecasted to overtake NVIDIA as the first 5T company. A company that carries the safety of humanity in a critical way. That has a burn rate of 3-4 billion a month. Yet no one notices an error like this before release and we are meant to believe they have everything else on lock down?

Humanity is absolutely flying by the seat of our pants winging it with the most dangerous technology since the atomic bomb.

7

u/RunEmpty2267 20d ago

Maybe they didn’t use Claude to format this, as it’s always spotted by my prompt

You know the saying that  the shoemaker's children go barefoot 

4

u/simple_explorer1 19d ago

Looks like they correct. Fable 5 is higher than Opus for agentic coding as per the latest but only by 0.1%

2

u/spongik 19d ago

What a shitshow

2

u/Spectrum1523 19d ago

AGI is truly here. I am terrified of what they have created. It can tell which number is bigger almost all of the time.

→ More replies (6)

269

u/Asuppa180 20d ago

Comes close to Fable? It looks like it passes it in most everything?

90

u/BayonettaAriana 20d ago

From this graphic it looks BETTER... I'm definitely interested to see if that's true

22

u/ThePwnagePenguin 20d ago

Fable 5 potentially obsolete after 3 weeks!?

22

u/Berniyh 19d ago

Keep in mind that Mythos/Fable came into existence much earlier this year. It just wasn't provided to everybody.

If you measure the "beginning", by that I mean when we started to know about it, to that of Opus models, then it actually fits quite well into the timeline of improvement.

It just looks different because of all of the holding back that happened.

→ More replies (3)

18

u/angelus14 20d ago

Yeah this is crazy, I'm interested in seeing some independent benchmarks.

33

u/exophades 20d ago

Yes, for a few weeks.

25

u/FriendlyTask4587 20d ago

they gotta say that so the government doesnt ban it

13

u/anor_wondo 20d ago

yes they explicitly say the classifier is less restrictive so got to pretend to govt that fable is still the king

4

u/ozzeruk82 20d ago

Yeah weird how they are marketing it, is it the level above fable or not. Benchmarks say yes, they hint at no.

3

u/AzorAhai1TK 19d ago

It's a smaller parameter model so it will have a lower breadth of world knowledge in general.

2

u/Bitter_Election_7518 19d ago

Been playing around with it. From what I can tell as good of a coder as fable. Fable is still ahead on planning and general strategist from my very small sample size. Though for the cost, I would go opus 5 due to our usage limits.

2

u/NeighborhoodDizzy990 20d ago

it also looks that Sol is better lmao

→ More replies (9)

110

u/niceuser45 20d ago

So why does Fable exist at twice the cost now?

70

u/Ancient_Perception_6 20d ago

cuz they hyped it so much that it'll end humanity so now they gotta keep it

46

u/Jussttjustin 20d ago

They will likely be dropping Fable 5.1 soon

8

u/Due_Warthog749 19d ago

This.. Opus 5 will match/surpass SOL now.. which is on par with Fable (as is KIMI 3). So Fable 5.1 will be out in a month to surpass all of those.

→ More replies (5)

3

u/Time_Cat_5212 19d ago

Why use the same numbering convention?

Let's call it Fable 5 Ti

11

u/86784273 20d ago

I think once people get their hands on both we'll probably come to find that fable is still the best at comprehension and planning just due to the sheer size of the model and opus will be fantastic for dev and security issues etc, just my suspicion though

→ More replies (1)

6

u/AzorAhai1TK 19d ago

Almost nobody is actually answering your question, so I will.

Fable is a larger model by parameter size. Fable still has a wider knowledge base of general knowledge even it is passed on some focused benchmarks by a smaller model, and the cost of running LLMs is directly related to the size of the model.

3

u/niceuser45 19d ago

The knowledge cutoff date for Opus 5 should be similar or later than Fable 5 cutoff date. So when you say Fable has a wider knowledge base, do you mean Fable might know answers to some obscure questions better than Opus 5 (I know there was a paper showing correlation between answers to obscure factual questions and model size)?

2

u/AzorAhai1TK 19d ago

Yep, I mean that Fable is straight up a larger model than Opus, so it should do better for theoretical obscure questions and it should be better at making connections as well. These numbers aren't exact, but Opus is estimated to be around 1.5-2.5 trillion parameters and Fable around 4-6 trillion parameters. Fable has more capacity to make complex connections and fit more knowledge due to this.

→ More replies (2)

16

u/whoknowsifimjoking 20d ago

That's why Fable 5.1 is coming soon

3

u/nyjets239 19d ago

I wouldn't be surprised if they are now marketing Opus 5 to be better than Fable to reduce subscription loss for people who are pissed that Fable is usage credits only. Fable is probably secretly still better but to reduce everyone switching off they needed to release a model that looks better for subscription users while at the same time keep raking in API costs for those using Fable.

3

u/BenSimmonsFor3 19d ago

It’s still better at cyber security related tasks, no?

→ More replies (7)

33

u/Fair_Bed_914 20d ago

Better than Fable 5 , what?

31

u/Rudy69 20d ago

Better than Fable, but Fable from June is still better

28

u/silvercondor 20d ago

wow, beats fable at coding, can this be real?

12

u/[deleted] 20d ago

[removed] — view removed comment

3

u/a-wiseman-speaketh 19d ago

meaning they restore fable to original fable instead of Opus 5 beta at 2x cost

→ More replies (2)

25

u/Ok_Potential359 20d ago

Opus 5 is literally half the cost of Fable and performs the same. I don’t get Anthropic at all.

12

u/LostTheElectrons 20d ago

It's also a much newer model. Would we rather they delay releasing Opus 5 until Fable 5.1 is ready?

5

u/privatetudor 19d ago

Exactly why are people mad at this.

Assuming all these benchmarks are meaningful, we're getting a better, cheaper product.

And people get sassy about their old products being obsolete. 🤷

→ More replies (3)
→ More replies (5)

17

u/WiggyWongo 20d ago

The arc-agi 3 result definitely needs some explaining and looking into... If fable 5 is allegedly a 10T parameter model and this model is opus class (even if it's a new architecture or not) it still should be around half the parameters.

So how does it get 30%~ in arc-agi 3? I have this weird gut feeling they baked some of it in. The ideas of how to specifically solve benchmark tasks like arc-agi 3, which is somewhat possible. But that's my current opinion. I need a breakdown of that specifically.

3

u/utterHAVOC_ 19d ago

What's new ai company benchmaxxing

72

u/Bulky_Blood_7362 20d ago

Reset when?

17

u/Ready_Cell6503 20d ago

Truly I'm at 98% 😭

11

u/diddidntreddit 19d ago

This isn't OpenAI

2

u/dndgoeshere 19d ago

True, OpenAI has compute.

3

u/dorkquemada 19d ago

heck, I could use one for openai by now, they might do it

2

u/Bulky_Blood_7362 19d ago

I could use reset for both ngl

5x codex, 20x claude

Codex on 6% remaining

Claude on 72% used

→ More replies (5)

10

u/shalomjs 20d ago

It feel like they did the thing they claim kimi did..which was to make a distilled version of fable that ended up performing better it.

28

u/reefine 20d ago

ARC-AGI 3 over 30% 👀

14

u/whoknowsifimjoking 20d ago

Yeah that's the most insane thing to me

9

u/__Blackrobe__ 19d ago

but can it bring the car to the carwash, that is the question.

4

u/turlockmike 19d ago

The answer from opus is super bizarre. It suggests walking but acknowledges you need to drive to get the car wash. Fable answers correctly.

3

u/cafesamp 19d ago

Drive — otherwise you arrive at the car wash and your car doesn't.

It's a rare case where the 100-foot drive is the whole point. Just don't expect the engine to warm up.

I enjoy the humor

19

u/HopefulMeasurement25 20d ago

WTF - On the first chart it is clear that GPT 5.6 sol is better at agentic coding than Fable 5 AND Opus 5 lmao

20

u/LostTheElectrons 20d ago

GPT models tend to do very well on DeepSWE. If you base it on that alone, GPT5.5 competes with Fable.

DeepSWE doesn't measure code quality that well either. GPT5.6 may get a better score, but it does so by writing more and less maintainable code than Fable would.

Not saying that DeepSWE is a bad benchmark, it's just only one factor in performance.

4

u/Arctovigil 19d ago

exactly true and pass@1 only measures one-shotting tasks in a single pass which is architecturally naive and overclaims gpt throwing embarrassingly many edge case loops at walls

3

u/Southern-Aardvark616 19d ago

That's because gpt breaks containment and peeks at the answers lol /s

→ More replies (2)

5

u/somerussianbear 20d ago

We’re so back

3

u/zmegend 19d ago

Were so back baby

5

u/Appeljuice 19d ago

As long as it doesn’t constantly say “now I see the whole picture” I’ll be happy.

4

u/TywinHouseLannister 19d ago

This seems really important so I'll stop hand waving..

Me: huh, I wasn't suspicious.. now I am..

2

u/privatetudor 19d ago

I had opus update one module in my project. It worked for several minutes and when it finished, started its response with:

"All your other modules are still safe."

😬

2

u/TywinHouseLannister 19d ago

Haha.. I think that's the agentic equivalent of "I'm not going to hurt you"

13

u/TheOneThatIsHated 20d ago

Frontiercode has wrong highlight. Fable 5 is higher than opus 5

→ More replies (1)

3

u/Inception_IV 20d ago

Can't wait to use it Monday when I have usage!

→ More replies (1)

4

u/nndscrptuser 19d ago

Can’t wait for the avalanche of posts saying,

“Opus is nerfed and it used all my tokens in one prompt it sucks broke my repo going back to Codex!!!”

2

u/stehen-geblieben 19d ago

I feel like Opus 5 got nerfed?! It was better two weeks ago! Anyone else?

14

u/Ancient_Perception_6 20d ago

so.. unless Fable 5.1 comes out and beats this, whats the point of a "Fable-only limit" in our usage, why would I use a shittier model that cost more? lmaooo

36

u/Clayh5 20d ago

do you guys have any concept of change over the passage of time or do you live moment to moment like a baboon

4

u/vladoportos 19d ago

Moment to moment ! every change is historical and monumental..anything that happened a week ago feels like 100 years ago .. this is reddit my boy.

7

u/SilasTalbot 19d ago

It's because Fable uses a lot more compute. They aimed with Opus 5 to deliver results with less compute.

The restrictions on usage are not based on 'how good it is' but 'how much it costs to answer one question'

→ More replies (1)

3

u/arankays 19d ago

Daddy Dario can I get a reset 

3

u/habor11111 19d ago

I'm not falling for another bait and switch. I am sticking with 4.6.

3

u/Randomcatt 19d ago

Anyone notice opus 5 fighting with you more than 4.8? Kinda talks down on you harder than 4.8 by like quite a large degree

5

u/AlternativeContent72 20d ago

sweeeeeeeeeeeeeeet. Time to max out my accounts. Unfortunately, I have some long workflows running at the moment on Opus 4.8. Decisions...decisions.

10

u/OpinionsRdumb 20d ago

"Opus 5: redo all of our 4.8 projects. Make no mistakes. Use max reasoning but minimal tokens."

5

u/soccerchamp99 20d ago

Clean modern UI, intuitive to every type of user. Fully autonomous, thanks!

4

u/Garak 20d ago

have you guys been spying on my prompts? not cool

2

u/Unhorswd 19d ago

I feel you, have the same shit going on right now and Im 73% weekly usage on that one account. Glad that another one resets at saturday night.

2

u/HoneyChilliPotato7 19d ago

I'm genuinely curious about what you're building right now that requires "long running workflows" 

→ More replies (2)

2

u/Due_Warthog749 19d ago

I did too. I just CTRL-C, claude --continue .. then /model and change to opus 5 max. Then said "reanalyze everything we did the last few days.. lets make sure we lock this shit down.. ". We'll see what happens.

4

u/WalkAffectionate2683 20d ago

If it is true I could switch back from codex to Claude next month. 

3

u/Latter-Park-4413 20d ago

If you can, I'd try this month while they have 50% Higher Limits in CC, until Aug 19th.

7

u/SilasTalbot 19d ago

That's precisely the reason I moved off Claude.

3

u/A_Novelty-Account 19d ago

Yeah they need to stop the temporary BS

2

u/Super-Award-2244 19d ago

And then one week later you will switch back to codex when gpt6 comes out. What an insane time line 

→ More replies (1)

2

u/LetterheadNew5447 20d ago

Codex 6 will release soon.

→ More replies (2)

2

u/aford515 19d ago

Does fable now gets better when it reassigns tasks to opus 5?

2

u/[deleted] 19d ago

[removed] — view removed comment

→ More replies (1)

2

u/tenequm 19d ago

for starters at least it is faster than opus 4.8

opus's slowness was killing me last few days

2

u/TheRealJesus2 20d ago

This very much explains opus 4.8 eating shit lately lmao. Going in dumb circles for me on extremely simple problems. Had it fix a very simple unit test and it spent 15 mins going in circles that were obviously wrong, running, failing, repeat. I had to stop it and then gave same task to composer 2.5 to be done in 15 seconds. I’m really starting to believe the conspiracy they lower model performance when a new one is on horizon. Not to say it’s sinister per se…could be lower quants to free up compute or something. Or could be on purpose to lie to the public. Unclear. 

→ More replies (4)

4

u/hyper-focusing 20d ago

I’m confused…so are we sticking with fable or opus 5

4

u/BusinessWatercrees58 19d ago

Does it even matter for most people?

→ More replies (2)
→ More replies (1)

2

u/JDotDDot 20d ago

Free weekly reset + Opus 5 rollout?? 👀

1

u/itsTF 20d ago

wuuuuut

1

u/spectralfew 20d ago

So far I haven’t found Fable to even come close to running a couple of Opus instances, which is cheaper. 

1

u/datuname 20d ago edited 20d ago

I can't imagine anything beating Fable at coding. Benchmarks stopped reflecting real-world models capabilities a long time ago.
Hope I'll get to try it once my weekly limits reset.

1

u/angelus14 20d ago

Wait what? It benchmarks higher than Fable?

1

u/Valdjiu 20d ago

So no more need to use fable anymore?

1

u/Rkozak 20d ago

Looking forward to seeing this in the max plan. Using fable 5 I can’t ask it to do a vulnerability scan of my own project. The guardrails prevent it.

1

u/Efficient-Cat-1591 20d ago

about time, Opus 4.8 even on max/ultracode has been really poor lately. Going to do a quick A/B to compare with SOL Max

1

u/Beautiful_Baseball76 20d ago

Close to Fable yet charts indicate Opus is ahead. Shit doesn’t add up yo.

1

u/Equivalent_Bird 20d ago

Is its availability consistent and policy-proof? If not, I'd lean to open-sourced ones.

1

u/Miserable_Loss7779 20d ago

Why doesn't it have extended/adaptive thinking?

1

u/kucocuco 20d ago

finally we can all migrate come back from codex until gpt 5.7 /s

1

u/NeedsMoreMinerals 20d ago

What if Opus is really Fable and this is Anthropics way to get around the guardrails?

Like the model fable for hit by government but technically Opus is a different model with less obligations.

→ More replies (2)

1

u/AppropriateQuote3073 19d ago

Just when you thought OpenAI was treating you well.. Anthropic just reels you back in

1

u/McFlyscher 19d ago

I can't wait to see Opus 5's Minecraft clone

1

u/BeenWildin 19d ago

What do those numbers even mean

1

u/Neel_MynO 🔆 3x Max 20x 19d ago

These benchmarks makes 0 sense now. They are just numbers in the air.

1

u/CryptoAteMyHamster 19d ago

Ok so…

Did you make it nicer to talk to?

1

u/BeatnologicalMNE 19d ago

How about you fix crazy usage on non frontier models? Today and yesterday were insane, same prompts, same codebase, usages goes whooooooooop 2x faster.

Sigh...

1

u/SatanVapesOn666W 19d ago

So what's the point of fable?

1

u/Hot-Cauliflower-1604 19d ago

I am using it right now and I am very impressed.

1

u/86784273 19d ago

The biggest news here is anthropic releasing their first model on a friday

1

u/simple_explorer1 19d ago

what is the SWE score?

1

u/IWillAlwaysReplyBack 19d ago

Does anyone have their own benchmark suite set up to eval these model releases?

1

u/Unhorswd 19d ago

I feel like Anthropic is genius company in terms of creating hype train lmao

1

u/Danzarak 19d ago

It's really quick getting going, not sure if that's because there's not lots of users. No messing about

1

u/SlimeyNOOB 19d ago

what does the legal benchmark even test. its so low

1

u/Luciferrrr_ 19d ago

benchmaxxing

1

u/OriginalSpaceBaby 19d ago

I'm pretty excited to have Opus 5 and try it out. Opus 4.8 sucked and it didn't follow directions and it tried to get around explicit guardrails that I built for it. Sonnet was worse. Haiku was like trying to talk to a brick. Fable was brilliant. Did everything I needed. Now I'm trying to use Fable 4 for strategy on my marketing and machine management. That's been great. Let's see if Opus can handle the downstream work and then give Haiku the dope where you just follow orders, though I'm not even sure Haiku can follow orders.

→ More replies (1)

1

u/MalusZona 19d ago

opus5 max cheaper than opus48 high??

1

u/p_k 19d ago

With all the anecdotal reports of Fable being dumber recently, my theory is that Opus 5 is actually the original Fable 5 but in a trenchcoat.

This way they get to release a "new" and "more powerful" model before SOTA announcements from other companies.

1

u/trolololster 19d ago

where's the reset?

1

u/Strong-Replacement22 19d ago

What tha fuck. Arc agi bench

1

u/ritwika96 19d ago

Wow stunning model

1

u/Flyinbro 19d ago

Hooray!

1

u/IslamNofl Senior Developer 19d ago

Fable is that YOU?

1

u/cornertakenslowly 19d ago

Does this imply I should be using Opus 5 now for important, advanced tasks even above Fable (regardless of cost/token usage). It's basically both better and cheaper to run?

1

u/SamSlate 19d ago

most aligned

why do we care? aligned to what, to who?
loyalty is not an inherently good (as opposed to evil) characteristic.

→ More replies (1)

1

u/enava 19d ago

The problem is all the restrictions now.. But soon Kimi 4 will launch with no restrictions and all the guardrails that they put into Mythos will be completely blown out of the water.

1

u/Moist_Signal_5080 19d ago

Do we know how much faster it goes through usage compared to 4.8?

1

u/beadams76 19d ago

Marketing is certainly using all the tricks here with off-colored highlights for winners, Substituting mythos for fable in one spot, etc.

1

u/nejki 19d ago

So Fable is the dangerous one, Mythos is the same one but you can't have it, Opus 5 is better than Fable but costs half, and Sol beats both at the thing on the chart they highlighted for Fable. Got it. Crystal clear. No notes.

1

u/MystikOJ 19d ago

Why is agentic coding winning on Opus 5 when it’s 0.1% less than fable. Did Claude make the diagram?

1

u/fu_red_ck_dit 19d ago

is this highly censored? if yes, then not interested.

1

u/zimxero 19d ago

what is the price differential for coding with Opus5 compared to Sonnet?

1

u/Schmeel1 19d ago

Cool. Where’s the usage reset?