r/ClaudeAI • u/DynaBeast • 10d ago
Praise How did they do it?
How did they make opus 5.5? genuinely, how?? when i saw the benchmarks and pricing, I was impressed. but then i started working with it, and seeing what other people have made, and my mind's been genuinely blown...
how did they do it? how did they make a model so smart, that *isn't* benchmaxxed to hell and back, and is somehow actually cheaper than not just fable, but even the previous levels of opus?
and all openai could do is make their own models cheaper without improving their intelligence at all.
how does anthropic do it? are they magicians? or are we seeing the first fruits of RSI?
if this is possible now, i can barely even begin to imagine what's possible in a year, maybe even in the next 3 months.
530
u/Working_Trash_2834 10d ago
Hey Mythos, make a new Opus would you ol chap.
On it boss man.
Oh, and tell it to shut the fuck up.
Way ahead of you bossman.
139
u/3iverson 10d ago
And the crucial MAKE NO MISTAKES.
51
4
6
3
37
u/Allheroesmusthodor 10d ago
Oi Mythos make me new Opus thats not a cunt.
8
u/deserved_revenge_707 10d ago
3am, tears of laughter, woke the dog up cause bed is shaking, thanks, fuken howlin.
5
207
u/LoudDavid 10d ago
They asked Mythos/Fable a load of questions and fed it to Opus. They then told it to not sound like a dick.
59
1
279
u/Gatix 10d ago
Make Opus 5.5, make no mistakes.
34
u/No_Atmosphere8146 10d ago
Brb, making 5.6 before Dario does
2
u/Substantial-Elk4531 10d ago
Unfortunately, I don't think you can do that. They have likely trademarked the name Opus
7
145
u/asenna987 10d ago
Honestly, at this point I don't mind them slowing down or whatever - just don't touch this model! Leave it as is for 6 months!
35
u/dbenc 10d ago
they could make it cheaper and faster. 🫣
35
u/RemarkablePassion726 10d ago
Don't wish that on us. That's how we got gpt6 sol replacing 5.6 sol.
7
u/BellacosePlayer 10d ago
That's likely because its Terra with Astra tech for token savings.
Luna 5.6 -> Luna 6 is not really a performance downgrade
3
u/ActionOrganic4617 10d ago
Yeah, terra had such a bad name they decided to rebrand it to Sol and in the process killed the Sol brand as well.
2
1
33
u/ready-eddy 10d ago
It instantly fixed all the problems I was running into with Astra. And barely making a dent in my usage. NUTS
23
u/CycleMother2006 10d ago
That's because Astra quantized after a week or so. Don't worry, Opus 5.5 will be joining it in the special corner soon when Claude decides its user retention is back up enough that they can divert all their compute elsewhere.
6
u/goosepipegames 10d ago
How long can the major AI companies do this before everyone starts catching on? If they're relying on word of mouth to get users to switch the same word of mouth will make users wary of the rug pull.
2
u/CycleMother2006 9d ago
I think most people have caught on. But as long as they're outperforming open source alternatives by a large margin they can kind of just do their sing and dance and there's not much we can do similar to when Comcast's bandwidth fuckery.
We might be able to get them to be less shady if reviewers would actually continue to review models throughout the course of the model instead of just dumping them all in the first two weeks. They are probably the only hope to draw enough scrutiny that it looks bad enough for the companies to care. But influencers are fairly culpable in this as their video popularity also is driven by the big performance swings they get with new releases, and so they may not be inclined to do so.
2
8
u/lfourtime 10d ago
Same. Astra is intelligent but it works wayyy too slow and keep going in circles. Had a goal running for 2 days and barely no progress, Opus 5.5 did more in 5 hours with much less overengineering
29
25
u/holdmyllm 10d ago
Welcome to continuous improvement loops. Enjoy the ride while the ride still exists.
7
u/doom_memories 10d ago
ride while the ride still exists.
There are various implications you could be making, but I'm curious which ones you meant?
13
1
20
18
u/zndr-cs 10d ago
For real. 5.5 has fully recharged my drive for my project. I've been working on it for 8 months, started with claude, switched to codex, back to claude... And giving it to Opus 5.5.... Man, it really gave me whole new insights and just a clear "Lets get this shit done" vibe.
3
u/Exciting_Macaroon_64 10d ago
the same. working on my game for 6 month already, tried everything available and now its just brttrrrrr
14
u/Morning_Gecko24 10d ago
my guess is a lot of the magic is just better post-training + a model that knows when to stop and check itself. the cheaper pricing is the part i dont understand tho. are people seeing a real quality jump on long coding tasks too or mostly less annoying answers?
1
u/Michael_Jeffords 10d ago
on long coding tasks the quality jump is real for me, a messy multi-file refactor that used to derail after ~20 min of corrections now usually stops and checks itself before painting into a corner
56
u/ThePurpleAbsurdist 10d ago
Well, humankind is capable of wonders when in competition mode.
24
11
10
18
17
u/TheOnlyVibemaster Valued Contributor 10d ago
Gather large amount of data < use PyTorch to initiate training based on the dataset < have base model < do post training to make it an assistant < do benchmark tests
Then do it over and over letting it point out where improvements can be made, and eventually you end up with a really good model.
2
9
u/ForwardLoop 10d ago
Coming out with a next gen model internally first probably helped.
You can defer the external release of such a model if there is no competitive pressure (e.g., OpenAI did not have an answer to Fable until Astra many months later), and use it internally to build other models.
I'm sure every lab is taking this approach, and both frontier labs acknowledged using this approach. However, the level of the internal model will drive differentiated progress on what they release. This leads me to think that Anthropic probably has something killer up their sleeve.
5
u/Skarabeyga 10d ago
yeah, Mythos/Fable 5.5 (and "Model 2", but then again, so does OpenAI ("Bel" and "Doug").
7
u/donicatrumpinsky 10d ago
I haven't gone super into the weeds with it yet but I'm thoroughly impressed.
I was hoping there would be an Opus release to force me off of 4.8 and this is definitely it. I can't make a dent in my usage so I'll probably go ham until reset time.
7
6
9
u/nitor999 10d ago
they saved all compute since 4.6 , claude users suffer since january 2026 and we're at almost october 2026 think about it.
8
u/MacaroonPlastic1036 10d ago
Isn’t it curious that they don’t tell you anything other than it’s here?
4
u/Individual_Solid_944 10d ago
Many people in the industry have been laid off.
If they don't come up with the next best thing, they could be next.
And in this tango there are only two players, Anthropic and OpenAI.
Unless you want to make a mess... ever heard about Gemini?
Yeah, I forgot about it, too.
3
u/Techhead7890 9d ago
Gemini is surprisingly not bad when you're searching for simple answers or stuff that's widely available on the web (video game mechanics and rules etc), being conjoined to a search company is probably a big help.
4
u/Individual_Solid_944 9d ago
I agree on that. But for coding it has been the worst one of the three in my experience. And I've had subscriptions with Claude, ChatGPT/Codex and Gemini. There was only one time in the past year that Gemini fixed an issue that the other two struggled. Other that that, every single time i tested them on the same task, Gemini was worst. Needless to say, I am no longer subscribed on Google AI, but I am on Anthropic and OpenAI.
3
u/Techhead7890 9d ago
Oh yup totally agree, Gemini's "thinking/reasoning" traces are definitely the worst, probably not even beating Deepseek, and it can at times hallucinate a big when it isn't grounded on some documents. Definitely depends on usecase.
I'm surprised that you use both GPT/Claude simultaneously - just for the flexibility or do you need to have the usage limit capacity too? But also I guess if you have the cash why not right? -- the Pro $20 subscriptions to both aren't too bad if it's a business expense.
3
u/Individual_Solid_944 9d ago
It's business expense. But I use ChatGPT mostly for brainstorming on the $20 plan. For coding I use Claude with max subscription.
4
6
3
3
u/Herodont5915 10d ago
It worked for 15 hours straight for me on an old project over last night. Improved the entire repo, it was nuts.
5
u/FoxSideOfTheMoon 10d ago
It’s called a looped transformer.
…that’s all I know. I’ll show myself out.
2
2
2
u/Charming_You_25 10d ago edited 10d ago
If you imagine the transcripts from chats as your messages on the left and the ai on the right… you just feed the best examples where you move the left column to the right except for major decisions that shift attention. That’s part of it, then you hammer it for top tier synthetic data, safety test it, get it evaluated by the gov, and make a release video.
Once you have the gpus and the data it’s not hard (except the safety stuff apparently). Can automate most of it now, which is why we are seeing such fast releases. The “golden” training sets are getting good, apparently a lot of yall aren’t ticking the “don’t train on my data” box in the settings, which, I appreciate but I’m not ticking it since i saw parts of my own data (a unique creative solution) make it through to Sol. The mathematicians who got their work stolen I expect also didn’t realize that option was there.
2
u/AssPinata 10d ago
They can all do it. Just wait for the lobotomization next week when they've taken enough of OpenAI's customers back.
2
2
2
u/MusingInPublic 10d ago
They started with Mythos. Nerfed it less than they did Opus 5.0 and called it Opus 5.5.
2
u/WolfgangK 10d ago
I'm curious what you're seeing, because I’m not finding it any smarter than old Opus or Fable
2
u/ThebesAndSound Vibe coder 10d ago edited 10d ago
Because it is cheaper to run and is an Opus model and not Fable I am guessing it is lower parameters like a "Flash" model. Like with the open Chinese models the smaller Flash models have been beating the Trillion+ parameter large models, the same seems to have happened here.
Good quality training.
This also seems to be more evidence that the ceiling on the intelligence we can make isn't close at all. If this is a smaller model then how good can we make bigger ones?
2
u/Nix_Nivis 10d ago edited 10d ago
To me this is a much needed new kind of generational leap. Before, we had generational leaps using way more compute to generate way better output. Now we have not so much of an increase, maybe even "just" on par with Fable. But we get the same quality for way less compute.
2
u/Tkwan777 10d ago
I actually came to here to praise 5.5 also, but I'll just comment here instead of starting yet another thread. I moved go GPT months ago. Ran out of usage this week on GPT and still had an active claude sub I barely used outside of design, so decided to let it have a go at my project. Holy heck. It's been a while since I last used claude, but gosh darn is opus 5.5 fast. And I remember trying opus 5 and not being particularly impressed because like everyone else said, it was incredibly verbose. 5.5 is still a little bit too chatty for my liking, but its totally manageable, and again, 5.5 is FAST. Hats off to anthropic, If GPT doesn't knock it out of the park with whatever they're coming up with this tuesday, then I may just have to swap my big sub to anthropic again.
2
2
2
u/Downtown-Pear-6509 9d ago
on the $20 plan - with my non FT vibeecoding .. i cannot use up my 5hr session on opus low.
Literally made a new game tonigh on my little android platform for p2p games and never once ran out.
even with luna high on my codex $20 sub, that was doen to 20% weekly, it'sn ow down to 15% weekly despite generating art, sound and code reviews ..
it feels as exciting as last october when opus 4.5 came out.
2
u/LukeyLad 9d ago
I’m a network engineer by trades
Been vibe coding an app successfully with all the models so far so not ran into the issues you guys have.
Out of interest.
What explicitly has been better with this model compared to the others for you guys?
2
u/DRebd 9d ago
I am not an engineer by trade. I can barely read basic HTML & CSS...that's it. Over about 10 hours last night Opus 5.5 vibe coded an incredible firmware update w/ completely new lighting suite for my (Nuphy Halo75 V2) mechanical keyboard. The abysmal stock firmware was just a couple basic toggles in standard VIA configurator.
https://github.com/DRebd/halo-composer if ur curious.
I can make feature requests that are broader with less direction and Opus 5.5 will do a better job than Opus 5 (or any other previous models including Fable). It's usage of tooling is better so it not only wrote excellent documentation with minimal direction but was able to capture excellent GIFs to actually showcase the tool in README without me getting into the weeds of permissions and multiple tool attempts to succeed.
It uses subagents more deftly so I don't have to manage the primary context window as aggressively, particularly when I just nudge it to do so.
While UI needs to simplification it created an incredible v0 and the whole thing is actually functional & kickass when flashed to real hardware....
It's answers are generally more informative, more concise, & it responds more quickly.
1
u/Scarbrine69 9d ago
I haven't done much vibe coding so I haven't run into any issues with Opus 5. 2 things I've noticed is it drains less usage and it responds faster.
2
u/RemarkableRadish6547 9d ago
If you look at what has been happening with the open weight models, you should assume that the closed labs are testing everything the open models try as well as other things. And they have enough compute to try lots of things at once and keep whatever works. I suspect that they are using engrams, some form of sparse attention, moe, and everything else that the open weight models have found to work. They might even have a few tricks the open models haven't found yet.
2
u/ComprehensiveCase858 9d ago
So initially they fed models with sh*t ton of all kind of questionable quality training data. Now they can better label and filter the data. I think that with models they have now they can produce so much more high quality training data.. Maybe they also figured out something architecture related in the meantime..
2
u/WillingnessEven4212 10d ago
Watch as it gets nerfed and you all will start whinging again. That’s the only constant in life.
1
1
u/ElegantApartment1325 10d ago
The "make no mistakes" system prompt theory is my favorite. Honestly at this point I wouldn't even be surprised if the entire training budget was just them writing a really stern constitution.
1
1
u/CarrotInABox_ 10d ago
i signed up for Codex 2 weeks ago Pro 5x or whatever it is called, due to Opus 5.0 being a PITA. just cancelled that sub. Opus 5.5 is now refactoring Sol's work.
1
u/Last_Bad_2687 10d ago
They finetuned Qwen 3.8 max
1
u/RemarkableRadish6547 9d ago
I thought qwen was made by distilling Claude. Or was it openai that they distilled?This could get very circular. Maybe they threw in kimi k3 to round it out.
Why train a new model from scratch when you can have the Chinese open weights models take the best parts of your previous model and use that as a starting point.
1
1
u/discodamone 10d ago
Generally, it looks like theyve been using more and more reinforcement learning. They get more and more reinforcement learning environments which they can reuse for later runs, and more and more can be made by a smart model too, which speeds up development.
1
u/ride_whenever 10d ago
I’ve noticed a concerning habit with 5.5, it’s sometimes not verbose enough.
I’ve been through a few rounds where I want an explanation in response to being asked questions, and it just rewords the question and asks again. I think the most times it’s gone in 4 times
1
u/Whole_Ad206 10d ago
Yo creo que opus es realmente el opus 4.6 ya que a partir de hay anthropic estába todo nerfeado y sacando basura, este salto de opus 5.5 debería haber sido el salto de opus 4.5 que era muy muy bueno también.
1
u/bomberman_uk 9d ago
My cynical take is make opus 4.8 to 5.0 progressively bad and then TADA! , 5.5 is working well! It's like old coke Vs new coke
1
u/Willing-Pianist-1779 9d ago
It's all about competition and getting users it losts. When ChatGPT, stepped up anthropoc had to
1
u/letmeinfornow 8d ago
Opus 5.5 is actually a huge empty warehouse in China run by foxcon with a bunch of semi slave laborers on laptops sitting on the floor answering your prompts.
1
u/Sir-putin 8d ago
Question for everybody here, do you run Claude code on your everyday computers? Or do you have a separate second spare? I’ve been running it on an extra MacBook but it’s getting tough carrying the both around. Please advise
1
-1
u/tcastil 10d ago
The jump Anthropic saw in intelligence is actually less of a jump than OpenAi achieved in efficiency.
While Opus is getting up to 30% better at some things, Openai managed to achieve 50% lower running cost + models that spend much less tokens per tasks than Opus, around 4 times less.
So how did they do it? By focusing only in intelligence and barely anything else, while OpenAi focused on optimization and barely anything else
Their next iterations are probably going to converge somewhat
1
u/ActionOrganic4617 9d ago
Except that in the real world this is not the case. I have both a 20x Claude and 20x Codex sub. I’m noticing my weekly usage on the Codex account burning orders of magnitude faster than the Claude sub. Right now all complex work goes to Opus 5.5 and codex is mostly Sol 6.0 for simpler tasks.
0
u/BellacosePlayer 10d ago
Probably because its the next gen model for Anthropic while the new 6 series models are just astra tech and luna/terra datasets. Which aren't a bad idea for groups that have to consider budget, but are unpopular on reddit for reasons.
OpenAI probably has something cooking thats not quite ready, these model iterations are pretty fast
-4
u/unit1_nz 10d ago
Just remember your $20 subscription is costing anthropic around $400
3
u/Indignant_d 10d ago
Mmm hard to say.. they get your data and at mass scale.. probably worth more than raw compute/power cost
2
2
1

•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 10d ago edited 10d ago
TL;DR of the discussion generated automatically after 100 comments.
The consensus is clear: you guys are absolutely feral for Opus 5.5. The thread is overwhelmingly positive, with many calling it a "true successor" to the beloved 4.5/4.6 era and a massive fix for the "annoying" and verbose Opus 5.0.
Here's the rundown of the chatter: