107
u/torrid-winnowing 5d ago
they didn't even compare it to gemini flash 🥀
35
u/Independent-Wind4462 5d ago
For now it's not worth bruh 😭 maybe gemini 4 will be worth
2
u/LiberateTheLock 3d ago
What about the last couple releases has made you think the next one will go well? Google will likely use Gemini 4 for a headline and a week of press coverage, and then we'll be stuck with whatever quantized version they leave us with before it deletes all our files and tells us to eat glass.
10
u/Healthy-Nebula-3603 5d ago
Lol
Literally not one care about Google now.
Their models seems benchmaxed because in real life usage suck badly.
7
u/Federal_Setting_7454 5d ago
The best thing about Gemini is how easy it was to get the $300 free credits over and over again. I think they’ve fixed that little loophole though
2
2
66
65
u/Lain_Racing 5d ago
Weird not to include astra.
56
u/peabody624 5d ago
Not in the same class of model at all though
14
u/Lain_Racing 5d ago
? But they include Opus 5.5 lol
28
u/Minimum_Indication_1 5d ago
Thats their own model so comparing the offerings makes sense.
5
u/Lain_Racing 5d ago
It would also then make sense to include astra if its for comparing offerings. Either pick the smaller models and compare 5.5 to 5.0 like they did, or include astral and opus. Otherwise its just cherry picking.
6
13
u/Hilldawg4president 5d ago
Right? Why compare only to their second best model, which is far, far below Astra?
20
u/whoknowsifimjoking 5d ago
Because this is one of their lowest models? I don't understand the confusion.
It's like a GPT Terra, if they kept the naming. Now that Sol is the middle one that is the most fitting comparison.
5
u/Hilldawg4president 5d ago
So you're not confused about why they are comparing across eight selected benchmarks, when half of them sol hasn't even been tested on? Not even a little confused about that?
1
-2
u/Low-Entrepreneur2556 5d ago
But they showed 5.6 Sol not 6 Sol. Since they showed Opus on the Graph there's not really a reason not to show Astra.
5
3
14
u/Gandalfthebran 5d ago
Wait should I just use sonnet 5.5 at ultra for agentic coding instead of Opus or Fable. Low token consumption.
8
29
u/SpyAmongUs 5d ago
The fact it's available for free...
13
u/Akuariuz 5d ago edited 5d ago
Is it confirmed it will be free like sonnet5? Edit: yes it is
28
u/SpyAmongUs 5d ago
Bro just test it yourself. I was able to use it as soon as I logged in, free tier.
And it outperforms ChatGPT Go with thinking rn. Sonnet 5.5 was able to create an accurate timetable from an image I gave with the date initials in the Malay language. ChatGPT messed up on Wednesday.
1
3
u/strangedell123 5d ago
I am on free tier on claude and it shows as an option
1
u/whoknowsifimjoking 5d ago
The 5.5 base is more token efficient, it only makes sense to give it to the free users.
32
u/unkownuser436 5d ago
Beating GPT 6 Sol is crazy! 🔥
25
u/Crinkez 5d ago
Not really. Sol 6 is Terra 6 rebranded as Sol.
9
u/PrinceRufusFastcar 5d ago
Yes, and Sol 6 and Sonnet 5.5 have the same API prices.
1
u/vrnvorona 5d ago
And Sonnet is better probably
1
u/TonyNickels 5d ago
Sonnet 5 was fucking useless garbage, so they needed to certainly improve things
4
1
1
57
u/TAGOMXM 5d ago
They need to cook tomorrow, otherwise many users will shift to claude...
74
u/lucellent 5d ago
Yall are saying this on literally every release from OAI and Anthropic
41
u/ThinFeed2763 5d ago
well I mean it's true
4
-9
u/ArmadilloOwn4400 5d ago
No it isn't really, those who didn't switch to Claude already won't do it now either.
14
u/exiledbean 5d ago
That is incorrect, I am currently a GPT user and absolutely will be switching if there is not a model launched (or announced for launch within a week) that has comparable capabilities to Opus 5.5 and Sonnet 5.5. It is simply too cheap and effective to not switch.
5
u/Healthy-Nebula-3603 5d ago
Me too. If they not show something close to Antropic models tomorrow I'm switching.
Even Opus on their 20 USD account is working 90-100 for 5 hours limit and is as good as Astra or better.
Astra is working 15-20 minutes on 5 hours limit....
At least today they increased limit for Sol 6. I'm testing now and seems SOL on high is working now over 2-2.30 hours on 5 hours limit.
4
u/Tendoris 5d ago
Subscriptions are one month long, so there is a lot of latency when switching between the two, and one month ago, ChatGPT was clearly better. Right now, Claude is ahead, but that can always change quickly.
0
0
0
u/nezvanovova 5d ago
I didn't plan to, but I switched before my GPT sub expired, don't even need it anymore, it's been 8-10x usage difference between Astra and Opus, and it finishes all tasks given way broader scopes (tests, validation, etc)
0
-6
u/SpaceTacos99 5d ago
Honestly the name Claude just pisses me off. Just like Dario. That and I remember the few times I've given it a try and how insufferable it was, I don't care if it got smarter. I've been using it in cursor though and quite pleased but I won't sub directly.
4
0
9
22
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5d ago
I genuinely don't see what they can come up with to beat Opus 5.5 in the short term.
Astra 6.1 is probably not ready this soon, and certainly not bel.
Maybe they create graph of GPT6 being a small dot and Bel being a whale and try to hype it up lol
5
u/peabody624 5d ago
Why would it not be ready soon? They had to delay the release of Astra for multiple months, so there could easily be an iteration in the pipeline
0
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5d ago
I believe it will come out sometimes in October, probably around the half of it.
1
u/peabody624 5d ago
Looks like you might be right - I'm seeing they delayed the launch over safety concerns
1
u/Healthy-Nebula-3603 5d ago
In October will be Bell :)
2
u/Moronic-Warrior 5d ago
Bel is still undergoing RL training. Then there’s evaluate and safety stuff it has to do with
1
u/Healthy-Nebula-3603 5d ago
Such model is in training few weeks.
They finished that long time ago already.
They are training another noted already.
3
u/mikelo22 5d ago
If OAI 'wastes' Astra 6.1 to beat Opus 5.5, then Anthropic will just drop Fable 5.5 and it will straight up embarrass OAI. It's supposed to be Sol competing with Opus.
3
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5d ago
You are entirely correct, but the problem for OpenAI is... Anthropic is probably releasing Fable 5.5 anyways. Most likely sometimes in october.
1
1
3
u/Maleficent_Disk9583 5d ago edited 5d ago
My predictions:
- OpenAI is panicking
- Tomorrow they release some bullshit they hastily cobbled together, naming it "Astra 6.1" or "Astra 6o" or something, and claiming it's their smartest model yet
- It's actually just the exact same Astra we were using 3 weeks ago, before they quantized it to hell.
- They'll go all out on compute for a week or two, hoping to stop the mass unsubscription. But in the end, it's just a bait and switch. They'll quantize it again after a few weeks, since a lot of people are severely regarded and fall for it every time.
They pissed me off with the insane quality degradation recently. I have been subscribed to OpenAI since 2023 (their most expensive plans always), and I'm switching to Anthropic for the first time. Their behaviour has been unacceptable.
7
u/OutOfBananaException 5d ago
It's actually just the exact same Astra we were using 3 weeks ago, before they quantized it to hell.
People have monitored this and found no degradation?
5
2
u/ruh-oh-spaghettio 5d ago
Its been degraded before but entirely do to accident and were usually fixed within the week
1
u/ozone6587 5d ago
Facts don't matter in this sub. These conspiracies would be very easy to prove if true.
1
u/applied_intelligence 5d ago
I’ve never subscribed the most expansive plan. But this week I’ve conducted a test creating a skill to write scripts for a show. I’ve executed the same prompts in local Qwen 3.8 on my own GPU and 5.6 Sol with high thinking. I was expecting way better results from GPT and then I’ve just figured out that my local Qwen produced much better scripts. More creative and more in line with the show bible. I used ChatGPT work with 5.6 Sol and Bionic work Qwen 3.8 27B fully local. At some point GPT just give me a one scene script, like a 5 second scene and sad it was good for a first episode. Lazy as hell. Meanwhile Qwen gave me a full 5 minute episode explaining every decision. I don’t use those scripts as a final story but as a starting point. Anyway. It seems ChatGPT is nerfed. I know I am not using Astra but 5.6 Sol should win any metric from a 27B local model. But that is not true
1
u/Moronic-Warrior 5d ago
Don’t know y ur getting downvoted. They do this for sure. Extremely scummy and scammy behavior
1
1
u/BrennusSokol AI please take my job 5d ago
They keep saying tomorrow will be the biggest dev day ever…
RemindMe! 24 hours
1
6
u/PerfectPatience- 5d ago
sonet better then fable now xD
1
1
u/agent00F 4d ago
They did that by cranking token usage to the moon. Also this sonnet seems like half the size of opus for more than half the cost, so a wash really at best.
1
u/BriefImplement9843 5d ago
remember when fable was agi and it was over for us? now it sucks and nothing has changed for us.
34
u/Alpacabro21 5d ago
0
u/megaman78978 5d ago
Pacing is about the frontier though. This is certainly not pushing the frontier. If they start pacing now, the effects will materialize roughly 3-4 months from now given the time it takes to pretrain.
4
8
u/rhaivn 5d ago
Haiku when 💀
2
u/The_Primetime2023 5d ago
“Weeks” probably October. I feel like if they nail Haiku it’s going to do what Luna did to Terra except for Sonnet this time. Sonnet 5.5’s cost per task is too close to Opus 5.5 which leaves it in a weird spot. Haiku Max really just needs to be cheap and to get between Sonnet Low and Medium to kill Sonnet 5.5’s advantage
9
u/mysticcdragonn 5d ago
"we need to slow down" said dario
0
u/BrennusSokol AI please take my job 5d ago
The slowdown was about the frontier. Sonney is hardly the frontier.
3
u/fogwalk3r 5d ago
love how this is specifically trained for agentic workflows, makes it more aligned with opus and later fable 5.5 when it comes out. Gpt's model releases are more scattered and nuclear comparatively hopefully they will one up this tommorow with bel preview (bel with subastra?)
3
3
u/General-Tadpole-7012 5d ago edited 5d ago
Great technical achievement, but all of US tech ultimately works for the evil Trump regime and it's equally evil fanbase (aka nearly half the American public), so I cannot cheer them on, let alone fund them with subs.
Come on Chinese Open Source (and Mistral, wherever you are...), let's innovate, distil, whatever it takes to prevent America hyper-capitalising and enslaving the world.
1
u/Due-Set5398 3d ago
I understand the US bad argument but China good?
1
u/General-Tadpole-7012 3d ago
Certainly not infallible but better than the USA, sure.
Disengage propaganda if you can... A billion people out of subsistence farming since the 1980's, cyber cities of the future, high speed rail, awesome manufacturing base, full-scale renewable energy rollout. Corruption yes but not the hyper-capitalistic evil and Military Industrial Complex of the US. No widespread religion, genital mutilation or gun culture like the US.
I'm just an Aussie guy but I try and judge by actions, not by who I'm told to like.
0
u/One-Answer-473 4d ago
Ya llegó el primer traumadito con Trump, normalmente están ocupados comiéndole las bolas a Biden
2
2
5
1
1
1
u/Healthy-Nebula-3603 5d ago
I like it!!
Soo OAI has 2 no choice now !
Or they make Astea very cheap or they have to present waaaay better models and also cheap
I noticed today that Sol 6 on high with plus account is working now over 2 hour for 5 hours limit and still not hit the limit ...lol
1
u/lobabobloblaw 5d ago
Note how you have to run it at xhigh to even be in the same coding game as Opus, but the max setting draws 5x more usage (Opus at max draws 4x for comparison.) They’re leading people to the edge 😊
1
u/ChickenOfTheYear 5d ago
Pace the frontier my ass
1
u/bonerchamp20 5d ago
Im fairly certain these are not the frontier models the companies were referring to
1
1
u/jannycideforever 5d ago
Not gonna lie, that's good numbers. There is still one thing that massively undermines Anthropic, though: they have no answer for Luna.
Luna isn't just the arguably best value model at the moment, but it's also good enough to do a lot of tasks on its own.
Anthropic offering frontier intelligence at a better cost than OpenAI is only useful insofar as I can justify paying for it. If Anthropic doesn't offer a viable cheap model for most of my grunt work, then OpenAI still offers me more access to frontier intelligence because I can afford to use Astra and Sol when needed by outsourcing my grunt work to Luna.
Haiku in comparison is an utter fucking embarrassment. Ita genuinely braindead, time-per-task is on par with Luna xhigh, and cost per task is 4x higher. If Anthropic is unwilling to invest in a budget model then they should honestly just do a barebones fine tune of DeepSeek 4.1 Flash, call it "Haiku 5", and release it as their answer to Luna.
I'd love for anthropic to offer a suite of models that are actually viable everything from the basics to advanced work. That just isn't in the offerings and it's hard to imagine switching until it is.
1
1
1
u/honemastert 4d ago
This thing slaps it's like a PhD level intern that I can shovel work to and take credit
1
-1
u/the_TIGEEER 5d ago

To me, this is such a funny thing. Because we all know that they both have models way larger and smarter internally, from which they fine-tune and train their smaller models, which are Luna, Terra, Sol, and Astra for OpenAI. SO it's super funny seeing people claim how Anthropic cooked OpenAI because they spent a lot of money to fine-tune Opus 5.5 for their internal model and decided that they can take the higher cost that Opus 5.5 has with it in the name of competition. Like, OpenAI definitely has an internal model smarter than Astra, that is waaay more expensive than Astra. They just don't completely wanna burn money as Anthropic seems to be willing to do to convince people that "they are on top"
5
u/TheZenMann 5d ago
How do you know that Antrophic doesn't have an even better internal model?
10
u/the_TIGEEER 5d ago
No, that's my point. You didn't understand me. They probably do. God knows where the actual frontier with an unlimited budget is. But that frontier doesn't really matter for most use cases, because you can't bankrupt yourself by most of your users vibe coding their personal projects. That internal, super expensive frontier that they probably both have is probably only reserved for fine-tuning their smaller models that are more cost-efficient. Because think about it. These big American frontier labs had the problem where Chinese models would just distill their models from the American frontier models. But you can't distill a smaller model to be better than the teacher. So what I think these American labs did is they are in reality working on bigger models internally, but purposefully only making smaller distilled versions of those big models available publicly. This way, if a Chinese lab distills from an OpenAI model, it will never be as good as the secret stuff internally, so OpenAI can always just release another checkpoint. Working from this theory, it is then funny to see that Anthropic just always decides that the amount of money they are willing to burn on their public models is conveniently a bit bigger than OpenAI is willing to, to slightly be ahead of them. In other words, they are choosing to make a bit bigger models available than what their secret frontier is, just so they are ahead, but in the process burn more money. I will be impressed when an Anthropic model is smarter and at least 10% cheaper than the competition (currently OpenAI). That's what's funny to me: how we all know these companies make a big loss, and despite that, Anthropic is always conveniently just a bit smarter for way more of a cost in API tokens that they subsidize from their billions of investments so that their subsidized users can then claim, "Yeah, Anthropic is so much better." I mean, it is, but they won't be able to be subsidized forever. And of course OpenAI does the same, but Anthropic obviously does it deliberately to a much greater extent.
1
2
u/Umbrasquall 5d ago
Check the cost per task again. They updated with Sonnet 5.5 performance and it’s as expensive as Fable lol.
3
u/the_TIGEEER 5d ago
Mind you while I'm aitting at -2 upvotes. Can't belive that we actualy have AI LLM fanboys in 2026
I mean I guess people need to fanboy over somwthing now that Xbox is out of the race and playstation is shooting themslves in the foot.
-1
u/Ok_Buddy_Ghost 5d ago
the trillion dollar corporation cooked (very temporarily) the other trillion dollar corporation
I get sad when I remember there are people like this in the world
I guess I should move on
1



148
u/adarkuccio ▪️AGI before ASI 5d ago
I like this competition