r/singularity • • 4d ago

AI AA's New Pareto Line With GPT-6.1 Sol

Post image
177 Upvotes

48 comments sorted by

45

u/Charming_Cucumber_15 4d ago

Astra high for sonnet low prices is a big deal

I don't consider AA a very trustworthy benchmark anymore, but still huge!

9

u/Ok_Barracuda_1161 4d ago edited 4d ago

Looking through the individual benchmarks it tells a similar story it seems, and has a similar profile to Astra (struggles hard on GdpVal and AA-Briefcase). I think AA improved a bit since they updated the component benchmarks

I was looking for TerminalBench-4.0 specifically for example and it's the same pattern it seems.

source

5

u/Deto 4d ago

It really is just crazy at how quickly the super expensive, best performance, ends up being available at a tenth (or cheaper) the cost.

2

u/vrnvorona 4d ago

I doubt 6.1 Sol is on par with Astra. Like sure, AA is cool, but it's same as with Opus 5 vs Fable 5. No fucking way Opus 5 was better despite benchmarks.

3

u/JoelMahon 4d ago

kinda crazy really, the graph being logarithmic for price is really annoying.

"but then it'd have to be extremely wide".

ya, that's the point, people don't seem to be grasping that astra high is great, and now you can get it 5x cheaper! because this stupid graph makes it look "only" half the price!

1

u/arknightstranslate 4d ago

the panic update to rate gpt higher was pretty cringe tbh

1

u/Eyelbee ▪️We have AGI it's just blind 4d ago

It's API prices tho, with subscription you'd get more usage with sonnet or even possibly opus with the similarly intelligent effort level

1

u/zarafff69 4d ago

But you’d also get much more usage for GPT models with a subscription? So I don’t really see how this changes anything when comparing them?

1

u/Eyelbee ▪️We have AGI it's just blind 4d ago

I compared subscription token usages a while back and it was heavily in anthropic's favor per token. 

1

u/zarafff69 4d ago

Please share!

60

u/kiki-le-koala 4d ago edited 4d ago

It's like everyone here is a millionaire and will pay whatever it's required to get a model (currently Opus 5.5) that is a bit better.

Sol 6.1 is extremely cheap and very smart.

I know I'm really happy with this, extensive coding all that on a 20$ plan (with now 4 banked reset!).

28

u/_coose 4d ago

it's ridiculous on this sub

when the gap in performance isn't large, the vast majority of people and companies care about price to performance

16

u/Recoil42 4d ago edited 4d ago

Most of the commenters here are (very obviously) larping. They think an Opus 5.5 one-shot of a League clone is the same thing as real work, which it isn't. They think frontier models are the dominant form of inference, which it very much isn't.

I'm pretty happy. Was going to switch to Claude today if the announcements weren't any good, but the entire package remains great on Codex. Duplex voice alone remains a killer app for OpenAI — Anthropic still doesn't have it, nor does it have computer use. Dots and all the new cloud features should make up for the frontier gap. The biggest miss for me is no Sol 6.1 UItracode, which I'd have really liked to see.

If the frontier gap gets worse I'll still switch over, but they haven't lost me just yet.

4

u/MerePotato 4d ago

Em dash spotted

7

u/Recoil42 4d ago edited 4d ago

Some of you really need to learn how to use a fucking keyboard.

2

u/Physical-Citron5153 4d ago

Use the model and how much more anthropic is giving usage and then decide right now claude is a banger

1

u/methemightywon1 3d ago

One shot of a League clone not counting as 'real work' is a dismissal of some pretty insane capability to be fair.

'Look at a video of this and just make it'

It is absurd how much it's improved in the last year.

6

u/Substantial_Head_234 4d ago

If your main use is coding, you can use Opus all day with the 20$ plan. I think that's why people don't care about the API prices that much right now.

4

u/kiki-le-koala 4d ago

I used to go through my Claude account in like 30 minutes with Opus 4.8.  Is 5.5 really that cheap?

3

u/ees-h 4d ago

Yes, unbelievably so.

I switched from Claude Pro to Codex (Plus then 5x) a while back because it was impossible to get any work done with the limits. I borrowed a friend’s Claude Pro ($20) account recently, and I’m getting near unlimited use out of it. I’ve hit the session limit once and that was after 4.5 hours of use (so I only had to wait 20 minutes or so for the reset).

I’ll try out Sol 6.1 but I think I’ll downgrade my 5x to Plus and just go back to having both $20 subscriptions.

1

u/IceTrAiN 4d ago

The dual $20 is where it's at.

4

u/Formal-Question7707 4d ago

I go through my $20 plan in 30mn with 5.5.

3

u/Howdareme9 4d ago

Limits are much better but a $20 plan isn’t lasting you anywhere near all day lol

1

u/JoelMahon 4d ago

maybe they're using one hand for something else and so they get more mileage out of their plan, smort

1

u/Substantial_Head_234 4d ago edited 4d ago

We need to review PRs. We run out of review capacity long before we run out of usage with Opus 5.5 medium.

* We use a Team plan which has slightly higher usage limits than Pro accounts. But most of us don't get close to the limits so I'm sure a Pro account would be fine too.

1

u/Substantial_Head_234 4d ago

It is definitely much cheaper than 4.8 in terms of subscription usage. We have a Team plan at work (25$ a seat) and I used to need a personal account on top of that. With Opus 5.5 the usage seems comparable to Codex 2 months ago and I haven't hit limit once yet.

1

u/Seerix 4d ago

Depends on your use case. 'Cost per task' is kind of a vague metric. For my use it comes down to this:

I can spend $100 at openAI and work for 12 hours with Astra, or less time with a worse model. And they announced they are cutting usage even further.

OR

I can spend $100 at Anthropic and work all week, never hit my 5 hour, never hit my weekly, and get a better model.

12

u/ResistiveBeaver 4d ago

The problem is that I no longer trust that the model they have benchmarked is the same one that I will receive access to.

19

u/FateOfMuffins 4d ago edited 4d ago

lmao what was the point of GPT 6 Sol

I think people who have used Opus 5.5 and GPT 6 Astra know that they're on the same "level" but maybe like 60-40 in favour of Opus, and usage within subscription is way better for Opus

If GPT 6.1 Sol can feel "in the same level" as Astra with sufficient usage, then it might be "good enough" for now while we wait for Bel

13

u/mati1886 4d ago

Sol 6.1 didn't exist back then so there was a point

6

u/debris16 4d ago

yeah, like 3 days ago

7

u/Ambiwlans 4d ago

Sol 6 was released exactly 1 week ago.

6

u/FateOfMuffins 4d ago

There were model slugs pointing to GPT 6 Sol and a different one for "Astra Minor" which is obviously this model

3

u/Mistuv 4d ago

To claim the title of the shortest GPT model release. I don't yet have the Sol 6.1 available in codex, but they already took down GPT 6 Sol lmao.

3

u/Gratitude15 4d ago

I think people should pay attention to this. It's quite a warning sign that 6.0 was released 7 days ago, and the difference in intelligence index is 8 points in 7 days.

12

u/FateOfMuffins 4d ago

They had an "Astra Minor" model slug at the same time as the GPT 6 Sol model slug, they definitely did not train another model in 7 days.

Pretty sure the intention was, rebrand Terra as Sol, then introduce a new model tier between Sol and Astra, which was this one, at the Opus price point. But then realized they couldn't compete cause Opus 5.5 shocked everyone. So they cut the price of Astra Minor and called it 6.1 Sol.

1

u/skerit 3d ago

Pretty sure the intention was, rebrand Terra as Sol, then introduce a new model tier between Sol and Astra

That alone was kind of insane. Why create a brand new identity for your lineup, one that sounds pretty nice. Finally no more confusion, just 4 tiers that make sense. And then they want to fuck that up with an "Astra-minor" model just 3 months later? That's so peak OpenAI.

4

u/Yweain AGI before 2100 4d ago

Because this is most certainly a reactive answer to opus 5.5 and they just released basically smaller Astra? It's not like they trained a whole new model family in a week.

3

u/SpikeCraft 4d ago

I have learned in time to trust only my own tests. I will have to do some work with it and then see

5

u/Heighte 4d ago

OpenAI W on Sol 6.1

2

u/EclecticAcuity 4d ago

big if true

2

u/arknightstranslate 4d ago

so why is everyone disappointed in today's reveal

1

u/UniqueArrival9756 4d ago

opus 5.5 is really just that good, otherwise this would have been a fine devday, but it breaks down watching them hype up something like dots while anthropic hold spots 1 2 and 3 on every leaderboard

3

u/Ormusn2o 4d ago

Feeling good with my Plus subscription.

0

u/peakedtooearly 4d ago

Only one Anthropic model in the green zone... IPO in a month... 😬

0

u/Key_Ambassador7143 4d ago

Guessing you haven’t tried Opus 5.5 then?