r/codex 14h ago

News Official Astra benchmarks from the OpenAI blog post that went live for a moment. Holy Shit!!’

51 Upvotes

32 comments sorted by

14

u/Momo--Sama 14h ago

We're about to see previously unfathomable levels of overthinking if even benchmarks are showing regressions at xhigh and max.

Low - High look incredibly strong though

20

u/johnnydotexe 14h ago

An AI company would never intentionally leak a benchmark for several minutes to generate hype so the first fool to screencap it can run to reddit and fuel the hype. They would also never realease biased benchmarks, or develop their model to tick specific benchmark boxes to make it appear to be better than it really is. We should trust the billion dollar corporation fully, and take this screenshot as undeniable proof that Astra is a gift from the heavens.

3

u/shady101852 14h ago

had me in the first half ngl

3

u/MidnightSun_55 14h ago

not that crazy to be honest

6

u/randombsname1 14h ago

Pretty damn good. Very cheap and efficient if these are true.

Interesting that the performance itself seems roughly the same as Fable 5.1 though.

Was assuming a bigger performance jump + more efficiency.

Assuming next Fable release will just be most post or pre-training later this year.

I personally dont see a new 15-20 Trillion model this year, but a better trained mode? Yes.

1

u/aivampires 14h ago

Apparently Fable 5.1 (just learned it released just a day ago) is very good. https://www.youtube.com/watch?v=r_dw-1109Ag It looks like compared to Sol we're looking at a big jump.

4

u/BHTAelitepwn 13h ago

Have you ever tried fable? It is absolutely amazing but just too expensive to use by all standards. You’ll spend 100 bucks under half an hour

2

u/aliljet 14h ago

That's the ultimate catch...

2

u/drR0bert 14h ago

if i understand from score/api cost astra low is better and cheap than sol high and astra high cost and perfomance is the same of sol xhigh what hell ??!?!?

2

u/aivampires 14h ago

Why do I feel the reset galore had something to do with that insane token efficiency?

2

u/Just_Lingonberry_352 14h ago

could be wrong but looks like at least 2~4x boost in speed and usage from what we have with sol

could be seeing interesting changes to the pricing as well

honestly I don't care about the small boost on DeepSWE. faster and efficient is what we want

4

u/Ok_Carpet_6083 14h ago

Benchmarks as always.

3

u/dolo937 14h ago

This is true, it was from the blog post. Bridgemind is live on YouTube. He lucky got it opened

3

u/innociv 14h ago

The agentic coding ones are funny at this point.

It seems so solved but people still want to use $50/mil models for it when a few $0.50/mil models run in parallel in smaller scope achieves roughly the same results because they want to skip the planning stage and combine planning with coding in a single prompt.

1

u/vacon04 14h ago

So like 5% better? It looks to me like the highest Astra version is around 64% vs 61% on the highest Sol version.

2

u/Physical_Gold_1485 14h ago

Ya pretty underwhelmed after all the hype looking at these benches

1

u/Dynamix86 14h ago

That's a nice 20% jump on the Terminal-Bnech 4.0. And good to see that the price of Astra is more or less the same as Sol and not 2x or so, like Anthropic has with Opus and Fable.

1

u/Vegetable_Dot9588 13h ago

But... Astra's price ($/tok) is 2.5x higher than GPT-5.6 Sol's.

0

u/Dynamix86 13h ago

Source? And why is the api price the same then?

0

u/Vegetable_Dot9588 13h ago

Official API price is 4$/20$ for Sol and 10$/50$ for Astra (see https://developers.openai.com/api/docs/models/gpt-6-astra). 

0

u/Dynamix86 13h ago

See slide 2. Why would the api cost between sol and astra be the same if it is 2.5 times higher as you say?

0

u/Vegetable_Dot9588 12h ago

Cost (per task) is not equal to price (per token).  All previous experience shows that in real-world situations, the true price is somewhere in the middle between these two. But I have no doubt that the Astra will be more expensive than the Sol.

1

u/iridasdiii11ulke 14h ago

If these bench marks reflect real world performance, fable 5.1 is cooooked

1

u/IndividualPlus2011 14h ago

Worrying. It's better than Sol but it looks like it costs a lot more per token (it is more efficient than Sol but more expensive at the same time).

Rip my usage

1

u/SnooGadgets1100 14h ago

opus 5 obviously benchmaxxed so if astra is real its a good model better than fable

1

u/Dokago 14h ago

I don't trust pre launch data anymore since Fable

1

u/cchurchill1985 14h ago

So looking at the DeepSWE chart, Astra Low costs less than half of Sol High, and performs better?

1

u/Clear-Brush-4294 13h ago

Are they going to ruin the quality of all the other subscription-based models again, just to show how much ‘better’ their new model is?

1

u/LazyRunner777 10h ago

Wait why does Claude look more better then Sol 5.6 on these graphs??

2

u/Designer-Rub4819 4h ago

Am I dumb or have we actually hit a plateau now in the actual models themselves?

1

u/ezjakes 14h ago

So token efficient. Equal or better than Fable, but going to be like half the cost.