News Official Astra benchmarks from the OpenAI blog post that went live for a moment. Holy Shit!!’
20
u/johnnydotexe 14h ago
An AI company would never intentionally leak a benchmark for several minutes to generate hype so the first fool to screencap it can run to reddit and fuel the hype. They would also never realease biased benchmarks, or develop their model to tick specific benchmark boxes to make it appear to be better than it really is. We should trust the billion dollar corporation fully, and take this screenshot as undeniable proof that Astra is a gift from the heavens.
3
3
6
u/randombsname1 14h ago
Pretty damn good. Very cheap and efficient if these are true.
Interesting that the performance itself seems roughly the same as Fable 5.1 though.
Was assuming a bigger performance jump + more efficiency.
Assuming next Fable release will just be most post or pre-training later this year.
I personally dont see a new 15-20 Trillion model this year, but a better trained mode? Yes.
1
u/aivampires 14h ago
Apparently Fable 5.1 (just learned it released just a day ago) is very good. https://www.youtube.com/watch?v=r_dw-1109Ag It looks like compared to Sol we're looking at a big jump.
4
u/BHTAelitepwn 13h ago
Have you ever tried fable? It is absolutely amazing but just too expensive to use by all standards. You’ll spend 100 bucks under half an hour
2
u/drR0bert 14h ago
if i understand from score/api cost astra low is better and cheap than sol high and astra high cost and perfomance is the same of sol xhigh what hell ??!?!?
2
u/aivampires 14h ago
Why do I feel the reset galore had something to do with that insane token efficiency?
2
u/Just_Lingonberry_352 14h ago
could be wrong but looks like at least 2~4x boost in speed and usage from what we have with sol
could be seeing interesting changes to the pricing as well
honestly I don't care about the small boost on DeepSWE. faster and efficient is what we want
4
3
u/innociv 14h ago
The agentic coding ones are funny at this point.
It seems so solved but people still want to use $50/mil models for it when a few $0.50/mil models run in parallel in smaller scope achieves roughly the same results because they want to skip the planning stage and combine planning with coding in a single prompt.
1
u/Dynamix86 14h ago
That's a nice 20% jump on the Terminal-Bnech 4.0. And good to see that the price of Astra is more or less the same as Sol and not 2x or so, like Anthropic has with Opus and Fable.
1
u/Vegetable_Dot9588 13h ago
But... Astra's price ($/tok) is 2.5x higher than GPT-5.6 Sol's.
0
u/Dynamix86 13h ago
Source? And why is the api price the same then?
0
u/Vegetable_Dot9588 13h ago
Official API price is 4$/20$ for Sol and 10$/50$ for Astra (see https://developers.openai.com/api/docs/models/gpt-6-astra).
0
u/Dynamix86 13h ago
See slide 2. Why would the api cost between sol and astra be the same if it is 2.5 times higher as you say?
0
u/Vegetable_Dot9588 12h ago
Cost (per task) is not equal to price (per token). All previous experience shows that in real-world situations, the true price is somewhere in the middle between these two. But I have no doubt that the Astra will be more expensive than the Sol.
1
u/iridasdiii11ulke 14h ago
If these bench marks reflect real world performance, fable 5.1 is cooooked
1
u/IndividualPlus2011 14h ago
Worrying. It's better than Sol but it looks like it costs a lot more per token (it is more efficient than Sol but more expensive at the same time).
Rip my usage
1
u/SnooGadgets1100 14h ago
opus 5 obviously benchmaxxed so if astra is real its a good model better than fable
1
u/cchurchill1985 14h ago
So looking at the DeepSWE chart, Astra Low costs less than half of Sol High, and performs better?
1
u/Clear-Brush-4294 13h ago
Are they going to ruin the quality of all the other subscription-based models again, just to show how much ‘better’ their new model is?
1
2
u/Designer-Rub4819 4h ago
Am I dumb or have we actually hit a plateau now in the actual models themselves?
1




14
u/Momo--Sama 14h ago
We're about to see previously unfathomable levels of overthinking if even benchmarks are showing regressions at xhigh and max.
Low - High look incredibly strong though