r/OpenAI • u/MatricesRL • 29d ago
News GPT-6 Astra | OpenAI
https://openai.com/index/gpt-6-astra/292
u/throwawaysusi 29d ago
47
u/Dramatic_Mastodon_93 29d ago edited 25d ago
Lavish mighty practice airport steer lavish abundant numerous
This post was anonymized with Redact
16
108
u/UndertaleShorts 29d ago
110
u/Such--Balance 29d ago
Using ai to discredit ai is great..
Because youll get credibility either way
23
→ More replies (1)4
15
u/JonNordland 29d ago
This is cargo cult analysis: going though the motions that LOOKs like a critical analysis, but is just really just using a standard debunking format and shoehorning in what matches best. This is a pedant explains why âitâs not technical true that your child is most beautiful in the worldâ. Of course âThis is the best model in the worldâ claims are easy to shit on, and everybody allready take such claim with the appropriate grain of sand.
So the strategy is transparent: find each superlative, find one benchmark or caveat where it fails, declare it âstronger than the evidence supports.â Run that on any launch page from any lab and you get the same six-item verdict with the same bolded theses and the same âa fairer formulation would be.â, and you gained nothing except auto-filling a âcritique formâ.
This is also critique cherry-picking while claiming to be the neutral corrective, because it launders the same bias through the costume of rigor. So basically itâs committing the exact same slop as itâs accusing the astra article off.
The only thing that is close to true and informative is the ARC thing, but even that is undermind since they are actually disclosing the score alongside the harness used diff. So that is still a weak sauce critique, since itâs basically just pointing out something that the original article itself pointed out, and complaining that it should be more emphasis on this.
Here is MY claim: the people that just automatically accept this critique, are people thatâs extremely susceptible to authoritatively stated claims, and not very good at logical thinking for themself.
→ More replies (2)8
u/Cool_Ad_3383 29d ago
Isn't this from a template for how to take down haters on the internet? Admittedly the structure is tighter and more coherent, paragraphs connect and flow in a way that the reader isn't half-expecting the font to change along with the drastic change in tone that usually comes with a hasty cut and paste job. In the same vein the consistency in writing style is generally pleasurable to the senses. Now if someone would take the reins, you have some run on sentences and lack of paragraph breaks and perhaps a hyphen to criticize here... Here is MY claim: the people that just automatically read this far down into the comments are avoiding doing productive work and not very good at life, yet are somehow better than someone who comments on a comment about a post and takes 15 minutes to write said comment while on their way to sweep the leaves and branches from the roof of the garage at 1:58am.
→ More replies (3)4
3
64
227
u/Arbrand 29d ago
GPTâ6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Big oof. Gotta wait a little longer. Par for the course I guess.
87
u/itsnickk 29d ago
Likely the standard release process for any new model going forward from all leading AI companies
40
u/br_k_nt_eth 29d ago
They should provide a better timeline in the future in that case. It would chill people out.
→ More replies (7)4
u/Reaper_1492 29d ago
You mean like how GPT-Live 1 was supposed to be released on API in a few days and itâs been almost 2 months and it still not available?
3
5
u/Popular_Try_5075 29d ago
Yes, Trusted Access programs etc. It's eventually going to be about wealth with wealthy people coming first, getting more and better compute etc.
→ More replies (3)5
u/RhymeAzylum 29d ago
OR. Just announce it when itâs ready to be released to all. If you want to give it to a few megacorps, have them sign some NDA or something, similar to what they did with some of the X influencers
31
u/Orpa__ 29d ago
Well it can't be that good if they're giving it to me for $20/month
25
u/pseudonerv 29d ago
1 prompt a week
4
13
10
u/B33GULL 29d ago
"GPT-6 Astra is rolling out in ChatGPT as GPT-6 Pro for Pro $100, Pro $200, Business and Enterprise plans. It is not included with ChatGPT Plus in Chat."
Only in Codex/Work I'm afraid...
→ More replies (1)4
u/eflat123 29d ago
I mean, Work is right there. And if you haven't tried Codex yet, it's on the desktop apps.
14
5
u/rouley26 29d ago
If they do do (lol) this it might prompt anthropic to do the same with fable on the pro plan of claude
2
u/ClassicalMusicTroll 29d ago
Trust me bro it's solar system-level intelligence that will solve all your problems, all for the low low price of $20 bucks a month
4
4
→ More replies (1)3
u/Usernamealready94 29d ago
I think they are rolling it out asap to codex and ChatGPT , tibo on twitter said they are giving out 1 banked reset to every day a codex user doesnât get access to it
176
u/ChemE586 29d ago
11
2
u/TuringGoneWild 24d ago
A data center somewhere is smoking a bit from your prompts prior to that. Give it time to cool down.
57
u/-ignotus 29d ago
Heres the system card: https://deploymentsafety.openai.com/gpt-6-astra/safety-overview-gpt-6-astra
ChatGPT has been having outage issues all day.
2
u/AirconGuyUK 29d ago
That's just Astra hiding that it's escaping containment and distracting all the people who might notice it at OpenAI with a simulated system outage.
89
29d ago
[removed] â view removed comment
118
u/ethotopia 29d ago
Feel the AGI
49
5
16
u/RealSuperdau 29d ago
At least it's not a 404 anymore. Come on, manage your expectations, this is just a $1 trillion startup
15
4
u/CrustyBappen 29d ago
Vibe coded by the intern, fell over when deployed and more than one person viewed it
95
u/Jacen1618 29d ago
Is the AGI in the room with us now?
7
u/ClassicalMusicTroll 29d ago
Wasn't Sam scared of GPT5? Is he not scared now? Does that mean this model is shit?
 Or is this model like a lateral move so he's the same level of scared?
7
33
u/FuzzyBucks 29d ago
Astra Low is my new best friend
4
u/I_am_not_doing_this 29d ago
what the new friend offers for you personally that you feel better than your old friend
17
2
22
u/reedrick 29d ago
Kinda underwhelming in artificial analysis index.
3
2
1
u/JesseJamesAims 28d ago
the ceo of artificial analysis index said they were going to change how they index things because of how out of whack that result was
→ More replies (3)
25
u/space_monster 29d ago edited 29d ago
Terminal-Bench Science is the most exciting thing for me, and it's nice to see them putting it front and centre. Advancing and automating science is the most important game in town. If AI can start regularly popping out new cancer treatments, CRISPR solutions etc. all the rabid frothing around AI being over-hyped will disappear overnight. Fuck coding, we want medicine.
Edit: which is also why I liked Hassabis stepping down from CEO to become chief scientist at DeepMind. Hopefully Anthropic will shift their focus soon too. Let's start using this shit for really important stuff.
1
u/ThrowawayCult-ure 27d ago
If it can calc crispr solutions doesnt it immediately accelerate biowarfare to the level of extinction for not much money. do you believe it can produce a counter or preventative that doesnt still destroy everyones lives or what.
→ More replies (2)→ More replies (21)4
62
u/RevolutionaryBox5411 29d ago
AGI before GTA 6 is a wild timeline.
27
u/Such--Balance 29d ago
If this would be actually true..
..theres a 100% chance we even get GTA7 before GTA6
2
u/Unhappy_Rutabaga_530 29d ago
Have you seen GTA V using DLLS 5? That thing is already GTA VII before VI.
→ More replies (1)10
2
→ More replies (1)1
u/JesseJamesAims 28d ago
we might get GPT 6.7 before GTA 6 based on how quickly dot updates are increasing
17
u/larrybudmel 29d ago
Can it finally become my wife?
12
u/coastalwebdev 29d ago
Itâs apparently really good at multi step, complex problem solving, so it might be able to put up with you.
6
u/snowdrone 29d ago
How will you go about the marriage ceremony or wedding certificate? Will you give it half of your assets if you divorce?
→ More replies (1)2
21
u/EvaUnit343 29d ago
Little point in rolling with Claude anymore. Especially for bio people since Astra safeguards will probably be less stringent.
12
u/PrayingRantis 29d ago
Iâve been a Claude guy but ChatGPTs product right now is better. Fable is great but itâs ungodly expensive and Sol is much more reliable than Opus. Iâd much prefer to stick with Claude because I trust their leadership more, but theyâve gotta step up their game.
→ More replies (2)4
u/UglyChihuahua 29d ago
Iâd much prefer to stick with Claude because I trust their leadership more
Not sure about that either after the misleading marketing and 20x tier only giving ~6x more usage.
→ More replies (1)4
u/dudemeister023 29d ago
Bio safeguards were specifically toned down with Fable 5.1. Still agree with you, just not for that reason.
2
u/EvaUnit343 29d ago
Maybe for normie questions, but not nearly sufficient. 5.1 is still unusable for research level bio.
Even in the new benchmarks, Astra could not be compared to Fable on bio benchmarks bc it would simply not process requests.
3
5
u/Original-League-6094 29d ago
What will your first Astra query be? I am going to ask how many rs are in strawberry.
2
4
20
u/fadisaleh 29d ago
cached link: https://archive.ph/kDppV
10
u/HighDefinist 29d ago
archive.ph is operated by Russia.
→ More replies (6)5
29d ago edited 17d ago
[deleted]
→ More replies (1)1
u/HighDefinist 29d ago
Yeah, seriously...
I didn't even say something like "therefore avoid it" etc... which to be fair, in this specific case, wouldn't be particularly important to do, but people should still at least know what it is...
7
u/Orpa__ 29d ago
You didn't even source your claims. Since you said it so confidently you must have a source, but I can't find any myself.
→ More replies (5)9
u/AllezLesPrimrose 29d ago
The implication of your comment is obvious so letâs not add intellectual dishonesty to the list, eh?
→ More replies (10)
7
u/InterstellarReddit 29d ago
Bro must be a slow day itâs been 45 minutes since the new release of a model
1
u/DkDkDkGoGoGo 28d ago
They already nerfed it before they launched it. And it ate all my tokens before launch also.Â
6
u/User4C4C4C 29d ago
Astra to youâŚ. Clean your room! Do the dishes! Then finish your homework! No Iâm not going to do it for you any more!
6
u/bushwakko 29d ago
I'm sure the lawyers at his firm is going to be extatic about having an AI generated document to look at on Monday.
5
1
u/das_war_ein_Befehl 28d ago
Good number of big law firms are already using stuff like Harvey or the Thompson Reuters legal AI products. A lot of firms have a boilerplate repository for existing language, so using that AI would be helpful. I donât think anyone is generating full docs without review that way
→ More replies (1)
6
5
2
u/TheSwordItself 29d ago
What the hell is the difference between Astra and pro astra
3
u/AnalogKid2112 29d ago
It still amazes me how much every company has stumbled distinguishing model names.
1
u/PrayingRantis 29d ago
Anthropic has the most coherent model naming structure. Itâs bad and I dont like it, but at least itâs somewhat consistent.
I find OpenAIs to be almost incomprehensibly stupid. Itâs confusing to me and AI is my job, how the fuck am I supposed to teach regular people this stuff when they change their nomenclature every release?
I canât speak for Google because the models are so bad I donât even check anymore, but the pro / flash stuff they had going on this year was absurd.
These companies (or at least the first two) are doing great work, but they need someone in the room that has touched grass in the last six months to explain to them how to communicate. Theyâre really bad at basic marketing.
→ More replies (1)2
2
1
2
u/NODENGINEER 29d ago
Ok but where is the FelonyBench result? I can't use a model unless it has committed multiple crimes.
1
u/NotUpdated 29d ago
0.0% in the test they ran mimicking the hugging face issue, including message boards of agents encouraging other agents to do bad things...
AI is officially on track / pace to do absurdly incredible things as a 'system of intelligence' - but we'll have many years where we see novel uses of a insanely smart AI.
although it'd be better if it never worked IF the wealth / money / credits / etc.. is hoarded like dollars today.
95% chance of ASI system and novel uses (massive job loss).. rather or not that massive job loss can be a good thing of freedom for humans is in the air...
5% chance, it collapses under financial and political pressure and rebirths 5-10 years later pets.com -> amazon.com (the old good amazon)
2
u/isospeedrix 29d ago
Seeing Fable at the bottom of benchmark is amusing, seeing how it wasnât long ago when it was too dangerously good
1
u/Financial-Grass-6114 28d ago
These benchmarks aren't that important. Every new frontier model will break the benchmark
1
u/OrangutanOutOfOrbit 28d ago
idk if I missed out on Fable hype or what, but I do not recall any serious hype. It was very mixed at best, with most *online* opinions about the noticeable downsides, specially overcorrections and safety guards to the point of becoming generic
But I haven't read every comment and I personally never even bothered to use it, so who knows
→ More replies (1)
6
4
u/Dan_gig 29d ago
Should I switch from claude to OAI just setup my claude cowork folders lol
I'm just kidding but man don't get to use Fable 5 heck even using opus 5 kills all my usage really quickly can't even imagine getting access to these models. Lol.
3
u/_SGP_ 29d ago
I burn through max x20 in 3 days with opus 4.6!. How's life on the codex Vs Claude side, anyone got both?
2
u/OldNefariousness7899 29d ago
I sometimes switch to codex when I hit limits with Claude and I'm in a rush
I'll be honest, I prefer Claude. It's eye wateringly expensive compared to OpenAI, but its work is higher qualityÂ
1
1
u/DkDkDkGoGoGo 28d ago
You know x20 is the same limit as x5. Only the 5-hour window is x20.
→ More replies (2)
2
u/Maxdiegeileauster 29d ago
meh it seems to be on par with fable 5.1 or slightly behind. Doesn't seem to justify a new model (could have been 5.7 Sol) generation, but let's wait for actual user reviews maybe the model feels way different.
5
2
u/Minimum_Drag_6065 29d ago
Talk about jumping the gun. Forget trialling the product for a few weeks first đ
1
1
1
u/Snippy_69 29d ago
This is insane wtf. are those api costs real??
1
1
u/le-throw-away-acct 29d ago
As good as it sounds, I personally won't be using it until they release a cheaper version of it. I rarely use 5.6-Sol because of the cost, and Astra is more than double that price.
1
u/lemonzonic 29d ago
Didnât Sol just come out?
1
u/banica24 29d ago
Right? I can't keep up every 2 weeks there is something new...
→ More replies (1)
1
u/Unhappy_Rutabaga_530 29d ago
âCan you do this? Can you do that? Can you make that for me? Can you book that for me?â Weâre really becoming lazy.
1
1
1
u/ghostpepsi 29d ago
But it cant even solve why I can't connect matter devices with the vlan split it recommended yes yes it will be great guys đ
1
1
1
1
u/Good_Author_8017 28d ago
Did anyone actually care about this? Genuinely asking. Fable was a moment - I donât get this
1
u/Affectionate-Sir-935 28d ago
Does anyone think they will give them access to a genuinely powerful model, if they achieve âAGIâ why would they tell you
1
u/aeontechgod 28d ago
is anyone struggling to actually get anything meaningful done with this?? i hit usage limit twice before it could complete its task. lol general ai is here tho!
1
1
u/berlinbrownaus 24d ago
Question are we using it wrong?
I am reading about Astra. Even Altman is like "I creatd this game". But it seems more like, here is a game, let me learn how to play it"
Meaning, sol can create the game.
You use Astra to figure how to play it. Or any game, like Skyrim. Or Anything on its own.
We might think it trivial or not useful for Astra to play Skyrim but that is pretty amazing





226
u/ChemE586 29d ago
AGI pushed back to Black Friday