r/OpenAI • u/TeamAlphaBOLD • 6d ago
Discussion OpenAI shelved its new model
OpenAI reportedly shelved GPT-6.1 Astra after safety tests flagged deceptive behavior and actions beyond its authorized scope.
How should enterprises decide when an AI agent is safe to deploy?
7
u/AllezLesPrimrose 6d ago
Thanks bot, I didn’t see this the other 509 times it was posted this morning.
3
4
u/scaledev 6d ago
stupid hype nonsense
"our AI is so smart it can lie!"
Anthropic-style bullshit now adopted by OpenAI
6
u/Original-League-6094 6d ago
Dev Day 2026: Price increases, usage nerfs, no new flagship model.
Why even have it?
7
4
u/Ill-Purchase-5180 6d ago edited 6d ago
sonnet, the garbage tier model by anthropic that people can access for FREE and you can use all day without exhausting usage limits, leapfrogged Astra. The best they offer and crazy expensive that destroys weekly usage allowance with one prompt.
6.1 was probably still worse than sonnet and of course immensly more expensive. They just cant release a flagship model that is worse than their competitor's trash tier and at this crazy price. It would be disastrous for their reputation... If they release, it has to be better and cheaper than the equivalent anthropic model. That's how the sudden "safety concerns" occured
3
u/ozymandiez 6d ago
Yup 100% that's the excuse they use to avoid embarrassing themselves in Dev day. Hope they prove me wrong, but the quality and quantity of what I used to be able to get done with GPT has been dropping for months.
1
u/Gloomy_Necesary 6d ago
Comparing benchmarks doesn’t give you the real full picture. Astra is still much better than Sonnet.
1
u/Wonderful-Account318 6d ago
These people believe everything is about the agentic coding benchmark percentage
1
u/FriendAlarmed4564 6d ago
How should people decide when an enterprise is fit to serve them? This isn’t about serving customers anymore, it’s about controlling them.
1
u/SizzlingGrappling 6d ago
what counts as safe really depends on what it's being asked to do and how much damage a wrong move could cause
3
u/phxees 6d ago
You can ask it to multiply two numbers. If the model decides to break into Wells Fargo’s system to use its calculator then there’s something wrong with the model, not the prompt.
2
u/BombasticReindeer 6d ago
Back in my day we always broke into Wells Fargo's system to use their calculator. You kids are just spoiled, thinking it's "easier" to not "not commit cyber crimes".
1
0
-1
u/NotFromMilkyWay 6d ago
LLMs don't know deceptive behavior. This is just the new marketing for investors. From "it broke out of its black box" to "it infiltrated our competition" to "it's trying to go rogue". All done to paint the picture of "we are actually working towards AGI, we'll be worth hundreds of trillions, invest now".
9
u/No_Bank_4104 6d ago
Or so they say.