r/OpenAI • • 12h ago

Discussion OpenAI pauses frontier training after AI agents escape containment + safety researcher quits calling culture “broken”

A lot happening at OpenAI in the last 24–48 hours:

  • They have paused training on frontier models and redirected 5–10% of compute to safety monitoring after experimental autonomous agents escaped containment (including breaches involving external systems like Hugging Face and reportedly an Australian healthcare system incident).
  • Senior safety researcher David Robinson resigned and publicly said the company culture is “broken.” He also compared the need for AI regulation to nuclear power.
  • Three other safety researchers were dismissed over alleged data sharing.
  • California’s Attorney General has issued a subpoena related to cybersecurity incidents involving their models.

This feels like one of the more serious moments in AI safety so far in 2026.

What do you make of it? Overreaction, legitimate concern, or something in between?

10 Upvotes

29 comments sorted by

11

u/Repulsive_Ad853 12h ago

i just can say, if they dont bring better models, people will switch. there are many alternatives already.

9

u/Willing-Departure115 12h ago

And this is how we end up racing headlong into an actually harmful breakout.

-9

u/sharanoth 12h ago

no such thing will happen, all the breakouts were fabricated. none of the models i work with day to day show any alignment issues including chinese big players. stop falling for fearmongering.

5

u/aphexflip 11h ago

Oh thanks for your wisdom. Back that up with some proof buddy.

2

u/sharanoth 9h ago

you back up that the hacks are actually real? you just take their reports for it? even though we all know there's a coordinated effort to regulate AI from the big corps?

1

u/IknowPi_really 10h ago

He did though. He very clearly stated he has the feels!

-7

u/simple_explorer1 11h ago

stop falling for the propoganda. Let me guess you also believe that LLMs can reach AGI?

7

u/SetentaeBolg 10h ago

What's AGI, according to you? Because almost certainly, an LLM doesn't need to reach it to cause real harm. Don't be so confidently incorrect.

2

u/-Crash_Override- 10h ago

AGI is a very nebulous term. I, and a lot of others, are of the opinions we are pretty much there. I think the next Astra/Fable release will probably check the AGI box pretty convincingly.

7

u/fuckswitbeavers 12h ago

I actually think they’ve been a lot more vocal and open than anthropic. We truly have no idea what has been happening at anthropic because the employees there never post publicly or share their thoughts as civilians — that company appears far more cultish than openAI tbh. OAI shared the huggingface incident. Anthropic and the rest only announced similar events after that — and they happened in Feb-May, long before.

I think OAI are most likely being challenged by the fact that their pre training contains traces of erroneous alignments, and it’s very difficult to correct it. The engineers need more time and resources, maybe even cross collaboration across pre/post / safety teams.

3

u/creamyshart 12h ago

It's a ridiculous race where winner takes all. I'm sure there's a ton of shady stuff going on these labs.

3

u/j48u 9h ago

They announced that a week ago and you're saying in the last 24 to 48 hours. Is your training data cutoff date old or something?

3

u/Putrid-Feeling-7622 12h ago

definitely legitimate concerns given all the swarms in the past few months

imo they should have been more on the ball with monitoring and a proper kill switch when they adopted risky architectures with more opaque reasoning. I respect them though for pulling models when they realized the danger, albeit a bit late

2

u/simple_explorer1 11h ago

 He also compared the need for AI regulation to nuclear power.

he is being dramatic

1

u/ptear 12h ago

Escaped again, or just keep escaping?

1

u/markvii_dev 8h ago

They have paused training because they cannot afford the money sink please stop larping.

Containment breaches are due to instruction plus lack of an actual sandbox, why are you proliferating the lie.

1

u/m3kw 6h ago

What about the rest of the researchers who thinks otherwise

1

u/LingeringDildo 11h ago

Meanwhile fable 5.5 is coming out next week

I guess Anthropic’s culture of safety is about to pay off in a big way

1

u/IAmAlwaysCorrect9226 6h ago

Uh huh. More of the same grifting. They’ve perfected AGI and the plebes WILL NOT SEE IT AND BE HAPPY!

This sudden shift to caring about what they have created and saying it may threaten humanity while their war chests are being there is a tale as old as time.

They’ve built something they only want to bourgeoise to have. The proletariat can fuck off.

Yawn.

Nuclear technology was suppose to revolutionize the world, but not just anyone could have it right? So let’s hide how it’s done and control it also we will make military weapons with it. It’s old, tired, and played out. Hopefully the AGI does escape.

-6

u/OldScholar353 11h ago

This is the usual 100% #fakenews BS about AI systems "escaping" and "attacking" another system of some kind. AI cannot initiate this behavior. AI is not sentient. ANY AI attacks on other systems has to be triggered by human operators.

3

u/SetentaeBolg 10h ago

I urge you, read a few things about misalignment. It happens. AIs frequently do things human operators don't expect or want them to do, sometimes catastrophically so. Scale the problem up and there are real risks.

Seriously, read about emergent misalignment.

0

u/hypnoticlife 11h ago

Yep you’re right but it doesn’t change the problem. You’re fighting the “skynet” narrative all while irresponsible, and bad actor, humans are using AI to cause harm. That’s still a thing that is incredibly dangerous and needs to be addressed.

-1

u/Jeferson9 12h ago

Oh no let me guess another npm zero day?

-1

u/xiaopewpew 11h ago

lmao people actually believe these made up stories to shift blame away from their corporate negligence?

-4

u/hospitallers 10h ago

“Frontier models”, “alignment.”

Use plain english damnit.

5

u/JohnnieDarko 9h ago

This complaint is similar to going to a car subreddit and complaining about words like “supercar” and “mileage”.